feat(classifier): add multilingual ECP inherence classifier POC

This commit is contained in:
2026-08-20 00:51:02 -03:00
parent d371b81aa4
commit 67cc40f91a
175 changed files with 30399 additions and 703 deletions
@@ -0,0 +1,61 @@
# POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
**Purpose**: Validate requirement quality, clarity, and completeness for the POC implementation before code execution
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md) | [plan.md](../plan.md)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
---
## 1. Requirement Completeness & Scope Boundaries
- [x] CHK001 Are the required and optional fields of the ECP Snapshot explicitly specified with defaults? [Completeness, Spec §FR-003]
- [x] CHK002 Is the CLI execution contract completely specified with argument names, flags, and file path requirements? [Completeness, Spec §FR-002]
- [x] CHK003 Are explicit out-of-scope boundaries (no REST API, no live Neo4j, no worker queues) clearly stated to prevent scope creep? [Scope, Spec §Assumptions & Scope]
- [x] CHK004 Is the behavior for missing `--output` (printing to stdout) explicitly defined? [Completeness, Spec §FR-002]
---
## 2. Requirement Clarity & Decision Semantics
- [x] CHK005 Are the 4 decision categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) defined with unambiguous conditions? [Clarity, Spec §FR-005]
- [x] CHK006 Is the deterministic derivation rule for `is_inherent` (`true` for DIRECT/CONTEXTUAL, `false` for TANGENTIAL/NOT_RELATED) mathematically unambiguous? [Clarity, Spec §FR-006]
- [x] CHK007 Is the 3-tier hybrid execution order explicitly defined such that clear cases never invoke Tier 2/Tier 3? [Clarity, Spec §FR-004]
- [x] CHK008 Are all standardized error codes (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`) enumerated? [Clarity, Spec §FR-008]
---
## 3. Requirement Consistency & Alignment
- [x] CHK009 Do the example filenames in `spec.md`, `plan.md`, `quickstart.md`, and `tasks.md` align consistently without discrepancies? [Consistency, Plan §Project Structure]
- [x] CHK010 Is the benchmark test file name (`tests/test_benchmark_24.py`) consistent across all documentation artifacts? [Consistency, Quickstart §3]
- [x] CHK011 Are dependencies strictly limited to `pytest` without mandatory external AI/cloud packages? [Consistency, Plan §Technical Context]
---
## 4. Acceptance Criteria & Measurability
- [x] CHK012 Is the POC benchmark suite quantified with an explicit test count (24 cases = 6 languages × 4 decision types)? [Measurability, Spec §SC-002]
- [x] CHK013 Is the accuracy target (≥ 90% precision) bounded specifically to the controlled 24-case benchmark suite? [Measurability, Spec §SC-002]
- [x] CHK014 Is the execution latency target (< 200ms for Tier 1) testable and quantified? [Measurability, Spec §SC-004]
---
## 5. Scenario & Edge Case Coverage
- [x] CHK015 Are requirements specified for malformed Markdown or empty content files? [Edge Cases, Spec §FR-008]
- [x] CHK016 Are requirements specified for corrupted JSON or missing required ECP fields? [Edge Cases, Spec §FR-008]
- [x] CHK017 Are negative anchor suppression rules specified for homonym disambiguation? [Coverage, Spec §FR-005]
- [x] CHK018 Are graph snapshot relationship matches (weight, scope, relation_type) accounted for in contextual inherence decisions? [Coverage, Spec §FR-003, §FR-005]
---
## Notes
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied.
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation.
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers.
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`.
@@ -0,0 +1,40 @@
# Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-08-19
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- Initial requirements documented from user audio brief and refined with POC constraints.
- ECP defined as Entity Context Profile with materialized JSON snapshots.
- Tier 1 (deterministic rules) defined as mandatory core; Tier 2 (embeddings) and Tier 3 (LLM) defined as optional/configurable (LLM off by default, no API key required).
- Decision rules explicitly codified for all 4 categories and derived `is_inherent` boolean.
- Success criteria scoped to a controlled 24-case POC benchmark suite (6 languages × 4 decision types).
- Structured error codes defined (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`).
- Explicit out-of-scope declarations enforced (No REST API, No package publishing, No queues/workers, No runtime DB/Neo4j).
@@ -0,0 +1,42 @@
# CLI Contract & Interface Specification (POC)
**Feature**: `001-multilingual-entity-classifier`
**Status**: Completed
---
## 1. Command Line Interface
```bash
python classify.py --ecp <path-to-ecp.json> --content <path-to-content.md> [--output <path-to-result.json>] [--enable-embeddings] [--enable-llm]
```
### 1.1 Arguments & Options
| Parameter | Type | Required | Description |
|---|---|---|---|
| `--ecp` | Path (`string`) | **Yes** | Absolute or relative path to the ECP Snapshot JSON file |
| `--content` | Path (`string`) | **Yes** | Absolute or relative path to the Markdown content file |
| `--output`, `-o` | Path (`string`) | No | Destination path to write result JSON. If omitted, prints JSON to `stdout`. |
| `--enable-embeddings` | Flag (`bool`) | No | Enables Tier 2 local multilingual vector similarity adapter (default: false). |
| `--enable-llm` | Flag (`bool`) | No | Enables Tier 3 LLM fallback adapter for ambiguous cases (default: false). |
| `--version`, `-v` | Flag (`bool`) | No | Displays version and exits. |
| `--help`, `-h` | Flag (`bool`) | No | Displays help message. |
---
## 2. Standard Streams & Exit Codes
### 2.1 Exit Codes
- `0`: Success (classification completed normally, result JSON written to file or stdout).
- `1`: Validation / Processing Error (invalid arguments, malformed input, missing fields, structured error JSON printed to stderr or output file).
### 2.2 Standard Output (`stdout`) / Standard Error (`stderr`)
- If `--output` is provided:
- Success result is saved to the specified file.
- On error, error JSON is written to the output file (if accessible) and emitted to `stderr`.
- If `--output` is NOT provided:
- On success: formatted JSON is printed directly to `stdout`.
- On error: structured error JSON is printed to `stderr`.
@@ -0,0 +1,205 @@
# Data Models & Schemas (POC)
**Feature**: `001-multilingual-entity-classifier`
**Status**: Completed
---
## 1. Input Schemas
### 1.1 ECP Snapshot Schema (`snapshot.json`)
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "ECPSnapshot",
"type": "object",
"required": [
"target_entity_id",
"target_name",
"aliases",
"domain",
"anchors"
],
"properties": {
"target_entity_id": {
"type": "string",
"description": "Unique identifier of the entity"
},
"target_name": {
"type": "string",
"description": "Canonical name of the entity"
},
"aliases": {
"type": "array",
"items": { "type": "string" },
"description": "Multilingual aliases, acronyms, and trade names"
},
"domain": {
"type": "string",
"description": "Primary industry, category, or domain of the entity"
},
"anchors": {
"type": "array",
"items": { "type": "string" },
"description": "Key domain concepts, topics, and contextual keywords"
},
"negative_anchors": {
"type": "array",
"items": { "type": "string" },
"default": [],
"description": "Homonym disambiguators or exclusion terms"
},
"graph_version": {
"type": "string",
"default": "1.0.0",
"description": "Version timestamp or hash of the source graph"
},
"related_entities": {
"type": "array",
"default": [],
"items": {
"type": "object",
"required": ["entity_id", "name", "relation_type", "weight"],
"properties": {
"entity_id": { "type": "string" },
"name": { "type": "string" },
"aliases": { "type": "array", "items": { "type": "string" }, "default": [] },
"relation_type": { "type": "string" },
"weight": { "type": "number", "minimum": 0.0, "maximum": 1.0 },
"scope": { "type": "string", "default": "general" },
"confidence": { "type": "number", "minimum": 0.0, "maximum": 1.0, "default": 1.0 }
}
}
}
}
}
```
### 1.2 Content Item Schema (`content.md`)
- **Format**: Plain Markdown document (`.md` or `.txt`).
- **Validation Rules**:
- Must not be empty (minimum 5 non-whitespace characters).
- Must contain readable text in UTF-8 encoding.
---
## 2. Output Schemas
### 2.1 Classification Success Result Schema (`result.json`)
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "ClassificationResult",
"type": "object",
"required": [
"decision",
"is_inherent",
"confidence",
"detected_language",
"matched_anchors",
"negative_matches",
"graph_matches",
"evidence",
"rationale",
"warnings"
],
"properties": {
"decision": {
"type": "string",
"enum": [
"DIRECT_INHERENT",
"CONTEXTUAL_INHERENT",
"TANGENTIAL",
"NOT_RELATED"
]
},
"is_inherent": {
"type": "boolean",
"description": "Derived: true if DIRECT_INHERENT or CONTEXTUAL_INHERENT; false if TANGENTIAL or NOT_RELATED"
},
"confidence": {
"type": "number",
"minimum": 0.0,
"maximum": 1.0,
"description": "Normalized confidence score of the classification"
},
"detected_language": {
"type": "string",
"description": "Detected ISO-639-1 code (pt, en, es, de, it, fr, etc.)"
},
"matched_anchors": {
"type": "array",
"items": { "type": "string" },
"description": "Entity aliases and direct anchors identified in text"
},
"negative_matches": {
"type": "array",
"items": { "type": "string" },
"description": "Negative anchors found in text"
},
"graph_matches": {
"type": "array",
"items": {
"type": "object",
"properties": {
"entity_id": { "type": "string" },
"name": { "type": "string" },
"relation_type": { "type": "string" },
"weight": { "type": "number" }
}
},
"description": "Related entities from snapshot matched in text"
},
"evidence": {
"type": "array",
"items": { "type": "string" },
"description": "Extracted textual snippets from Markdown justifying the decision"
},
"rationale": {
"type": "string",
"description": "Concise explanation of the classification verdict"
},
"warnings": {
"type": "array",
"items": { "type": "string" },
"description": "Non-fatal warnings encountered during processing"
}
}
}
```
### 2.2 Error Result Schema
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "ClassificationError",
"type": "object",
"required": [
"error_code",
"message",
"details"
],
"properties": {
"error_code": {
"type": "string",
"enum": [
"invalid_ecp_json",
"invalid_markdown",
"unsupported_language",
"empty_content",
"missing_required_field"
]
},
"message": {
"type": "string"
},
"details": {
"type": "object"
}
}
}
```
@@ -0,0 +1,128 @@
# Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
**Branch**: `001-multilingual-entity-classifier` | **Date**: 2026-08-19 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `specs/001-multilingual-entity-classifier/spec.md`
---
## Summary
Implement a lightweight, standalone Python CLI tool (`classify.py`) that evaluates whether a Markdown document is inherent to a target entity defined by an Entity Context Profile (ECP Snapshot JSON). The architecture implements a Tier 1 deterministic matching engine for 6 core languages (PT, EN, ES, DE, IT, FR), with clean optional adapter hooks for Tier 2 (embeddings) and Tier 3 (LLM fallback), fully validated against a controlled 24-case benchmark suite.
---
## Technical Context
**Language/Version**: Python 3.10+ (standard library for Tier 1 core execution).
**Primary Dependencies**:
- Core: standard library (`json`, `re`, `argparse`, `pathlib`, `typing`).
- Testing & Validation (mandatory): `pytest>=7.0` (only required dependency in `requirements.txt`).
- Optional (Tier 2 / Tier 3 adapters): `sentence-transformers`, `httpx` / `openai` (strictly optional, disabled by default).
**Storage**: None (file-in / file-out via CLI, no database).
**Testing**: `pytest` running unit tests and the 24-case controlled benchmark suite (`6 languages × 4 decision types`) via `tests/test_benchmark_24.py`.
**Target Platform**: Cross-platform (Windows / Linux / macOS).
**Project Type**: Standalone CLI script & modular core library (`classify.py` + `src/`).
**Performance Goals**: < 200ms execution time for Tier 1 deterministic evaluation on documents < 2,000 words.
**Constraints**:
- Zero mandatory external network calls or cloud API keys required to execute the POC or pass tests.
- Zero server/API dependencies (no FastAPI, no workers, no queues).
- Decoupled from live graph databases (consumes pre-materialized ECP JSON snapshots).
**Scale/Scope**: POC scope with minimum 24 explicit benchmark test fixtures.
---
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle / Gate | Status | Notes |
|---|---|---|
| **I. Library / Script First** | **PASS** | Standalone Python module structure, easily importable and script-callable. |
| **II. CLI Interface** | **PASS** | Clean standard CLI: `python classify.py --ecp <json> --content <md> --output <json>`, supports file paths and stdout output. |
| **III. Test-First (TDD)** | **PASS** | 24-case controlled benchmark matrix defined before code implementation. |
| **IV. Simplicity & YAGNI** | **PASS** | No premature REST API, no DB, no worker queues, no heavy frameworks. |
---
## Project Structure
### Documentation (this feature)
```text
specs/001-multilingual-entity-classifier/
├── spec.md # Feature specification
├── plan.md # Implementation plan (this file)
├── research.md # Technical research & decisions
├── data-model.md # Schemas & data contracts
├── quickstart.md # Validation & usage guide
├── checklists/
│ └── requirements.md # Quality checklist
├── contracts/
│ └── cli-contract.md # CLI input/output contract
└── tasks.md # Implementation tasks (/speckit-tasks command)
```
### Source Code (repository root)
```text
classify.py # Main CLI entrypoint script
src/
├── __init__.py
├── models.py # Dataclasses & schema validators (ECPSnapshot, Result, Error)
├── language.py # Lightweight multilingual detector & normalizer (6 languages)
├── parser.py # Markdown content parser & excerpt extractor
├── classifier.py # Core classification engine & decision logic (Tier 1 core)
└── adapters/
├── __init__.py
├── base.py # Base abstract adapter interfaces
├── embeddings.py # Optional Tier 2 embeddings adapter (disabled by default)
└── llm.py # Optional Tier 3 LLM fallback adapter (disabled by default)
examples/
├── ecp_petrobras.json # Example ECP Snapshot (PT)
├── ecp_volkswagen.json # Example ECP Snapshot (DE)
├── ecp_apple.json # Example ECP Snapshot (ES/EN)
├── content_presal_pt.md # Example Markdown Content (PT)
├── content_northvolt_de.md # Example Markdown Content (DE)
└── content_tangential_es.md # Example Markdown Content (ES)
tests/
├── __init__.py
├── test_cli.py # CLI argument parsing, flags, file I/O, stdout emission, exit codes
├── test_models.py # ECP snapshot parsing and structured error handling
├── test_language.py # Language detection and text normalization tests
├── test_classifier.py # Unit tests for decision rules (DIRECT, CONTEXTUAL, TANGENTIAL, NOT_RELATED)
├── test_benchmark_24.py # Controlled 24-case benchmark runner (6 languages x 4 decisions)
└── fixtures/
└── benchmark_24/ # 24 paired test cases (ecp_*.json + content_*.md + expected_*.json)
├── pt/
├── en/
├── es/
├── de/
├── it/
└── fr/
```
**Structure Decision**: Single modular project with root CLI `classify.py` and clear separation of models, language normalization, deterministic classification, and optional adapter stubs under `src/`.
---
## Complexity Tracking
> No constitution violations detected. Design enforces absolute simplicity and strict POC constraints.
| Component | Choice | Simpler Alternative Rejected Because |
|---|---|---|
| CLI vs REST API | Standalone CLI | REST API adds unnecessary network latency, FastAPI dependencies, and server management for a POC. |
| Deterministic Core vs Pure LLM | Deterministic Tier 1 Core | Pure LLM is expensive, non-deterministic, slow, and requires mandatory external API keys. |
| In-Memory Snapshot vs Live Graph Query | Materialized JSON Snapshot | Live Neo4j queries couple classifier runtime to external infrastructure and break test reproducibility. |
@@ -0,0 +1,109 @@
# Quickstart & Validation Guide (POC)
**Feature**: `001-multilingual-entity-classifier`
**Status**: Completed
---
## 1. Prerequisites & Installation
- Python 3.10+
- Virtual environment (optional, standard library only for Tier 1):
```bash
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/Mac
pip install pytest
```
---
## 2. Basic CLI Usage Examples
### 2.1 Direct Inherence (Portuguese)
```bash
python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/petrobras_result.json
```
**Expected Output (`out/petrobras_result.json`)**:
```json
{
"decision": "DIRECT_INHERENT",
"is_inherent": true,
"confidence": 0.95,
"detected_language": "pt",
"matched_anchors": ["Petrobras", "Petróleo Brasileiro S.A.", "pré-sal"],
"negative_matches": [],
"graph_matches": [],
"evidence": ["...a Petrobras anunciou ampliação da produção na camada pré-sal..."],
"rationale": "Direct match of target entity aliases with high contextual anchor density.",
"warnings": []
}
```
### 2.2 Contextual Inherence via Graph Snapshot (German)
```bash
python classify.py --ecp examples/ecp_volkswagen.json --content examples/content_northvolt_de.md
```
**Expected Output (`stdout`)**:
```json
{
"decision": "CONTEXTUAL_INHERENT",
"is_inherent": true,
"confidence": 0.88,
"detected_language": "de",
"matched_anchors": [],
"negative_matches": [],
"graph_matches": [
{
"entity_id": "ent_northvolt",
"name": "Northvolt",
"relation_type": "SUPPLIER_OF",
"weight": 0.85
}
],
"evidence": ["...Northvolt liefert Batteriezellen für europäische Elektrofahrzeuge..."],
"rationale": "Matched connected entity Northvolt from ECP snapshot with strong supplier relationship to Volkswagen.",
"warnings": []
}
```
### 2.3 Tangential Mention (Spanish)
```bash
python classify.py --ecp examples/ecp_apple.json --content examples/content_tangential_es.md
```
**Expected Output**:
```json
{
"decision": "TANGENTIAL",
"is_inherent": false,
"confidence": 0.35,
"detected_language": "es",
"matched_anchors": ["Apple"],
"negative_matches": [],
"graph_matches": [],
"evidence": ["...el dilema fue como la manzana de la discordia en la reunión..."],
"rationale": "Single passing mention without supporting tech domain anchors or entity context.",
"warnings": ["Low contextual density for target entity."]
}
```
---
## 3. Running the Controlled 24-Case Benchmark
Execute the automated test suite measuring classification precision across all 6 languages and 4 decision types:
```bash
pytest tests/test_benchmark_24.py -v
```
**Benchmark Matrix (6 × 4 = 24 test pairs)**:
- 🇧🇷 **Portuguese (`pt`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
- 🇺🇸 **English (`en`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
- 🇪🇸 **Spanish (`es`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
- 🇩🇪 **German (`de`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
- 🇮🇹 **Italian (`it`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
- 🇫🇷 **French (`fr`)**: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`
@@ -0,0 +1,61 @@
# Technical Research & Architecture Decisions (POC)
**Feature**: `001-multilingual-entity-classifier`
**Status**: Completed
---
## 1. Technical Decisions & Tradeoffs
### Decision 1: Execution Engine & CLI Architecture
- **Decision**: Standalone Python 3 script (`classify.py`) with clean modular components under `src/`.
- **Rationale**: Keeps the POC lightweight, zero-boilerplate, directly testable via standard command line and `pytest`, perfectly aligned with Ponytail and SpecKit principles.
- **Alternatives Considered**:
- *FastAPI REST Service*: Rejected (violates POC simplicity, introduces server overhead, unnecessary network latency for batch/CLI evaluation).
- *Publishable Package / Setuptools*: Rejected (premature abstraction before core classification logic is validated).
### Decision 2: Multilingual Language Detection & Normalization (Tier 1 Core)
- **Decision**: Lightweight regex + heuristic n-gram / stopword profile detection for the 6 core languages (PT, EN, ES, DE, IT, FR), with fallback to standard library / optional `langdetect` or `lingua` if installed. Text normalization converts diacritics/casing for robust deterministic matching while preserving original excerpt offsets for `evidence`.
- **Rationale**: Guarantees zero heavy mandatory dependencies for basic Tier 1 execution, sub-millisecond detection latency, and 100% offline capability.
- **Alternatives Considered**:
- *Heavy Transformer-based language ID*: Rejected (large model download, excessive latency for short/medium texts).
### Decision 3: Materialized ECP Snapshot Contract & Matching Logic
- **Decision**: Parse self-contained ECP Snapshot JSON containing:
- `target_entity_id`, `target_name`, `aliases`, `domain`, `anchors`, `negative_anchors`, `graph_version`, `related_entities`.
- Matching pipeline evaluates:
1. Direct entity match: `aliases` + `target_name` in Markdown.
2. Negative anchor presence: if negative terms dominate context → downgrade or mark `NOT_RELATED`.
3. Graph snapshot match: `related_entities` found in Markdown with associated `relation_type`, `weight`, `confidence`.
4. Domain/Anchor density: computes normalized score (0.0 to 1.0) and assigns one of the 4 decision categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`).
- **Rationale**: Completely decouples inference from live graph database queries, enabling deterministic, fast, reproducible, and audit-friendly decisions.
### Decision 4: Tier 2 (Embeddings) & Tier 3 (LLM) Optional Adapters
- **Decision**: Implement Tier 2 and Tier 3 as decoupled adapter interfaces (`EmbeddingProviderInterface`, `LLMProviderInterface`).
- By default, `Tier2` and `Tier3` are disabled (`--enable-embeddings` / `--enable-llm` CLI flags).
- If Tier 1 has high confidence or unambiguous negative match, Tier 2/Tier 3 are skipped entirely.
- Zero mandatory external API keys are required to execute or test the POC.
- **Rationale**: Satisfies the 3-tier hybrid requirement without imposing heavy dependencies (e.g. PyTorch/HuggingFace) or external API keys onto the baseline test suite.
### Decision 5: Controlled 24-Case POC Benchmark Suite
- **Decision**: Construct a fixture suite of 24 controlled test cases:
- 6 Languages: `pt`, `en`, `es`, `de`, `it`, `fr`
- 4 Decision Types per language: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`
- Total: 6 × 4 = 24 benchmark pairs of `(ecp_snapshot.json, content.md, expected_result.json)`.
- **Rationale**: Provides an unambiguous, reproducible test bed to measure and verify the ≥ 90% precision requirement.
---
## 2. Standardized Error Handling Strategy
Standardized JSON error envelope emitted with non-zero exit code (1) when execution cannot complete:
```json
{
"error_code": "invalid_ecp_json | invalid_markdown | unsupported_language | empty_content | missing_required_field",
"message": "Human readable description",
"details": {
"field": "aliases",
"reason": "Field 'aliases' must be an array of strings"
}
}
```
@@ -0,0 +1,138 @@
# Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
**Feature Branch**: `001-multilingual-entity-classifier`
**Created**: 2026-08-19
**Status**: Draft / Documented (POC Scope)
**Input**: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."
---
## Clarifications
### Session 2026-08-19
- Q: O que é o ECP e qual o formato de entrada esperado? → A: **ECP significa Entity Context Profile**. A entrada para o classificador é um **ECP Snapshot materializado em formato JSON** (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato **Markdown**.
- Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um **script Python CLI** (`python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>`). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp.
- Q: Como o pipeline híbrido deve se comportar na POC? → A: O **Tier 1 (regras determinísticas)** é o núcleo obrigatório da POC. O **Tier 2 (embeddings multilíngues locais)** é opcional/configurável. O **Tier 3 (LLM fallback)** é opcional/configurável e **desativado por padrão** (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
- Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) com `is_inherent` como campo derivado booleano (`true` para `DIRECT_INHERENT` e `CONTEXTUAL_INHERENT`; `false` para `TANGENTIAL` e `NOT_RELATED`), acompanhado de `confidence`, `matched_anchors`, `negative_matches`, `graph_matches`, `evidence`, `rationale` e `warnings`.
- Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo **24 casos (6 idiomas × 4 tipos de decisão)**, medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.
---
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)
As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.
**Why this priority**: Core value proposition and execution model of the POC.
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json` and verify the output contains `decision: "DIRECT_INHERENT"` and `is_inherent: true`.
**Acceptance Scenarios**:
1. **Given** a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", **When** `classify.py` executes, **Then** it produces `decision: "DIRECT_INHERENT"`, `is_inherent: true`, and matched aliases in `matched_anchors`.
2. **Given** a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, **When** evaluated, **Then** it produces `decision: "CONTEXTUAL_INHERENT"`, `is_inherent: true`, and "Northvolt" listed in `graph_matches`.
3. **Given** a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, **When** evaluated against an ECP Snapshot for "Apple Inc.", **Then** it produces `decision: "TANGENTIAL"`, `is_inherent: false`.
4. **Given** an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", **When** evaluated, **Then** it produces `decision: "NOT_RELATED"`, `is_inherent: false`.
---
### User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)
As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.
**Why this priority**: Guarantees baseline reliability and debuggability across all target languages.
**Independent Test**: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.
**Acceptance Scenarios**:
1. **Given** valid Markdown documents in each of the 6 languages, **When** classified against matching ECP Snapshots, **Then** the system correctly identifies `detected_language` and applies language-aware tokenization/matching.
2. **Given** a corrupted JSON file or empty Markdown file, **When** `classify.py` executes, **Then** it outputs a structured JSON error response with appropriate error codes (`invalid_ecp_json`, `empty_content`, etc.) and non-zero exit code.
---
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- **FR-002**: System MUST provide a standalone Python CLI entrypoint:
```bash
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>
```
If `--output` is omitted, the result JSON MUST be emitted to standard output (`stdout`). Both `--ecp` and `--content` MUST be valid file paths.
- **FR-003**: System MUST parse and validate the materialized ECP Snapshot JSON structure:
- **Required fields**: `target_entity_id` (string), `target_name` (string), `aliases` (array of strings), `domain` (string), `anchors` (array of strings).
- **Optional fields with defaults**: `negative_anchors` (array of strings, default `[]`), `graph_version` (string, default `"1.0.0"`), `related_entities` (array of objects, default `[]`).
- **Related entity structure**: `entity_id` (string), `name` (string), `relation_type` (string), `weight` (number), `aliases` (array of strings, default `[]`), `scope` (string, default `"general"`), `confidence` (number, default `1.0`).
- **FR-004**: System MUST execute a 3-tier hybrid classification pipeline:
- **Tier 1 (Deterministic Rules - Mandatory in POC)**: Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
- **Tier 2 (Local Multilingual Embeddings - Optional/Configurable)**: Semantic similarity scoring using local embeddings (can be toggled via config/flag).
- **Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default)**: Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
- **FR-005**: System MUST enforce the following decision logic:
- `DIRECT_INHERENT`: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.
- `CONTEXTUAL_INHERENT`: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.
- `TANGENTIAL`: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.
- `NOT_RELATED`: Absence of relevant signals or dominant negative anchors.
- **FR-006**: System MUST derive the boolean `is_inherent` strictly from `decision`:
- `DIRECT_INHERENT` → `is_inherent: true`
- `CONTEXTUAL_INHERENT` → `is_inherent: true`
- `TANGENTIAL` → `is_inherent: false`
- `NOT_RELATED` → `is_inherent: false`
- **FR-007**: System MUST output a success JSON containing:
- `decision`: String (`DIRECT_INHERENT` | `CONTEXTUAL_INHERENT` | `TANGENTIAL` | `NOT_RELATED`)
- `is_inherent`: Boolean (`true` | `false`)
- `confidence`: Float (`0.0` to `1.0`)
- `detected_language`: String (ISO language code)
- `matched_anchors`: Array of strings
- `negative_matches`: Array of strings
- `graph_matches`: Array of objects/strings (matched related entities from snapshot)
- `evidence`: Array of strings (excerpts extracted from the Markdown)
- `rationale`: String (short explanation of the decision)
- `warnings`: Array of strings
- **FR-008**: System MUST output structured error responses for failure conditions:
- `error_code`: Enum (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`)
- `message`: Human-readable error description
- `details`: Object with debugging details
---
### Key Entities *(data models & domain entities)*
- **Entity Context Profile (ECP) Snapshot (JSON)**: Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
- **Content Item (Markdown)**: The input markdown document payload to be analyzed.
- **Inherence Assessment Result (JSON)**: The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
- **Classification Error (JSON)**: Structured failure payload with standardized error code.
---
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
- **SC-002**: Classification precision reaches at least **90% over a controlled POC benchmark suite of 24 cases** (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`).
- **SC-003**: 100% of outputs conform strictly to the specified JSON success/error schemas.
- **SC-004**: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.
---
## Assumptions & Scope
### Assumptions
- The 6 core target languages are Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
- Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.
### Explicit Out of Scope (POC)
- REST API / FastAPI microservices.
- PyPI package creation/distribution.
- Message queues (RabbitMQ, SQS, Celery) or background workers.
- Direct runtime database or graph database connections (Neo4j).
- Integration with ContentMachineApp.
- Mandatory external LLM API keys.
@@ -0,0 +1,113 @@
# Tasks: Multilingual NLP Entity Inherence Classifier (POC)
**Feature**: `001-multilingual-entity-classifier` | **Spec**: [spec.md](./spec.md) | **Plan**: [plan.md](./plan.md)
---
## Phase 1: Setup (Shared Infrastructure)
**Purpose**: Project initialization, directory structure, and environment configuration
- [x] T001 Create project directories (`src/`, `src/adapters/`, `examples/`, `tests/`, `tests/fixtures/benchmark_24/`)
- [x] T002 [P] Create `requirements.txt` with minimal test dependencies (`pytest>=7.0`) and optional dependencies commented
- [x] T003 [P] Create `.gitignore` validation and project configuration in `pyproject.toml` or `setup.cfg`
---
## Phase 2: Foundational (Blocking Prerequisites)
**Purpose**: Core data models, schema validation, and text processing utilities required by all user stories
**CRITICAL**: No classification engine logic can run until this phase is complete
- [x] T004 Implement data models and schema validation in `src/models.py` (`ECPSnapshot`, `ClassificationResult`, `ClassificationError`, and decision enums)
- [x] T005 [P] Implement Markdown parsing and evidence snippet extractor in `src/parser.py`
- [x] T006 [P] Implement lightweight multilingual language detector & diacritic normalizer for the 6 languages in `src/language.py`
- [x] T007 Implement unit tests for models, parser, and language detector in `tests/test_models.py` and `tests/test_language.py`
**Checkpoint**: Foundation ready - models, language detection, and Markdown parsing validated.
---
## Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP]
**Goal**: Deliver a functioning CLI (`classify.py`) executing Tier 1 deterministic evaluation (aliases, anchors, negative anchors, graph matches) on ECP Snapshots and Markdown documents.
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/test.json` and verify `decision: "DIRECT_INHERENT"`, `is_inherent: true`, confidence, and evidence.
### Tests for User Story 1
- [x] T008 [P] [US1] Create CLI contract and execution tests in `tests/test_cli.py` (testing `--ecp`, `--content`, `--output`, stdout fallback, and exit codes)
- [x] T009 [P] [US1] Create deterministic decision logic unit tests in `tests/test_classifier.py` for `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, and `NOT_RELATED`
### Implementation for User Story 1
- [x] T010 [US1] Implement Tier 1 deterministic matching engine in `src/classifier.py` (alias matching, negative anchor suppression, graph node resolution, scoring algorithm, and derived `is_inherent` boolean)
- [x] T011 [US1] Implement main CLI script `classify.py` handling argument parsing (`--ecp`, `--content`, `--output`, `--enable-embeddings`, `--enable-llm`), file I/O, error formatting, and stdout output
- [x] T012 [P] [US1] Create baseline example files: `examples/ecp_petrobras.json`, `examples/content_presal_pt.md`, `examples/ecp_volkswagen.json`, `examples/content_northvolt_de.md`, `examples/ecp_apple.json`, `examples/content_tangential_es.md`
- [x] T013 [US1] Verify end-to-end execution of CLI on example files and validate output JSON conforms strictly to schema
**Checkpoint**: User Story 1 complete! Standalone CLI works end-to-end locally with Tier 1 deterministic classification.
---
## Phase 4: User Story 2 - Multilingual 24-Case Controlled Benchmark Suite (Priority: P2)
**Goal**: Provide 24 paired test fixtures across all 6 languages (PT, EN, ES, DE, IT, FR) and 4 decision types, measuring ≥ 90% precision.
**Independent Test**: Run `pytest tests/test_benchmark_24.py -v` and achieve 100% pass rate across the 24 controlled benchmark fixtures.
### Implementation for User Story 2
- [x] T014 [P] [US2] Create Portuguese benchmark fixtures (4 cases: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`) in `tests/fixtures/benchmark_24/pt/`
- [x] T015 [P] [US2] Create English benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/en/`
- [x] T016 [P] [US2] Create Spanish benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/es/`
- [x] T017 [P] [US2] Create German benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/de/`
- [x] T018 [P] [US2] Create Italian benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/it/`
- [x] T019 [P] [US2] Create French benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/fr/`
- [x] T020 [US2] Implement parameterized benchmark runner in `tests/test_benchmark_24.py` verifying precision, score thresholds, evidence extraction, and language detection across all 24 cases
**Checkpoint**: User Story 2 complete! Multilingual robustness validated across 24 test cases with ≥ 90% precision.
---
## Phase 5: Optional Adapters - Tier 2 (Embeddings) & Tier 3 (LLM) (Priority: P3)
**Goal**: Provide clean, decoupled adapter interfaces for local embeddings and LLM fallback without introducing mandatory runtime dependencies.
**Independent Test**: Validate that adapter interfaces load correctly and remain inert/disabled by default unless explicit CLI flags are provided.
### Implementation for User Story 3
- [x] T021 [P] [US3] Create base adapter interface in `src/adapters/base.py` (`BaseNLPAdapter`)
- [x] T022 [P] [US3] Implement optional local vector similarity adapter stub in `src/adapters/embeddings.py` (activated via `--enable-embeddings`)
- [x] T023 [P] [US3] Implement optional LLM fallback adapter stub in `src/adapters/llm.py` (activated via `--enable-llm`, disabled by default)
- [x] T024 [US3] Wire adapter hooks into `src/classifier.py` ensuring Tier 1 handles clear cases without invoking Tier 2/Tier 3
**Checkpoint**: User Story 3 complete! Extensibility contracts established with zero breaking changes or mandatory cloud dependencies.
---
## Phase 6: Polish & Validation
**Purpose**: Quickstart verification, code hygiene, and documentation audit
- [x] T025 Execute all quickstart validation scenarios defined in `specs/001-multilingual-entity-classifier/quickstart.md`
- [x] T026 [P] Run full test suite (`pytest -v --tb=short`)
- [x] T027 Run `graphify update .` to index new implementation files into knowledge graph
---
## Dependencies & Execution Order
```mermaid
flowchart TD
Setup[Phase 1: Setup T001-T003] --> Foundational[Phase 2: Foundational T004-T007]
Foundational --> US1[Phase 3: US1 Core CLI & Tier 1 T008-T013]
US1 --> US2[Phase 4: US2 24-Case Benchmark T014-T020]
US1 --> US3[Phase 5: US3 Optional Adapters T021-T024]
US2 --> Polish[Phase 6: Polish T025-T027]
US3 --> Polish
```
### Implementation Strategy
1. **MVP (Phase 1 + 2 + 3)**: Delivers working `classify.py` with Tier 1 deterministic engine and example files.
2. **Benchmark Verification (Phase 4)**: Guarantees multilingual compliance on 24 controlled test fixtures.
3. **Adapter Stubs (Phase 5)**: Prepares the codebase for future embedding/LLM extensions cleanly.