# Feature Specification: Multilingual NLP Entity Inherence Classifier (POC) **Feature Branch**: `001-multilingual-entity-classifier` **Created**: 2026-08-19 **Status**: Draft / Documented (POC Scope) **Input**: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade." --- ## Clarifications ### Session 2026-08-19 - Q: O que é o ECP e qual o formato de entrada esperado? → A: **ECP significa Entity Context Profile**. A entrada para o classificador é um **ECP Snapshot materializado em formato JSON** (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato **Markdown**. - Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um **script Python CLI** (`python classify.py --ecp --content --output `). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp. - Q: Como o pipeline híbrido deve se comportar na POC? → A: O **Tier 1 (regras determinísticas)** é o núcleo obrigatório da POC. O **Tier 2 (embeddings multilíngues locais)** é opcional/configurável. O **Tier 3 (LLM fallback)** é opcional/configurável e **desativado por padrão** (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC). - Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) com `is_inherent` como campo derivado booleano (`true` para `DIRECT_INHERENT` e `CONTEXTUAL_INHERENT`; `false` para `TANGENTIAL` e `NOT_RELATED`), acompanhado de `confidence`, `matched_anchors`, `negative_matches`, `graph_matches`, `evidence`, `rationale` e `warnings`. - Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo **24 casos (6 idiomas × 4 tipos de decisão)**, medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC. --- ## User Scenarios & Testing *(mandatory)* ### User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1) As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity. **Why this priority**: Core value proposition and execution model of the POC. **Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json` and verify the output contains `decision: "DIRECT_INHERENT"` and `is_inherent: true`. **Acceptance Scenarios**: 1. **Given** a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", **When** `classify.py` executes, **Then** it produces `decision: "DIRECT_INHERENT"`, `is_inherent: true`, and matched aliases in `matched_anchors`. 2. **Given** a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, **When** evaluated, **Then** it produces `decision: "CONTEXTUAL_INHERENT"`, `is_inherent: true`, and "Northvolt" listed in `graph_matches`. 3. **Given** a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, **When** evaluated against an ECP Snapshot for "Apple Inc.", **Then** it produces `decision: "TANGENTIAL"`, `is_inherent: false`. 4. **Given** an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", **When** evaluated, **Then** it produces `decision: "NOT_RELATED"`, `is_inherent: false`. --- ### User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2) As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs. **Why this priority**: Guarantees baseline reliability and debuggability across all target languages. **Independent Test**: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes. **Acceptance Scenarios**: 1. **Given** valid Markdown documents in each of the 6 languages, **When** classified against matching ECP Snapshots, **Then** the system correctly identifies `detected_language` and applies language-aware tokenization/matching. 2. **Given** a corrupted JSON file or empty Markdown file, **When** `classify.py` executes, **Then** it outputs a structured JSON error response with appropriate error codes (`invalid_ecp_json`, `empty_content`, etc.) and non-zero exit code. --- ## Requirements *(mandatory)* ### Functional Requirements - **FR-001**: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`). - **FR-002**: System MUST provide a standalone Python CLI entrypoint: ```bash python classify.py --ecp --content --output ``` If `--output` is omitted, the result JSON MUST be emitted to standard output (`stdout`). Both `--ecp` and `--content` MUST be valid file paths. - **FR-003**: System MUST parse and validate the materialized ECP Snapshot JSON structure: - **Required fields**: `target_entity_id` (string), `target_name` (string), `aliases` (array of strings), `domain` (string), `anchors` (array of strings). - **Optional fields with defaults**: `negative_anchors` (array of strings, default `[]`), `graph_version` (string, default `"1.0.0"`), `related_entities` (array of objects, default `[]`). - **Related entity structure**: `entity_id` (string), `name` (string), `relation_type` (string), `weight` (number), `aliases` (array of strings, default `[]`), `scope` (string, default `"general"`), `confidence` (number, default `1.0`). - **FR-004**: System MUST execute a 3-tier hybrid classification pipeline: - **Tier 1 (Deterministic Rules - Mandatory in POC)**: Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches. - **Tier 2 (Local Multilingual Embeddings - Optional/Configurable)**: Semantic similarity scoring using local embeddings (can be toggled via config/flag). - **Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default)**: Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC. - **FR-005**: System MUST enforce the following decision logic: - `DIRECT_INHERENT`: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors. - `CONTEXTUAL_INHERENT`: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence. - `TANGENTIAL`: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence. - `NOT_RELATED`: Absence of relevant signals or dominant negative anchors. - **FR-006**: System MUST derive the boolean `is_inherent` strictly from `decision`: - `DIRECT_INHERENT` → `is_inherent: true` - `CONTEXTUAL_INHERENT` → `is_inherent: true` - `TANGENTIAL` → `is_inherent: false` - `NOT_RELATED` → `is_inherent: false` - **FR-007**: System MUST output a success JSON containing: - `decision`: String (`DIRECT_INHERENT` | `CONTEXTUAL_INHERENT` | `TANGENTIAL` | `NOT_RELATED`) - `is_inherent`: Boolean (`true` | `false`) - `confidence`: Float (`0.0` to `1.0`) - `detected_language`: String (ISO language code) - `matched_anchors`: Array of strings - `negative_matches`: Array of strings - `graph_matches`: Array of objects/strings (matched related entities from snapshot) - `evidence`: Array of strings (excerpts extracted from the Markdown) - `rationale`: String (short explanation of the decision) - `warnings`: Array of strings - **FR-008**: System MUST output structured error responses for failure conditions: - `error_code`: Enum (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`) - `message`: Human-readable error description - `details`: Object with debugging details --- ### Key Entities *(data models & domain entities)* - **Entity Context Profile (ECP) Snapshot (JSON)**: Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities. - **Content Item (Markdown)**: The input markdown document payload to be analyzed. - **Inherence Assessment Result (JSON)**: The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale. - **Classification Error (JSON)**: Structured failure payload with standardized error code. --- ## Success Criteria *(mandatory)* ### Measurable Outcomes - **SC-001**: System executes successfully via Python CLI without external network/service dependencies when running Tier 1. - **SC-002**: Classification precision reaches at least **90% over a controlled POC benchmark suite of 24 cases** (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`). - **SC-003**: 100% of outputs conform strictly to the specified JSON success/error schemas. - **SC-004**: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words. --- ## Assumptions & Scope ### Assumptions - The 6 core target languages are Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`). - ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime). - Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning. ### Explicit Out of Scope (POC) - REST API / FastAPI microservices. - PyPI package creation/distribution. - Message queues (RabbitMQ, SQS, Celery) or background workers. - Direct runtime database or graph database connections (Neo4j). - Integration with ContentMachineApp. - Mandatory external LLM API keys.