10 KiB
Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
Feature Branch: 001-multilingual-entity-classifier
Created: 2026-08-19
Status: Draft / Documented (POC Scope)
Input: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."
Clarifications
Session 2026-08-19
- Q: O que é o ECP e qual o formato de entrada esperado? → A: ECP significa Entity Context Profile. A entrada para o classificador é um ECP Snapshot materializado em formato JSON (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato Markdown.
- Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um script Python CLI (
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp. - Q: Como o pipeline híbrido deve se comportar na POC? → A: O Tier 1 (regras determinísticas) é o núcleo obrigatório da POC. O Tier 2 (embeddings multilíngues locais) é opcional/configurável. O Tier 3 (LLM fallback) é opcional/configurável e desativado por padrão (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
- Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (
DIRECT_INHERENT,CONTEXTUAL_INHERENT,TANGENTIAL,NOT_RELATED) comis_inherentcomo campo derivado booleano (trueparaDIRECT_INHERENTeCONTEXTUAL_INHERENT;falseparaTANGENTIALeNOT_RELATED), acompanhado deconfidence,matched_anchors,negative_matches,graph_matches,evidence,rationaleewarnings. - Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo 24 casos (6 idiomas × 4 tipos de decisão), medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.
User Scenarios & Testing (mandatory)
User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)
As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.
Why this priority: Core value proposition and execution model of the POC.
Independent Test: Run python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json and verify the output contains decision: "DIRECT_INHERENT" and is_inherent: true.
Acceptance Scenarios:
- Given a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", When
classify.pyexecutes, Then it producesdecision: "DIRECT_INHERENT",is_inherent: true, and matched aliases inmatched_anchors. - Given a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, When evaluated, Then it produces
decision: "CONTEXTUAL_INHERENT",is_inherent: true, and "Northvolt" listed ingraph_matches. - Given a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, When evaluated against an ECP Snapshot for "Apple Inc.", Then it produces
decision: "TANGENTIAL",is_inherent: false. - Given an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", When evaluated, Then it produces
decision: "NOT_RELATED",is_inherent: false.
User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)
As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.
Why this priority: Guarantees baseline reliability and debuggability across all target languages.
Independent Test: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.
Acceptance Scenarios:
- Given valid Markdown documents in each of the 6 languages, When classified against matching ECP Snapshots, Then the system correctly identifies
detected_languageand applies language-aware tokenization/matching. - Given a corrupted JSON file or empty Markdown file, When
classify.pyexecutes, Then it outputs a structured JSON error response with appropriate error codes (invalid_ecp_json,empty_content, etc.) and non-zero exit code.
Requirements (mandatory)
Functional Requirements
- FR-001: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (
pt), English (en), Spanish (es), German (de), Italian (it), and French (fr). - FR-002: System MUST provide a standalone Python CLI entrypoint:
If
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>--outputis omitted, the result JSON MUST be emitted to standard output (stdout). Both--ecpand--contentMUST be valid file paths. - FR-003: System MUST parse and validate the materialized ECP Snapshot JSON structure:
- Required fields:
target_entity_id(string),target_name(string),aliases(array of strings),domain(string),anchors(array of strings). - Optional fields with defaults:
negative_anchors(array of strings, default[]),graph_version(string, default"1.0.0"),related_entities(array of objects, default[]). - Related entity structure:
entity_id(string),name(string),relation_type(string),weight(number),aliases(array of strings, default[]),scope(string, default"general"),confidence(number, default1.0).
- Required fields:
- FR-004: System MUST execute a 3-tier hybrid classification pipeline:
- Tier 1 (Deterministic Rules - Mandatory in POC): Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
- Tier 2 (Local Multilingual Embeddings - Optional/Configurable): Semantic similarity scoring using local embeddings (can be toggled via config/flag).
- Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default): Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
- FR-005: System MUST enforce the following decision logic:
DIRECT_INHERENT: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.CONTEXTUAL_INHERENT: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.TANGENTIAL: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.NOT_RELATED: Absence of relevant signals or dominant negative anchors.
- FR-006: System MUST derive the boolean
is_inherentstrictly fromdecision:DIRECT_INHERENT→is_inherent: trueCONTEXTUAL_INHERENT→is_inherent: trueTANGENTIAL→is_inherent: falseNOT_RELATED→is_inherent: false
- FR-007: System MUST output a success JSON containing:
decision: String (DIRECT_INHERENT|CONTEXTUAL_INHERENT|TANGENTIAL|NOT_RELATED)is_inherent: Boolean (true|false)confidence: Float (0.0to1.0)detected_language: String (ISO language code)matched_anchors: Array of stringsnegative_matches: Array of stringsgraph_matches: Array of objects/strings (matched related entities from snapshot)evidence: Array of strings (excerpts extracted from the Markdown)rationale: String (short explanation of the decision)warnings: Array of strings
- FR-008: System MUST output structured error responses for failure conditions:
error_code: Enum (invalid_ecp_json,invalid_markdown,unsupported_language,empty_content,missing_required_field)message: Human-readable error descriptiondetails: Object with debugging details
Key Entities (data models & domain entities)
- Entity Context Profile (ECP) Snapshot (JSON): Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
- Content Item (Markdown): The input markdown document payload to be analyzed.
- Inherence Assessment Result (JSON): The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
- Classification Error (JSON): Structured failure payload with standardized error code.
Success Criteria (mandatory)
Measurable Outcomes
- SC-001: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
- SC-002: Classification precision reaches at least 90% over a controlled POC benchmark suite of 24 cases (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types:
DIRECT_INHERENT,CONTEXTUAL_INHERENT,TANGENTIAL,NOT_RELATED). - SC-003: 100% of outputs conform strictly to the specified JSON success/error schemas.
- SC-004: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.
Assumptions & Scope
Assumptions
- The 6 core target languages are Portuguese (
pt), English (en), Spanish (es), German (de), Italian (it), and French (fr). - ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
- Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.
Explicit Out of Scope (POC)
- REST API / FastAPI microservices.
- PyPI package creation/distribution.
- Message queues (RabbitMQ, SQS, Celery) or background workers.
- Direct runtime database or graph database connections (Neo4j).
- Integration with ContentMachineApp.
- Mandatory external LLM API keys.