Files
TextNLPClassifierApp/specs/001-multilingual-entity-classifier/spec.md
T

10 KiB
Raw Blame History

Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)

Feature Branch: 001-multilingual-entity-classifier

Created: 2026-08-19

Status: Draft / Documented (POC Scope)

Input: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."


Clarifications

Session 2026-08-19

  • Q: O que é o ECP e qual o formato de entrada esperado? → A: ECP significa Entity Context Profile. A entrada para o classificador é um ECP Snapshot materializado em formato JSON (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato Markdown.
  • Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um script Python CLI (python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp.
  • Q: Como o pipeline híbrido deve se comportar na POC? → A: O Tier 1 (regras determinísticas) é o núcleo obrigatório da POC. O Tier 2 (embeddings multilíngues locais) é opcional/configurável. O Tier 3 (LLM fallback) é opcional/configurável e desativado por padrão (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
  • Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED) com is_inherent como campo derivado booleano (true para DIRECT_INHERENT e CONTEXTUAL_INHERENT; false para TANGENTIAL e NOT_RELATED), acompanhado de confidence, matched_anchors, negative_matches, graph_matches, evidence, rationale e warnings.
  • Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo 24 casos (6 idiomas × 4 tipos de decisão), medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.

User Scenarios & Testing (mandatory)

User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)

As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.

Why this priority: Core value proposition and execution model of the POC.

Independent Test: Run python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json and verify the output contains decision: "DIRECT_INHERENT" and is_inherent: true.

Acceptance Scenarios:

  1. Given a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", When classify.py executes, Then it produces decision: "DIRECT_INHERENT", is_inherent: true, and matched aliases in matched_anchors.
  2. Given a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, When evaluated, Then it produces decision: "CONTEXTUAL_INHERENT", is_inherent: true, and "Northvolt" listed in graph_matches.
  3. Given a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, When evaluated against an ECP Snapshot for "Apple Inc.", Then it produces decision: "TANGENTIAL", is_inherent: false.
  4. Given an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", When evaluated, Then it produces decision: "NOT_RELATED", is_inherent: false.

User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)

As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.

Why this priority: Guarantees baseline reliability and debuggability across all target languages.

Independent Test: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.

Acceptance Scenarios:

  1. Given valid Markdown documents in each of the 6 languages, When classified against matching ECP Snapshots, Then the system correctly identifies detected_language and applies language-aware tokenization/matching.
  2. Given a corrupted JSON file or empty Markdown file, When classify.py executes, Then it outputs a structured JSON error response with appropriate error codes (invalid_ecp_json, empty_content, etc.) and non-zero exit code.

Requirements (mandatory)

Functional Requirements

  • FR-001: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (pt), English (en), Spanish (es), German (de), Italian (it), and French (fr).
  • FR-002: System MUST provide a standalone Python CLI entrypoint:
    python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>
    
    If --output is omitted, the result JSON MUST be emitted to standard output (stdout). Both --ecp and --content MUST be valid file paths.
  • FR-003: System MUST parse and validate the materialized ECP Snapshot JSON structure:
    • Required fields: target_entity_id (string), target_name (string), aliases (array of strings), domain (string), anchors (array of strings).
    • Optional fields with defaults: negative_anchors (array of strings, default []), graph_version (string, default "1.0.0"), related_entities (array of objects, default []).
    • Related entity structure: entity_id (string), name (string), relation_type (string), weight (number), aliases (array of strings, default []), scope (string, default "general"), confidence (number, default 1.0).
  • FR-004: System MUST execute a 3-tier hybrid classification pipeline:
    • Tier 1 (Deterministic Rules - Mandatory in POC): Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
    • Tier 2 (Local Multilingual Embeddings - Optional/Configurable): Semantic similarity scoring using local embeddings (can be toggled via config/flag).
    • Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default): Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
  • FR-005: System MUST enforce the following decision logic:
    • DIRECT_INHERENT: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.
    • CONTEXTUAL_INHERENT: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.
    • TANGENTIAL: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.
    • NOT_RELATED: Absence of relevant signals or dominant negative anchors.
  • FR-006: System MUST derive the boolean is_inherent strictly from decision:
    • DIRECT_INHERENT → is_inherent: true
    • CONTEXTUAL_INHERENT → is_inherent: true
    • TANGENTIAL → is_inherent: false
    • NOT_RELATED → is_inherent: false
  • FR-007: System MUST output a success JSON containing:
    • decision: String (DIRECT_INHERENT | CONTEXTUAL_INHERENT | TANGENTIAL | NOT_RELATED)
    • is_inherent: Boolean (true | false)
    • confidence: Float (0.0 to 1.0)
    • detected_language: String (ISO language code)
    • matched_anchors: Array of strings
    • negative_matches: Array of strings
    • graph_matches: Array of objects/strings (matched related entities from snapshot)
    • evidence: Array of strings (excerpts extracted from the Markdown)
    • rationale: String (short explanation of the decision)
    • warnings: Array of strings
  • FR-008: System MUST output structured error responses for failure conditions:
    • error_code: Enum (invalid_ecp_json, invalid_markdown, unsupported_language, empty_content, missing_required_field)
    • message: Human-readable error description
    • details: Object with debugging details

Key Entities (data models & domain entities)

  • Entity Context Profile (ECP) Snapshot (JSON): Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
  • Content Item (Markdown): The input markdown document payload to be analyzed.
  • Inherence Assessment Result (JSON): The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
  • Classification Error (JSON): Structured failure payload with standardized error code.

Success Criteria (mandatory)

Measurable Outcomes

  • SC-001: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
  • SC-002: Classification precision reaches at least 90% over a controlled POC benchmark suite of 24 cases (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED).
  • SC-003: 100% of outputs conform strictly to the specified JSON success/error schemas.
  • SC-004: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.

Assumptions & Scope

Assumptions

  • The 6 core target languages are Portuguese (pt), English (en), Spanish (es), German (de), Italian (it), and French (fr).
  • ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
  • Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.

Explicit Out of Scope (POC)

  • REST API / FastAPI microservices.
  • PyPI package creation/distribution.
  • Message queues (RabbitMQ, SQS, Celery) or background workers.
  • Direct runtime database or graph database connections (Neo4j).
  • Integration with ContentMachineApp.
  • Mandatory external LLM API keys.