feat(classifier): add multilingual ECP inherence classifier POC

This commit is contained in:
2026-08-20 00:51:02 -03:00
parent d371b81aa4
commit 67cc40f91a
175 changed files with 30399 additions and 703 deletions
@@ -0,0 +1,138 @@
# Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
**Feature Branch**: `001-multilingual-entity-classifier`
**Created**: 2026-08-19
**Status**: Draft / Documented (POC Scope)
**Input**: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."
---
## Clarifications
### Session 2026-08-19
- Q: O que é o ECP e qual o formato de entrada esperado? → A: **ECP significa Entity Context Profile**. A entrada para o classificador é um **ECP Snapshot materializado em formato JSON** (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato **Markdown**.
- Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um **script Python CLI** (`python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>`). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp.
- Q: Como o pipeline híbrido deve se comportar na POC? → A: O **Tier 1 (regras determinísticas)** é o núcleo obrigatório da POC. O **Tier 2 (embeddings multilíngues locais)** é opcional/configurável. O **Tier 3 (LLM fallback)** é opcional/configurável e **desativado por padrão** (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
- Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) com `is_inherent` como campo derivado booleano (`true` para `DIRECT_INHERENT` e `CONTEXTUAL_INHERENT`; `false` para `TANGENTIAL` e `NOT_RELATED`), acompanhado de `confidence`, `matched_anchors`, `negative_matches`, `graph_matches`, `evidence`, `rationale` e `warnings`.
- Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo **24 casos (6 idiomas × 4 tipos de decisão)**, medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.
---
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)
As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.
**Why this priority**: Core value proposition and execution model of the POC.
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json` and verify the output contains `decision: "DIRECT_INHERENT"` and `is_inherent: true`.
**Acceptance Scenarios**:
1. **Given** a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", **When** `classify.py` executes, **Then** it produces `decision: "DIRECT_INHERENT"`, `is_inherent: true`, and matched aliases in `matched_anchors`.
2. **Given** a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, **When** evaluated, **Then** it produces `decision: "CONTEXTUAL_INHERENT"`, `is_inherent: true`, and "Northvolt" listed in `graph_matches`.
3. **Given** a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, **When** evaluated against an ECP Snapshot for "Apple Inc.", **Then** it produces `decision: "TANGENTIAL"`, `is_inherent: false`.
4. **Given** an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", **When** evaluated, **Then** it produces `decision: "NOT_RELATED"`, `is_inherent: false`.
---
### User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)
As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.
**Why this priority**: Guarantees baseline reliability and debuggability across all target languages.
**Independent Test**: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.
**Acceptance Scenarios**:
1. **Given** valid Markdown documents in each of the 6 languages, **When** classified against matching ECP Snapshots, **Then** the system correctly identifies `detected_language` and applies language-aware tokenization/matching.
2. **Given** a corrupted JSON file or empty Markdown file, **When** `classify.py` executes, **Then** it outputs a structured JSON error response with appropriate error codes (`invalid_ecp_json`, `empty_content`, etc.) and non-zero exit code.
---
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- **FR-002**: System MUST provide a standalone Python CLI entrypoint:
```bash
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>
```
If `--output` is omitted, the result JSON MUST be emitted to standard output (`stdout`). Both `--ecp` and `--content` MUST be valid file paths.
- **FR-003**: System MUST parse and validate the materialized ECP Snapshot JSON structure:
- **Required fields**: `target_entity_id` (string), `target_name` (string), `aliases` (array of strings), `domain` (string), `anchors` (array of strings).
- **Optional fields with defaults**: `negative_anchors` (array of strings, default `[]`), `graph_version` (string, default `"1.0.0"`), `related_entities` (array of objects, default `[]`).
- **Related entity structure**: `entity_id` (string), `name` (string), `relation_type` (string), `weight` (number), `aliases` (array of strings, default `[]`), `scope` (string, default `"general"`), `confidence` (number, default `1.0`).
- **FR-004**: System MUST execute a 3-tier hybrid classification pipeline:
- **Tier 1 (Deterministic Rules - Mandatory in POC)**: Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
- **Tier 2 (Local Multilingual Embeddings - Optional/Configurable)**: Semantic similarity scoring using local embeddings (can be toggled via config/flag).
- **Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default)**: Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
- **FR-005**: System MUST enforce the following decision logic:
- `DIRECT_INHERENT`: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.
- `CONTEXTUAL_INHERENT`: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.
- `TANGENTIAL`: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.
- `NOT_RELATED`: Absence of relevant signals or dominant negative anchors.
- **FR-006**: System MUST derive the boolean `is_inherent` strictly from `decision`:
- `DIRECT_INHERENT` → `is_inherent: true`
- `CONTEXTUAL_INHERENT` → `is_inherent: true`
- `TANGENTIAL` → `is_inherent: false`
- `NOT_RELATED` → `is_inherent: false`
- **FR-007**: System MUST output a success JSON containing:
- `decision`: String (`DIRECT_INHERENT` | `CONTEXTUAL_INHERENT` | `TANGENTIAL` | `NOT_RELATED`)
- `is_inherent`: Boolean (`true` | `false`)
- `confidence`: Float (`0.0` to `1.0`)
- `detected_language`: String (ISO language code)
- `matched_anchors`: Array of strings
- `negative_matches`: Array of strings
- `graph_matches`: Array of objects/strings (matched related entities from snapshot)
- `evidence`: Array of strings (excerpts extracted from the Markdown)
- `rationale`: String (short explanation of the decision)
- `warnings`: Array of strings
- **FR-008**: System MUST output structured error responses for failure conditions:
- `error_code`: Enum (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`)
- `message`: Human-readable error description
- `details`: Object with debugging details
---
### Key Entities *(data models & domain entities)*
- **Entity Context Profile (ECP) Snapshot (JSON)**: Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
- **Content Item (Markdown)**: The input markdown document payload to be analyzed.
- **Inherence Assessment Result (JSON)**: The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
- **Classification Error (JSON)**: Structured failure payload with standardized error code.
---
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
- **SC-002**: Classification precision reaches at least **90% over a controlled POC benchmark suite of 24 cases** (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`).
- **SC-003**: 100% of outputs conform strictly to the specified JSON success/error schemas.
- **SC-004**: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.
---
## Assumptions & Scope
### Assumptions
- The 6 core target languages are Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
- Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.
### Explicit Out of Scope (POC)
- REST API / FastAPI microservices.
- PyPI package creation/distribution.
- Message queues (RabbitMQ, SQS, Celery) or background workers.
- Direct runtime database or graph database connections (Neo4j).
- Integration with ContentMachineApp.
- Mandatory external LLM API keys.