139 lines
10 KiB
Markdown
139 lines
10 KiB
Markdown
# Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
||
|
||
**Feature Branch**: `001-multilingual-entity-classifier`
|
||
|
||
**Created**: 2026-08-19
|
||
|
||
**Status**: Draft / Documented (POC Scope)
|
||
|
||
**Input**: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."
|
||
|
||
---
|
||
|
||
## Clarifications
|
||
|
||
### Session 2026-08-19
|
||
|
||
- Q: O que é o ECP e qual o formato de entrada esperado? → A: **ECP significa Entity Context Profile**. A entrada para o classificador é um **ECP Snapshot materializado em formato JSON** (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato **Markdown**.
|
||
- Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um **script Python CLI** (`python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>`). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp.
|
||
- Q: Como o pipeline híbrido deve se comportar na POC? → A: O **Tier 1 (regras determinísticas)** é o núcleo obrigatório da POC. O **Tier 2 (embeddings multilíngues locais)** é opcional/configurável. O **Tier 3 (LLM fallback)** é opcional/configurável e **desativado por padrão** (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
|
||
- Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) com `is_inherent` como campo derivado booleano (`true` para `DIRECT_INHERENT` e `CONTEXTUAL_INHERENT`; `false` para `TANGENTIAL` e `NOT_RELATED`), acompanhado de `confidence`, `matched_anchors`, `negative_matches`, `graph_matches`, `evidence`, `rationale` e `warnings`.
|
||
- Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo **24 casos (6 idiomas × 4 tipos de decisão)**, medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.
|
||
|
||
---
|
||
|
||
## User Scenarios & Testing *(mandatory)*
|
||
|
||
### User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)
|
||
|
||
As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.
|
||
|
||
**Why this priority**: Core value proposition and execution model of the POC.
|
||
|
||
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json` and verify the output contains `decision: "DIRECT_INHERENT"` and `is_inherent: true`.
|
||
|
||
**Acceptance Scenarios**:
|
||
|
||
1. **Given** a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", **When** `classify.py` executes, **Then** it produces `decision: "DIRECT_INHERENT"`, `is_inherent: true`, and matched aliases in `matched_anchors`.
|
||
2. **Given** a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, **When** evaluated, **Then** it produces `decision: "CONTEXTUAL_INHERENT"`, `is_inherent: true`, and "Northvolt" listed in `graph_matches`.
|
||
3. **Given** a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, **When** evaluated against an ECP Snapshot for "Apple Inc.", **Then** it produces `decision: "TANGENTIAL"`, `is_inherent: false`.
|
||
4. **Given** an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", **When** evaluated, **Then** it produces `decision: "NOT_RELATED"`, `is_inherent: false`.
|
||
|
||
---
|
||
|
||
### User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)
|
||
|
||
As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.
|
||
|
||
**Why this priority**: Guarantees baseline reliability and debuggability across all target languages.
|
||
|
||
**Independent Test**: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.
|
||
|
||
**Acceptance Scenarios**:
|
||
|
||
1. **Given** valid Markdown documents in each of the 6 languages, **When** classified against matching ECP Snapshots, **Then** the system correctly identifies `detected_language` and applies language-aware tokenization/matching.
|
||
2. **Given** a corrupted JSON file or empty Markdown file, **When** `classify.py` executes, **Then** it outputs a structured JSON error response with appropriate error codes (`invalid_ecp_json`, `empty_content`, etc.) and non-zero exit code.
|
||
|
||
---
|
||
|
||
## Requirements *(mandatory)*
|
||
|
||
### Functional Requirements
|
||
|
||
- **FR-001**: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
|
||
- **FR-002**: System MUST provide a standalone Python CLI entrypoint:
|
||
```bash
|
||
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>
|
||
```
|
||
If `--output` is omitted, the result JSON MUST be emitted to standard output (`stdout`). Both `--ecp` and `--content` MUST be valid file paths.
|
||
- **FR-003**: System MUST parse and validate the materialized ECP Snapshot JSON structure:
|
||
- **Required fields**: `target_entity_id` (string), `target_name` (string), `aliases` (array of strings), `domain` (string), `anchors` (array of strings).
|
||
- **Optional fields with defaults**: `negative_anchors` (array of strings, default `[]`), `graph_version` (string, default `"1.0.0"`), `related_entities` (array of objects, default `[]`).
|
||
- **Related entity structure**: `entity_id` (string), `name` (string), `relation_type` (string), `weight` (number), `aliases` (array of strings, default `[]`), `scope` (string, default `"general"`), `confidence` (number, default `1.0`).
|
||
- **FR-004**: System MUST execute a 3-tier hybrid classification pipeline:
|
||
- **Tier 1 (Deterministic Rules - Mandatory in POC)**: Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
|
||
- **Tier 2 (Local Multilingual Embeddings - Optional/Configurable)**: Semantic similarity scoring using local embeddings (can be toggled via config/flag).
|
||
- **Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default)**: Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
|
||
- **FR-005**: System MUST enforce the following decision logic:
|
||
- `DIRECT_INHERENT`: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.
|
||
- `CONTEXTUAL_INHERENT`: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.
|
||
- `TANGENTIAL`: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.
|
||
- `NOT_RELATED`: Absence of relevant signals or dominant negative anchors.
|
||
- **FR-006**: System MUST derive the boolean `is_inherent` strictly from `decision`:
|
||
- `DIRECT_INHERENT` → `is_inherent: true`
|
||
- `CONTEXTUAL_INHERENT` → `is_inherent: true`
|
||
- `TANGENTIAL` → `is_inherent: false`
|
||
- `NOT_RELATED` → `is_inherent: false`
|
||
- **FR-007**: System MUST output a success JSON containing:
|
||
- `decision`: String (`DIRECT_INHERENT` | `CONTEXTUAL_INHERENT` | `TANGENTIAL` | `NOT_RELATED`)
|
||
- `is_inherent`: Boolean (`true` | `false`)
|
||
- `confidence`: Float (`0.0` to `1.0`)
|
||
- `detected_language`: String (ISO language code)
|
||
- `matched_anchors`: Array of strings
|
||
- `negative_matches`: Array of strings
|
||
- `graph_matches`: Array of objects/strings (matched related entities from snapshot)
|
||
- `evidence`: Array of strings (excerpts extracted from the Markdown)
|
||
- `rationale`: String (short explanation of the decision)
|
||
- `warnings`: Array of strings
|
||
- **FR-008**: System MUST output structured error responses for failure conditions:
|
||
- `error_code`: Enum (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`)
|
||
- `message`: Human-readable error description
|
||
- `details`: Object with debugging details
|
||
|
||
---
|
||
|
||
### Key Entities *(data models & domain entities)*
|
||
|
||
- **Entity Context Profile (ECP) Snapshot (JSON)**: Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
|
||
- **Content Item (Markdown)**: The input markdown document payload to be analyzed.
|
||
- **Inherence Assessment Result (JSON)**: The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
|
||
- **Classification Error (JSON)**: Structured failure payload with standardized error code.
|
||
|
||
---
|
||
|
||
## Success Criteria *(mandatory)*
|
||
|
||
### Measurable Outcomes
|
||
|
||
- **SC-001**: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
|
||
- **SC-002**: Classification precision reaches at least **90% over a controlled POC benchmark suite of 24 cases** (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`).
|
||
- **SC-003**: 100% of outputs conform strictly to the specified JSON success/error schemas.
|
||
- **SC-004**: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.
|
||
|
||
---
|
||
|
||
## Assumptions & Scope
|
||
|
||
### Assumptions
|
||
- The 6 core target languages are Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
|
||
- ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
|
||
- Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.
|
||
|
||
### Explicit Out of Scope (POC)
|
||
- REST API / FastAPI microservices.
|
||
- PyPI package creation/distribution.
|
||
- Message queues (RabbitMQ, SQS, Celery) or background workers.
|
||
- Direct runtime database or graph database connections (Neo4j).
|
||
- Integration with ContentMachineApp.
|
||
- Mandatory external LLM API keys.
|