Files
TextNLPClassifierApp/specs/001-multilingual-entity-classifier/spec.md
T

139 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
**Feature Branch**: `001-multilingual-entity-classifier`
**Created**: 2026-08-19
**Status**: Draft / Documented (POC Scope)
**Input**: User description: "Ferramenta de NLP em vários idiomas (português, inglês, espanhol, alemão, italiano, francês). Recebe um conteúdo em Markdown e um Entity Context Profile (ECP Snapshot em JSON), analisando se aquele conteúdo é inerente àquela entidade."
---
## Clarifications
### Session 2026-08-19
- Q: O que é o ECP e qual o formato de entrada esperado? → A: **ECP significa Entity Context Profile**. A entrada para o classificador é um **ECP Snapshot materializado em formato JSON** (autocontido, com entidade alvo, aliases, anchors e entidades relacionadas do grafo). O conteúdo a ser avaliado é fornecido em formato **Markdown**.
- Q: Qual a interface e escopo da POC? → A: A ferramenta de classificação roda exclusivamente como um **script Python CLI** (`python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>`). Ficam fora de escopo para a POC: REST API, pacote publicável, filas/workers e integrações em runtime com banco de grafos (Neo4j) ou ContentMachineApp.
- Q: Como o pipeline híbrido deve se comportar na POC? → A: O **Tier 1 (regras determinísticas)** é o núcleo obrigatório da POC. O **Tier 2 (embeddings multilíngues locais)** é opcional/configurável. O **Tier 3 (LLM fallback)** é opcional/configurável e **desativado por padrão** (casos óbvios não chamam LLM e nenhuma API key externa é necessária para executar a POC).
- Q: Quais as regras mínimas de decisão e o formato de saída? → A: Decisão categorizada em 4 níveis (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`) com `is_inherent` como campo derivado booleano (`true` para `DIRECT_INHERENT` e `CONTEXTUAL_INHERENT`; `false` para `TANGENTIAL` e `NOT_RELATED`), acompanhado de `confidence`, `matched_anchors`, `negative_matches`, `graph_matches`, `evidence`, `rationale` e `warnings`.
- Q: Como deve ser o benchmark de validação da POC? → A: Um conjunto de teste controlado com no mínimo **24 casos (6 idiomas × 4 tipos de decisão)**, medindo acurácia ≥ 90% especificamente sobre essa suíte de benchmark da POC.
---
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Classify Content Inherence via CLI for an ECP Snapshot (Priority: P1)
As a developer or automation script, I want to execute a Python CLI command passing an ECP Snapshot (JSON) and a document (Markdown), so that I get a structured JSON evaluation determining whether the document is inherent to the target entity.
**Why this priority**: Core value proposition and execution model of the POC.
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/result.json` and verify the output contains `decision: "DIRECT_INHERENT"` and `is_inherent: true`.
**Acceptance Scenarios**:
1. **Given** a Portuguese Markdown article discussing offshore oil exploration and an ECP Snapshot for "Petrobras", **When** `classify.py` executes, **Then** it produces `decision: "DIRECT_INHERENT"`, `is_inherent: true`, and matched aliases in `matched_anchors`.
2. **Given** a German Markdown article about battery supply chains and an ECP Snapshot for "Volkswagen Group" with a related entity "Northvolt" in its graph snapshot, **When** evaluated, **Then** it produces `decision: "CONTEXTUAL_INHERENT"`, `is_inherent: true`, and "Northvolt" listed in `graph_matches`.
3. **Given** a Spanish Markdown text that mentions "Apple" only in a passing metaphor ("la manzana de la discordia") without tech context, **When** evaluated against an ECP Snapshot for "Apple Inc.", **Then** it produces `decision: "TANGENTIAL"`, `is_inherent: false`.
4. **Given** an Italian recipe text and an ECP Snapshot for "Ferrari N.V.", **When** evaluated, **Then** it produces `decision: "NOT_RELATED"`, `is_inherent: false`.
---
### User Story 2 - Multilingual Language Support & Edge Validation (Priority: P2)
As an evaluator, I want the CLI to handle content across 6 languages (PT, EN, ES, DE, IT, FR) and return explicit, structured error diagnostics on invalid inputs.
**Why this priority**: Guarantees baseline reliability and debuggability across all target languages.
**Independent Test**: Execute the CLI against a test suite covering the 6 languages and against corrupted/empty inputs, verifying proper decisions and error codes.
**Acceptance Scenarios**:
1. **Given** valid Markdown documents in each of the 6 languages, **When** classified against matching ECP Snapshots, **Then** the system correctly identifies `detected_language` and applies language-aware tokenization/matching.
2. **Given** a corrupted JSON file or empty Markdown file, **When** `classify.py` executes, **Then** it outputs a structured JSON error response with appropriate error codes (`invalid_ecp_json`, `empty_content`, etc.) and non-zero exit code.
---
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST support multilingual NLP classification across at least 6 core languages: Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- **FR-002**: System MUST provide a standalone Python CLI entrypoint:
```bash
python classify.py --ecp <snapshot.json> --content <doc.md> --output <result.json>
```
If `--output` is omitted, the result JSON MUST be emitted to standard output (`stdout`). Both `--ecp` and `--content` MUST be valid file paths.
- **FR-003**: System MUST parse and validate the materialized ECP Snapshot JSON structure:
- **Required fields**: `target_entity_id` (string), `target_name` (string), `aliases` (array of strings), `domain` (string), `anchors` (array of strings).
- **Optional fields with defaults**: `negative_anchors` (array of strings, default `[]`), `graph_version` (string, default `"1.0.0"`), `related_entities` (array of objects, default `[]`).
- **Related entity structure**: `entity_id` (string), `name` (string), `relation_type` (string), `weight` (number), `aliases` (array of strings, default `[]`), `scope` (string, default `"general"`), `confidence` (number, default `1.0`).
- **FR-004**: System MUST execute a 3-tier hybrid classification pipeline:
- **Tier 1 (Deterministic Rules - Mandatory in POC)**: Token matching for aliases, anchors, negative anchors, language detection, and graph snapshot node matches.
- **Tier 2 (Local Multilingual Embeddings - Optional/Configurable)**: Semantic similarity scoring using local embeddings (can be toggled via config/flag).
- **Tier 3 (LLM Fallback - Optional/Configurable & Disabled by Default)**: Invoked only for unresolved ambiguous boundary cases when explicitly enabled. Clear cases MUST NOT call LLM. No external API key is required to run the POC.
- **FR-005**: System MUST enforce the following decision logic:
- `DIRECT_INHERENT`: Strong match of target entity aliases + compatible domain/context + no dominant negative anchors.
- `CONTEXTUAL_INHERENT`: Match of related entity from snapshot + scope criteria satisfied + sufficient relational weight/confidence.
- `TANGENTIAL`: Weak/isolated mention, relationship lacking surrounding context, or insufficient evidence.
- `NOT_RELATED`: Absence of relevant signals or dominant negative anchors.
- **FR-006**: System MUST derive the boolean `is_inherent` strictly from `decision`:
- `DIRECT_INHERENT` → `is_inherent: true`
- `CONTEXTUAL_INHERENT` → `is_inherent: true`
- `TANGENTIAL` → `is_inherent: false`
- `NOT_RELATED` → `is_inherent: false`
- **FR-007**: System MUST output a success JSON containing:
- `decision`: String (`DIRECT_INHERENT` | `CONTEXTUAL_INHERENT` | `TANGENTIAL` | `NOT_RELATED`)
- `is_inherent`: Boolean (`true` | `false`)
- `confidence`: Float (`0.0` to `1.0`)
- `detected_language`: String (ISO language code)
- `matched_anchors`: Array of strings
- `negative_matches`: Array of strings
- `graph_matches`: Array of objects/strings (matched related entities from snapshot)
- `evidence`: Array of strings (excerpts extracted from the Markdown)
- `rationale`: String (short explanation of the decision)
- `warnings`: Array of strings
- **FR-008**: System MUST output structured error responses for failure conditions:
- `error_code`: Enum (`invalid_ecp_json`, `invalid_markdown`, `unsupported_language`, `empty_content`, `missing_required_field`)
- `message`: Human-readable error description
- `details`: Object with debugging details
---
### Key Entities *(data models & domain entities)*
- **Entity Context Profile (ECP) Snapshot (JSON)**: Materialized, self-contained entity snapshot containing target entity metadata, aliases, anchors, negative anchors, and relevant connected graph entities.
- **Content Item (Markdown)**: The input markdown document payload to be analyzed.
- **Inherence Assessment Result (JSON)**: The structured result produced by the CLI containing classification verdict, derived boolean, confidence score, evidence snippets, and rationale.
- **Classification Error (JSON)**: Structured failure payload with standardized error code.
---
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: System executes successfully via Python CLI without external network/service dependencies when running Tier 1.
- **SC-002**: Classification precision reaches at least **90% over a controlled POC benchmark suite of 24 cases** (6 languages: PT, EN, ES, DE, IT, FR × 4 decision types: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`).
- **SC-003**: 100% of outputs conform strictly to the specified JSON success/error schemas.
- **SC-004**: CLI execution time for Tier 1 deterministic evaluation is under 200ms for documents under 2,000 words.
---
## Assumptions & Scope
### Assumptions
- The 6 core target languages are Portuguese (`pt`), English (`en`), Spanish (`es`), German (`de`), Italian (`it`), and French (`fr`).
- ECP Snapshots are pre-materialized JSON files (the ECP Manager / Neo4j is responsible for publishing snapshots; the classifier does not connect to Neo4j or external databases in runtime).
- Tier 1 deterministic logic is sufficient for baseline evaluation; LLM fallback is an optional enhancement for future edge-case tuning.
### Explicit Out of Scope (POC)
- REST API / FastAPI microservices.
- PyPI package creation/distribution.
- Message queues (RabbitMQ, SQS, Celery) or background workers.
- Direct runtime database or graph database connections (Neo4j).
- Integration with ContentMachineApp.
- Mandatory external LLM API keys.