feat(classifier): add multilingual ECP inherence classifier POC

This commit is contained in:
2026-08-20 00:51:02 -03:00
parent d371b81aa4
commit 67cc40f91a
175 changed files with 30399 additions and 703 deletions
@@ -0,0 +1,113 @@
# Tasks: Multilingual NLP Entity Inherence Classifier (POC)
**Feature**: `001-multilingual-entity-classifier` | **Spec**: [spec.md](./spec.md) | **Plan**: [plan.md](./plan.md)
---
## Phase 1: Setup (Shared Infrastructure)
**Purpose**: Project initialization, directory structure, and environment configuration
- [x] T001 Create project directories (`src/`, `src/adapters/`, `examples/`, `tests/`, `tests/fixtures/benchmark_24/`)
- [x] T002 [P] Create `requirements.txt` with minimal test dependencies (`pytest>=7.0`) and optional dependencies commented
- [x] T003 [P] Create `.gitignore` validation and project configuration in `pyproject.toml` or `setup.cfg`
---
## Phase 2: Foundational (Blocking Prerequisites)
**Purpose**: Core data models, schema validation, and text processing utilities required by all user stories
**CRITICAL**: No classification engine logic can run until this phase is complete
- [x] T004 Implement data models and schema validation in `src/models.py` (`ECPSnapshot`, `ClassificationResult`, `ClassificationError`, and decision enums)
- [x] T005 [P] Implement Markdown parsing and evidence snippet extractor in `src/parser.py`
- [x] T006 [P] Implement lightweight multilingual language detector & diacritic normalizer for the 6 languages in `src/language.py`
- [x] T007 Implement unit tests for models, parser, and language detector in `tests/test_models.py` and `tests/test_language.py`
**Checkpoint**: Foundation ready - models, language detection, and Markdown parsing validated.
---
## Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP]
**Goal**: Deliver a functioning CLI (`classify.py`) executing Tier 1 deterministic evaluation (aliases, anchors, negative anchors, graph matches) on ECP Snapshots and Markdown documents.
**Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/test.json` and verify `decision: "DIRECT_INHERENT"`, `is_inherent: true`, confidence, and evidence.
### Tests for User Story 1
- [x] T008 [P] [US1] Create CLI contract and execution tests in `tests/test_cli.py` (testing `--ecp`, `--content`, `--output`, stdout fallback, and exit codes)
- [x] T009 [P] [US1] Create deterministic decision logic unit tests in `tests/test_classifier.py` for `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, and `NOT_RELATED`
### Implementation for User Story 1
- [x] T010 [US1] Implement Tier 1 deterministic matching engine in `src/classifier.py` (alias matching, negative anchor suppression, graph node resolution, scoring algorithm, and derived `is_inherent` boolean)
- [x] T011 [US1] Implement main CLI script `classify.py` handling argument parsing (`--ecp`, `--content`, `--output`, `--enable-embeddings`, `--enable-llm`), file I/O, error formatting, and stdout output
- [x] T012 [P] [US1] Create baseline example files: `examples/ecp_petrobras.json`, `examples/content_presal_pt.md`, `examples/ecp_volkswagen.json`, `examples/content_northvolt_de.md`, `examples/ecp_apple.json`, `examples/content_tangential_es.md`
- [x] T013 [US1] Verify end-to-end execution of CLI on example files and validate output JSON conforms strictly to schema
**Checkpoint**: User Story 1 complete! Standalone CLI works end-to-end locally with Tier 1 deterministic classification.
---
## Phase 4: User Story 2 - Multilingual 24-Case Controlled Benchmark Suite (Priority: P2)
**Goal**: Provide 24 paired test fixtures across all 6 languages (PT, EN, ES, DE, IT, FR) and 4 decision types, measuring ≥ 90% precision.
**Independent Test**: Run `pytest tests/test_benchmark_24.py -v` and achieve 100% pass rate across the 24 controlled benchmark fixtures.
### Implementation for User Story 2
- [x] T014 [P] [US2] Create Portuguese benchmark fixtures (4 cases: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`) in `tests/fixtures/benchmark_24/pt/`
- [x] T015 [P] [US2] Create English benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/en/`
- [x] T016 [P] [US2] Create Spanish benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/es/`
- [x] T017 [P] [US2] Create German benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/de/`
- [x] T018 [P] [US2] Create Italian benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/it/`
- [x] T019 [P] [US2] Create French benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/fr/`
- [x] T020 [US2] Implement parameterized benchmark runner in `tests/test_benchmark_24.py` verifying precision, score thresholds, evidence extraction, and language detection across all 24 cases
**Checkpoint**: User Story 2 complete! Multilingual robustness validated across 24 test cases with ≥ 90% precision.
---
## Phase 5: Optional Adapters - Tier 2 (Embeddings) & Tier 3 (LLM) (Priority: P3)
**Goal**: Provide clean, decoupled adapter interfaces for local embeddings and LLM fallback without introducing mandatory runtime dependencies.
**Independent Test**: Validate that adapter interfaces load correctly and remain inert/disabled by default unless explicit CLI flags are provided.
### Implementation for User Story 3
- [x] T021 [P] [US3] Create base adapter interface in `src/adapters/base.py` (`BaseNLPAdapter`)
- [x] T022 [P] [US3] Implement optional local vector similarity adapter stub in `src/adapters/embeddings.py` (activated via `--enable-embeddings`)
- [x] T023 [P] [US3] Implement optional LLM fallback adapter stub in `src/adapters/llm.py` (activated via `--enable-llm`, disabled by default)
- [x] T024 [US3] Wire adapter hooks into `src/classifier.py` ensuring Tier 1 handles clear cases without invoking Tier 2/Tier 3
**Checkpoint**: User Story 3 complete! Extensibility contracts established with zero breaking changes or mandatory cloud dependencies.
---
## Phase 6: Polish & Validation
**Purpose**: Quickstart verification, code hygiene, and documentation audit
- [x] T025 Execute all quickstart validation scenarios defined in `specs/001-multilingual-entity-classifier/quickstart.md`
- [x] T026 [P] Run full test suite (`pytest -v --tb=short`)
- [x] T027 Run `graphify update .` to index new implementation files into knowledge graph
---
## Dependencies & Execution Order
```mermaid
flowchart TD
Setup[Phase 1: Setup T001-T003] --> Foundational[Phase 2: Foundational T004-T007]
Foundational --> US1[Phase 3: US1 Core CLI & Tier 1 T008-T013]
US1 --> US2[Phase 4: US2 24-Case Benchmark T014-T020]
US1 --> US3[Phase 5: US3 Optional Adapters T021-T024]
US2 --> Polish[Phase 6: Polish T025-T027]
US3 --> Polish
```
### Implementation Strategy
1. **MVP (Phase 1 + 2 + 3)**: Delivers working `classify.py` with Tier 1 deterministic engine and example files.
2. **Benchmark Verification (Phase 4)**: Guarantees multilingual compliance on 24 controlled test fixtures.
3. **Adapter Stubs (Phase 5)**: Prepares the codebase for future embedding/LLM extensions cleanly.