# Tasks: Multilingual NLP Entity Inherence Classifier (POC) **Feature**: `001-multilingual-entity-classifier` | **Spec**: [spec.md](./spec.md) | **Plan**: [plan.md](./plan.md) --- ## Phase 1: Setup (Shared Infrastructure) **Purpose**: Project initialization, directory structure, and environment configuration - [x] T001 Create project directories (`src/`, `src/adapters/`, `examples/`, `tests/`, `tests/fixtures/benchmark_24/`) - [x] T002 [P] Create `requirements.txt` with minimal test dependencies (`pytest>=7.0`) and optional dependencies commented - [x] T003 [P] Create `.gitignore` validation and project configuration in `pyproject.toml` or `setup.cfg` --- ## Phase 2: Foundational (Blocking Prerequisites) **Purpose**: Core data models, schema validation, and text processing utilities required by all user stories **CRITICAL**: No classification engine logic can run until this phase is complete - [x] T004 Implement data models and schema validation in `src/models.py` (`ECPSnapshot`, `ClassificationResult`, `ClassificationError`, and decision enums) - [x] T005 [P] Implement Markdown parsing and evidence snippet extractor in `src/parser.py` - [x] T006 [P] Implement lightweight multilingual language detector & diacritic normalizer for the 6 languages in `src/language.py` - [x] T007 Implement unit tests for models, parser, and language detector in `tests/test_models.py` and `tests/test_language.py` **Checkpoint**: Foundation ready - models, language detection, and Markdown parsing validated. --- ## Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] **Goal**: Deliver a functioning CLI (`classify.py`) executing Tier 1 deterministic evaluation (aliases, anchors, negative anchors, graph matches) on ECP Snapshots and Markdown documents. **Independent Test**: Run `python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/test.json` and verify `decision: "DIRECT_INHERENT"`, `is_inherent: true`, confidence, and evidence. ### Tests for User Story 1 - [x] T008 [P] [US1] Create CLI contract and execution tests in `tests/test_cli.py` (testing `--ecp`, `--content`, `--output`, stdout fallback, and exit codes) - [x] T009 [P] [US1] Create deterministic decision logic unit tests in `tests/test_classifier.py` for `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, and `NOT_RELATED` ### Implementation for User Story 1 - [x] T010 [US1] Implement Tier 1 deterministic matching engine in `src/classifier.py` (alias matching, negative anchor suppression, graph node resolution, scoring algorithm, and derived `is_inherent` boolean) - [x] T011 [US1] Implement main CLI script `classify.py` handling argument parsing (`--ecp`, `--content`, `--output`, `--enable-embeddings`, `--enable-llm`), file I/O, error formatting, and stdout output - [x] T012 [P] [US1] Create baseline example files: `examples/ecp_petrobras.json`, `examples/content_presal_pt.md`, `examples/ecp_volkswagen.json`, `examples/content_northvolt_de.md`, `examples/ecp_apple.json`, `examples/content_tangential_es.md` - [x] T013 [US1] Verify end-to-end execution of CLI on example files and validate output JSON conforms strictly to schema **Checkpoint**: User Story 1 complete! Standalone CLI works end-to-end locally with Tier 1 deterministic classification. --- ## Phase 4: User Story 2 - Multilingual 24-Case Controlled Benchmark Suite (Priority: P2) **Goal**: Provide 24 paired test fixtures across all 6 languages (PT, EN, ES, DE, IT, FR) and 4 decision types, measuring ≥ 90% precision. **Independent Test**: Run `pytest tests/test_benchmark_24.py -v` and achieve 100% pass rate across the 24 controlled benchmark fixtures. ### Implementation for User Story 2 - [x] T014 [P] [US2] Create Portuguese benchmark fixtures (4 cases: `DIRECT`, `CONTEXTUAL`, `TANGENTIAL`, `NOT_RELATED`) in `tests/fixtures/benchmark_24/pt/` - [x] T015 [P] [US2] Create English benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/en/` - [x] T016 [P] [US2] Create Spanish benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/es/` - [x] T017 [P] [US2] Create German benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/de/` - [x] T018 [P] [US2] Create Italian benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/it/` - [x] T019 [P] [US2] Create French benchmark fixtures (4 cases) in `tests/fixtures/benchmark_24/fr/` - [x] T020 [US2] Implement parameterized benchmark runner in `tests/test_benchmark_24.py` verifying precision, score thresholds, evidence extraction, and language detection across all 24 cases **Checkpoint**: User Story 2 complete! Multilingual robustness validated across 24 test cases with ≥ 90% precision. --- ## Phase 5: Optional Adapters - Tier 2 (Embeddings) & Tier 3 (LLM) (Priority: P3) **Goal**: Provide clean, decoupled adapter interfaces for local embeddings and LLM fallback without introducing mandatory runtime dependencies. **Independent Test**: Validate that adapter interfaces load correctly and remain inert/disabled by default unless explicit CLI flags are provided. ### Implementation for User Story 3 - [x] T021 [P] [US3] Create base adapter interface in `src/adapters/base.py` (`BaseNLPAdapter`) - [x] T022 [P] [US3] Implement optional local vector similarity adapter stub in `src/adapters/embeddings.py` (activated via `--enable-embeddings`) - [x] T023 [P] [US3] Implement optional LLM fallback adapter stub in `src/adapters/llm.py` (activated via `--enable-llm`, disabled by default) - [x] T024 [US3] Wire adapter hooks into `src/classifier.py` ensuring Tier 1 handles clear cases without invoking Tier 2/Tier 3 **Checkpoint**: User Story 3 complete! Extensibility contracts established with zero breaking changes or mandatory cloud dependencies. --- ## Phase 6: Polish & Validation **Purpose**: Quickstart verification, code hygiene, and documentation audit - [x] T025 Execute all quickstart validation scenarios defined in `specs/001-multilingual-entity-classifier/quickstart.md` - [x] T026 [P] Run full test suite (`pytest -v --tb=short`) - [x] T027 Run `graphify update .` to index new implementation files into knowledge graph --- ## Dependencies & Execution Order ```mermaid flowchart TD Setup[Phase 1: Setup T001-T003] --> Foundational[Phase 2: Foundational T004-T007] Foundational --> US1[Phase 3: US1 Core CLI & Tier 1 T008-T013] US1 --> US2[Phase 4: US2 24-Case Benchmark T014-T020] US1 --> US3[Phase 5: US3 Optional Adapters T021-T024] US2 --> Polish[Phase 6: Polish T025-T027] US3 --> Polish ``` ### Implementation Strategy 1. **MVP (Phase 1 + 2 + 3)**: Delivers working `classify.py` with Tier 1 deterministic engine and example files. 2. **Benchmark Verification (Phase 4)**: Guarantees multilingual compliance on 24 controlled test fixtures. 3. **Adapter Stubs (Phase 5)**: Prepares the codebase for future embedding/LLM extensions cleanly.