6.8 KiB
Tasks: Multilingual NLP Entity Inherence Classifier (POC)
Feature: 001-multilingual-entity-classifier | Spec: spec.md | Plan: plan.md
Phase 1: Setup (Shared Infrastructure)
Purpose: Project initialization, directory structure, and environment configuration
- T001 Create project directories (
src/,src/adapters/,examples/,tests/,tests/fixtures/benchmark_24/) - T002 [P] Create
requirements.txtwith minimal test dependencies (pytest>=7.0) and optional dependencies commented - T003 [P] Create
.gitignorevalidation and project configuration inpyproject.tomlorsetup.cfg
Phase 2: Foundational (Blocking Prerequisites)
Purpose: Core data models, schema validation, and text processing utilities required by all user stories
CRITICAL: No classification engine logic can run until this phase is complete
- T004 Implement data models and schema validation in
src/models.py(ECPSnapshot,ClassificationResult,ClassificationError, and decision enums) - T005 [P] Implement Markdown parsing and evidence snippet extractor in
src/parser.py - T006 [P] Implement lightweight multilingual language detector & diacritic normalizer for the 6 languages in
src/language.py - T007 Implement unit tests for models, parser, and language detector in
tests/test_models.pyandtests/test_language.py
Checkpoint: Foundation ready - models, language detection, and Markdown parsing validated.
Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP]
Goal: Deliver a functioning CLI (classify.py) executing Tier 1 deterministic evaluation (aliases, anchors, negative anchors, graph matches) on ECP Snapshots and Markdown documents.
Independent Test: Run python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/test.json and verify decision: "DIRECT_INHERENT", is_inherent: true, confidence, and evidence.
Tests for User Story 1
- T008 [P] [US1] Create CLI contract and execution tests in
tests/test_cli.py(testing--ecp,--content,--output, stdout fallback, and exit codes) - T009 [P] [US1] Create deterministic decision logic unit tests in
tests/test_classifier.pyforDIRECT_INHERENT,CONTEXTUAL_INHERENT,TANGENTIAL, andNOT_RELATED
Implementation for User Story 1
- T010 [US1] Implement Tier 1 deterministic matching engine in
src/classifier.py(alias matching, negative anchor suppression, graph node resolution, scoring algorithm, and derivedis_inherentboolean) - T011 [US1] Implement main CLI script
classify.pyhandling argument parsing (--ecp,--content,--output,--enable-embeddings,--enable-llm), file I/O, error formatting, and stdout output - T012 [P] [US1] Create baseline example files:
examples/ecp_petrobras.json,examples/content_presal_pt.md,examples/ecp_volkswagen.json,examples/content_northvolt_de.md,examples/ecp_apple.json,examples/content_tangential_es.md - T013 [US1] Verify end-to-end execution of CLI on example files and validate output JSON conforms strictly to schema
Checkpoint: User Story 1 complete! Standalone CLI works end-to-end locally with Tier 1 deterministic classification.
Phase 4: User Story 2 - Multilingual 24-Case Controlled Benchmark Suite (Priority: P2)
Goal: Provide 24 paired test fixtures across all 6 languages (PT, EN, ES, DE, IT, FR) and 4 decision types, measuring ≥ 90% precision.
Independent Test: Run pytest tests/test_benchmark_24.py -v and achieve 100% pass rate across the 24 controlled benchmark fixtures.
Implementation for User Story 2
- T014 [P] [US2] Create Portuguese benchmark fixtures (4 cases:
DIRECT,CONTEXTUAL,TANGENTIAL,NOT_RELATED) intests/fixtures/benchmark_24/pt/ - T015 [P] [US2] Create English benchmark fixtures (4 cases) in
tests/fixtures/benchmark_24/en/ - T016 [P] [US2] Create Spanish benchmark fixtures (4 cases) in
tests/fixtures/benchmark_24/es/ - T017 [P] [US2] Create German benchmark fixtures (4 cases) in
tests/fixtures/benchmark_24/de/ - T018 [P] [US2] Create Italian benchmark fixtures (4 cases) in
tests/fixtures/benchmark_24/it/ - T019 [P] [US2] Create French benchmark fixtures (4 cases) in
tests/fixtures/benchmark_24/fr/ - T020 [US2] Implement parameterized benchmark runner in
tests/test_benchmark_24.pyverifying precision, score thresholds, evidence extraction, and language detection across all 24 cases
Checkpoint: User Story 2 complete! Multilingual robustness validated across 24 test cases with ≥ 90% precision.
Phase 5: Optional Adapters - Tier 2 (Embeddings) & Tier 3 (LLM) (Priority: P3)
Goal: Provide clean, decoupled adapter interfaces for local embeddings and LLM fallback without introducing mandatory runtime dependencies.
Independent Test: Validate that adapter interfaces load correctly and remain inert/disabled by default unless explicit CLI flags are provided.
Implementation for User Story 3
- T021 [P] [US3] Create base adapter interface in
src/adapters/base.py(BaseNLPAdapter) - T022 [P] [US3] Implement optional local vector similarity adapter stub in
src/adapters/embeddings.py(activated via--enable-embeddings) - T023 [P] [US3] Implement optional LLM fallback adapter stub in
src/adapters/llm.py(activated via--enable-llm, disabled by default) - T024 [US3] Wire adapter hooks into
src/classifier.pyensuring Tier 1 handles clear cases without invoking Tier 2/Tier 3
Checkpoint: User Story 3 complete! Extensibility contracts established with zero breaking changes or mandatory cloud dependencies.
Phase 6: Polish & Validation
Purpose: Quickstart verification, code hygiene, and documentation audit
- T025 Execute all quickstart validation scenarios defined in
specs/001-multilingual-entity-classifier/quickstart.md - T026 [P] Run full test suite (
pytest -v --tb=short) - T027 Run
graphify update .to index new implementation files into knowledge graph
Dependencies & Execution Order
flowchart TD
Setup[Phase 1: Setup T001-T003] --> Foundational[Phase 2: Foundational T004-T007]
Foundational --> US1[Phase 3: US1 Core CLI & Tier 1 T008-T013]
US1 --> US2[Phase 4: US2 24-Case Benchmark T014-T020]
US1 --> US3[Phase 5: US3 Optional Adapters T021-T024]
US2 --> Polish[Phase 6: Polish T025-T027]
US3 --> Polish
Implementation Strategy
- MVP (Phase 1 + 2 + 3): Delivers working
classify.pywith Tier 1 deterministic engine and example files. - Benchmark Verification (Phase 4): Guarantees multilingual compliance on 24 controlled test fixtures.
- Adapter Stubs (Phase 5): Prepares the codebase for future embedding/LLM extensions cleanly.