Files

6.8 KiB

Tasks: Multilingual NLP Entity Inherence Classifier (POC)

Feature: 001-multilingual-entity-classifier | Spec: spec.md | Plan: plan.md


Phase 1: Setup (Shared Infrastructure)

Purpose: Project initialization, directory structure, and environment configuration

  • T001 Create project directories (src/, src/adapters/, examples/, tests/, tests/fixtures/benchmark_24/)
  • T002 [P] Create requirements.txt with minimal test dependencies (pytest>=7.0) and optional dependencies commented
  • T003 [P] Create .gitignore validation and project configuration in pyproject.toml or setup.cfg

Phase 2: Foundational (Blocking Prerequisites)

Purpose: Core data models, schema validation, and text processing utilities required by all user stories

CRITICAL: No classification engine logic can run until this phase is complete

  • T004 Implement data models and schema validation in src/models.py (ECPSnapshot, ClassificationResult, ClassificationError, and decision enums)
  • T005 [P] Implement Markdown parsing and evidence snippet extractor in src/parser.py
  • T006 [P] Implement lightweight multilingual language detector & diacritic normalizer for the 6 languages in src/language.py
  • T007 Implement unit tests for models, parser, and language detector in tests/test_models.py and tests/test_language.py

Checkpoint: Foundation ready - models, language detection, and Markdown parsing validated.


Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP]

Goal: Deliver a functioning CLI (classify.py) executing Tier 1 deterministic evaluation (aliases, anchors, negative anchors, graph matches) on ECP Snapshots and Markdown documents.

Independent Test: Run python classify.py --ecp examples/ecp_petrobras.json --content examples/content_presal_pt.md --output out/test.json and verify decision: "DIRECT_INHERENT", is_inherent: true, confidence, and evidence.

Tests for User Story 1

  • T008 [P] [US1] Create CLI contract and execution tests in tests/test_cli.py (testing --ecp, --content, --output, stdout fallback, and exit codes)
  • T009 [P] [US1] Create deterministic decision logic unit tests in tests/test_classifier.py for DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, and NOT_RELATED

Implementation for User Story 1

  • T010 [US1] Implement Tier 1 deterministic matching engine in src/classifier.py (alias matching, negative anchor suppression, graph node resolution, scoring algorithm, and derived is_inherent boolean)
  • T011 [US1] Implement main CLI script classify.py handling argument parsing (--ecp, --content, --output, --enable-embeddings, --enable-llm), file I/O, error formatting, and stdout output
  • T012 [P] [US1] Create baseline example files: examples/ecp_petrobras.json, examples/content_presal_pt.md, examples/ecp_volkswagen.json, examples/content_northvolt_de.md, examples/ecp_apple.json, examples/content_tangential_es.md
  • T013 [US1] Verify end-to-end execution of CLI on example files and validate output JSON conforms strictly to schema

Checkpoint: User Story 1 complete! Standalone CLI works end-to-end locally with Tier 1 deterministic classification.


Phase 4: User Story 2 - Multilingual 24-Case Controlled Benchmark Suite (Priority: P2)

Goal: Provide 24 paired test fixtures across all 6 languages (PT, EN, ES, DE, IT, FR) and 4 decision types, measuring ≥ 90% precision.

Independent Test: Run pytest tests/test_benchmark_24.py -v and achieve 100% pass rate across the 24 controlled benchmark fixtures.

Implementation for User Story 2

  • T014 [P] [US2] Create Portuguese benchmark fixtures (4 cases: DIRECT, CONTEXTUAL, TANGENTIAL, NOT_RELATED) in tests/fixtures/benchmark_24/pt/
  • T015 [P] [US2] Create English benchmark fixtures (4 cases) in tests/fixtures/benchmark_24/en/
  • T016 [P] [US2] Create Spanish benchmark fixtures (4 cases) in tests/fixtures/benchmark_24/es/
  • T017 [P] [US2] Create German benchmark fixtures (4 cases) in tests/fixtures/benchmark_24/de/
  • T018 [P] [US2] Create Italian benchmark fixtures (4 cases) in tests/fixtures/benchmark_24/it/
  • T019 [P] [US2] Create French benchmark fixtures (4 cases) in tests/fixtures/benchmark_24/fr/
  • T020 [US2] Implement parameterized benchmark runner in tests/test_benchmark_24.py verifying precision, score thresholds, evidence extraction, and language detection across all 24 cases

Checkpoint: User Story 2 complete! Multilingual robustness validated across 24 test cases with ≥ 90% precision.


Phase 5: Optional Adapters - Tier 2 (Embeddings) & Tier 3 (LLM) (Priority: P3)

Goal: Provide clean, decoupled adapter interfaces for local embeddings and LLM fallback without introducing mandatory runtime dependencies.

Independent Test: Validate that adapter interfaces load correctly and remain inert/disabled by default unless explicit CLI flags are provided.

Implementation for User Story 3

  • T021 [P] [US3] Create base adapter interface in src/adapters/base.py (BaseNLPAdapter)
  • T022 [P] [US3] Implement optional local vector similarity adapter stub in src/adapters/embeddings.py (activated via --enable-embeddings)
  • T023 [P] [US3] Implement optional LLM fallback adapter stub in src/adapters/llm.py (activated via --enable-llm, disabled by default)
  • T024 [US3] Wire adapter hooks into src/classifier.py ensuring Tier 1 handles clear cases without invoking Tier 2/Tier 3

Checkpoint: User Story 3 complete! Extensibility contracts established with zero breaking changes or mandatory cloud dependencies.


Phase 6: Polish & Validation

Purpose: Quickstart verification, code hygiene, and documentation audit

  • T025 Execute all quickstart validation scenarios defined in specs/001-multilingual-entity-classifier/quickstart.md
  • T026 [P] Run full test suite (pytest -v --tb=short)
  • T027 Run graphify update . to index new implementation files into knowledge graph

Dependencies & Execution Order

flowchart TD
    Setup[Phase 1: Setup T001-T003] --> Foundational[Phase 2: Foundational T004-T007]
    Foundational --> US1[Phase 3: US1 Core CLI & Tier 1 T008-T013]
    US1 --> US2[Phase 4: US2 24-Case Benchmark T014-T020]
    US1 --> US3[Phase 5: US3 Optional Adapters T021-T024]
    US2 --> Polish[Phase 6: Polish T025-T027]
    US3 --> Polish

Implementation Strategy

  1. MVP (Phase 1 + 2 + 3): Delivers working classify.py with Tier 1 deterministic engine and example files.
  2. Benchmark Verification (Phase 4): Guarantees multilingual compliance on 24 controlled test fixtures.
  3. Adapter Stubs (Phase 5): Prepares the codebase for future embedding/LLM extensions cleanly.