Files
TextNLPClassifierApp/specs/001-multilingual-entity-classifier/plan.md
T

6.1 KiB
Raw Blame History

Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)

Branch: 001-multilingual-entity-classifier | Date: 2026-08-19 | Spec: spec.md

Input: Feature specification from specs/001-multilingual-entity-classifier/spec.md


Summary

Implement a lightweight, standalone Python CLI tool (classify.py) that evaluates whether a Markdown document is inherent to a target entity defined by an Entity Context Profile (ECP Snapshot JSON). The architecture implements a Tier 1 deterministic matching engine for 6 core languages (PT, EN, ES, DE, IT, FR), with clean optional adapter hooks for Tier 2 (embeddings) and Tier 3 (LLM fallback), fully validated against a controlled 24-case benchmark suite.


Technical Context

Language/Version: Python 3.10+ (standard library for Tier 1 core execution).

Primary Dependencies:

  • Core: standard library (json, re, argparse, pathlib, typing).
  • Testing & Validation (mandatory): pytest>=7.0 (only required dependency in requirements.txt).
  • Optional (Tier 2 / Tier 3 adapters): sentence-transformers, httpx / openai (strictly optional, disabled by default).

Storage: None (file-in / file-out via CLI, no database).

Testing: pytest running unit tests and the 24-case controlled benchmark suite (6 languages × 4 decision types) via tests/test_benchmark_24.py.

Target Platform: Cross-platform (Windows / Linux / macOS).

Project Type: Standalone CLI script & modular core library (classify.py + src/).

Performance Goals: < 200ms execution time for Tier 1 deterministic evaluation on documents < 2,000 words.

Constraints:

  • Zero mandatory external network calls or cloud API keys required to execute the POC or pass tests.
  • Zero server/API dependencies (no FastAPI, no workers, no queues).
  • Decoupled from live graph databases (consumes pre-materialized ECP JSON snapshots).

Scale/Scope: POC scope with minimum 24 explicit benchmark test fixtures.


Constitution Check

GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.

Principle / Gate Status Notes
I. Library / Script First PASS Standalone Python module structure, easily importable and script-callable.
II. CLI Interface PASS Clean standard CLI: python classify.py --ecp <json> --content <md> --output <json>, supports file paths and stdout output.
III. Test-First (TDD) PASS 24-case controlled benchmark matrix defined before code implementation.
IV. Simplicity & YAGNI PASS No premature REST API, no DB, no worker queues, no heavy frameworks.

Project Structure

Documentation (this feature)

specs/001-multilingual-entity-classifier/
├── spec.md              # Feature specification
├── plan.md              # Implementation plan (this file)
├── research.md          # Technical research & decisions
├── data-model.md        # Schemas & data contracts
├── quickstart.md        # Validation & usage guide
├── checklists/
│   └── requirements.md  # Quality checklist
├── contracts/
│   └── cli-contract.md  # CLI input/output contract
└── tasks.md             # Implementation tasks (/speckit-tasks command)

Source Code (repository root)

classify.py                  # Main CLI entrypoint script

src/
├── __init__.py
├── models.py                # Dataclasses & schema validators (ECPSnapshot, Result, Error)
├── language.py              # Lightweight multilingual detector & normalizer (6 languages)
├── parser.py                # Markdown content parser & excerpt extractor
├── classifier.py            # Core classification engine & decision logic (Tier 1 core)
└── adapters/
    ├── __init__.py
    ├── base.py              # Base abstract adapter interfaces
    ├── embeddings.py        # Optional Tier 2 embeddings adapter (disabled by default)
    └── llm.py               # Optional Tier 3 LLM fallback adapter (disabled by default)

examples/
├── ecp_petrobras.json       # Example ECP Snapshot (PT)
├── ecp_volkswagen.json      # Example ECP Snapshot (DE)
├── ecp_apple.json           # Example ECP Snapshot (ES/EN)
├── content_presal_pt.md     # Example Markdown Content (PT)
├── content_northvolt_de.md  # Example Markdown Content (DE)
└── content_tangential_es.md # Example Markdown Content (ES)

tests/
├── __init__.py
├── test_cli.py              # CLI argument parsing, flags, file I/O, stdout emission, exit codes
├── test_models.py           # ECP snapshot parsing and structured error handling
├── test_language.py         # Language detection and text normalization tests
├── test_classifier.py       # Unit tests for decision rules (DIRECT, CONTEXTUAL, TANGENTIAL, NOT_RELATED)
├── test_benchmark_24.py     # Controlled 24-case benchmark runner (6 languages x 4 decisions)
└── fixtures/
    └── benchmark_24/        # 24 paired test cases (ecp_*.json + content_*.md + expected_*.json)
        ├── pt/
        ├── en/
        ├── es/
        ├── de/
        ├── it/
        └── fr/

Structure Decision: Single modular project with root CLI classify.py and clear separation of models, language normalization, deterministic classification, and optional adapter stubs under src/.


Complexity Tracking

No constitution violations detected. Design enforces absolute simplicity and strict POC constraints.

Component Choice Simpler Alternative Rejected Because
CLI vs REST API Standalone CLI REST API adds unnecessary network latency, FastAPI dependencies, and server management for a POC.
Deterministic Core vs Pure LLM Deterministic Tier 1 Core Pure LLM is expensive, non-deterministic, slow, and requires mandatory external API keys.
In-Memory Snapshot vs Live Graph Query Materialized JSON Snapshot Live Neo4j queries couple classifier runtime to external infrastructure and break test reproducibility.