Files
TextNLPClassifierApp/specs/001-multilingual-entity-classifier/plan.md
T

129 lines
6.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
**Branch**: `001-multilingual-entity-classifier` | **Date**: 2026-08-19 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `specs/001-multilingual-entity-classifier/spec.md`
---
## Summary
Implement a lightweight, standalone Python CLI tool (`classify.py`) that evaluates whether a Markdown document is inherent to a target entity defined by an Entity Context Profile (ECP Snapshot JSON). The architecture implements a Tier 1 deterministic matching engine for 6 core languages (PT, EN, ES, DE, IT, FR), with clean optional adapter hooks for Tier 2 (embeddings) and Tier 3 (LLM fallback), fully validated against a controlled 24-case benchmark suite.
---
## Technical Context
**Language/Version**: Python 3.10+ (standard library for Tier 1 core execution).
**Primary Dependencies**:
- Core: standard library (`json`, `re`, `argparse`, `pathlib`, `typing`).
- Testing & Validation (mandatory): `pytest>=7.0` (only required dependency in `requirements.txt`).
- Optional (Tier 2 / Tier 3 adapters): `sentence-transformers`, `httpx` / `openai` (strictly optional, disabled by default).
**Storage**: None (file-in / file-out via CLI, no database).
**Testing**: `pytest` running unit tests and the 24-case controlled benchmark suite (`6 languages × 4 decision types`) via `tests/test_benchmark_24.py`.
**Target Platform**: Cross-platform (Windows / Linux / macOS).
**Project Type**: Standalone CLI script & modular core library (`classify.py` + `src/`).
**Performance Goals**: < 200ms execution time for Tier 1 deterministic evaluation on documents < 2,000 words.
**Constraints**:
- Zero mandatory external network calls or cloud API keys required to execute the POC or pass tests.
- Zero server/API dependencies (no FastAPI, no workers, no queues).
- Decoupled from live graph databases (consumes pre-materialized ECP JSON snapshots).
**Scale/Scope**: POC scope with minimum 24 explicit benchmark test fixtures.
---
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle / Gate | Status | Notes |
|---|---|---|
| **I. Library / Script First** | **PASS** | Standalone Python module structure, easily importable and script-callable. |
| **II. CLI Interface** | **PASS** | Clean standard CLI: `python classify.py --ecp <json> --content <md> --output <json>`, supports file paths and stdout output. |
| **III. Test-First (TDD)** | **PASS** | 24-case controlled benchmark matrix defined before code implementation. |
| **IV. Simplicity & YAGNI** | **PASS** | No premature REST API, no DB, no worker queues, no heavy frameworks. |
---
## Project Structure
### Documentation (this feature)
```text
specs/001-multilingual-entity-classifier/
├── spec.md # Feature specification
├── plan.md # Implementation plan (this file)
├── research.md # Technical research & decisions
├── data-model.md # Schemas & data contracts
├── quickstart.md # Validation & usage guide
├── checklists/
│ └── requirements.md # Quality checklist
├── contracts/
│ └── cli-contract.md # CLI input/output contract
└── tasks.md # Implementation tasks (/speckit-tasks command)
```
### Source Code (repository root)
```text
classify.py # Main CLI entrypoint script
src/
├── __init__.py
├── models.py # Dataclasses & schema validators (ECPSnapshot, Result, Error)
├── language.py # Lightweight multilingual detector & normalizer (6 languages)
├── parser.py # Markdown content parser & excerpt extractor
├── classifier.py # Core classification engine & decision logic (Tier 1 core)
└── adapters/
├── __init__.py
├── base.py # Base abstract adapter interfaces
├── embeddings.py # Optional Tier 2 embeddings adapter (disabled by default)
└── llm.py # Optional Tier 3 LLM fallback adapter (disabled by default)
examples/
├── ecp_petrobras.json # Example ECP Snapshot (PT)
├── ecp_volkswagen.json # Example ECP Snapshot (DE)
├── ecp_apple.json # Example ECP Snapshot (ES/EN)
├── content_presal_pt.md # Example Markdown Content (PT)
├── content_northvolt_de.md # Example Markdown Content (DE)
└── content_tangential_es.md # Example Markdown Content (ES)
tests/
├── __init__.py
├── test_cli.py # CLI argument parsing, flags, file I/O, stdout emission, exit codes
├── test_models.py # ECP snapshot parsing and structured error handling
├── test_language.py # Language detection and text normalization tests
├── test_classifier.py # Unit tests for decision rules (DIRECT, CONTEXTUAL, TANGENTIAL, NOT_RELATED)
├── test_benchmark_24.py # Controlled 24-case benchmark runner (6 languages x 4 decisions)
└── fixtures/
└── benchmark_24/ # 24 paired test cases (ecp_*.json + content_*.md + expected_*.json)
├── pt/
├── en/
├── es/
├── de/
├── it/
└── fr/
```
**Structure Decision**: Single modular project with root CLI `classify.py` and clear separation of models, language normalization, deterministic classification, and optional adapter stubs under `src/`.
---
## Complexity Tracking
> No constitution violations detected. Design enforces absolute simplicity and strict POC constraints.
| Component | Choice | Simpler Alternative Rejected Because |
|---|---|---|
| CLI vs REST API | Standalone CLI | REST API adds unnecessary network latency, FastAPI dependencies, and server management for a POC. |
| Deterministic Core vs Pure LLM | Deterministic Tier 1 Core | Pure LLM is expensive, non-deterministic, slow, and requires mandatory external API keys. |
| In-Memory Snapshot vs Live Graph Query | Materialized JSON Snapshot | Live Neo4j queries couple classifier runtime to external infrastructure and break test reproducibility. |