feat(classifier): add multilingual ECP inherence classifier POC

This commit is contained in:
2026-08-20 00:51:02 -03:00
parent d371b81aa4
commit 67cc40f91a
175 changed files with 30399 additions and 703 deletions
@@ -0,0 +1,128 @@
# Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
**Branch**: `001-multilingual-entity-classifier` | **Date**: 2026-08-19 | **Spec**: [spec.md](./spec.md)
**Input**: Feature specification from `specs/001-multilingual-entity-classifier/spec.md`
---
## Summary
Implement a lightweight, standalone Python CLI tool (`classify.py`) that evaluates whether a Markdown document is inherent to a target entity defined by an Entity Context Profile (ECP Snapshot JSON). The architecture implements a Tier 1 deterministic matching engine for 6 core languages (PT, EN, ES, DE, IT, FR), with clean optional adapter hooks for Tier 2 (embeddings) and Tier 3 (LLM fallback), fully validated against a controlled 24-case benchmark suite.
---
## Technical Context
**Language/Version**: Python 3.10+ (standard library for Tier 1 core execution).
**Primary Dependencies**:
- Core: standard library (`json`, `re`, `argparse`, `pathlib`, `typing`).
- Testing & Validation (mandatory): `pytest>=7.0` (only required dependency in `requirements.txt`).
- Optional (Tier 2 / Tier 3 adapters): `sentence-transformers`, `httpx` / `openai` (strictly optional, disabled by default).
**Storage**: None (file-in / file-out via CLI, no database).
**Testing**: `pytest` running unit tests and the 24-case controlled benchmark suite (`6 languages × 4 decision types`) via `tests/test_benchmark_24.py`.
**Target Platform**: Cross-platform (Windows / Linux / macOS).
**Project Type**: Standalone CLI script & modular core library (`classify.py` + `src/`).
**Performance Goals**: < 200ms execution time for Tier 1 deterministic evaluation on documents < 2,000 words.
**Constraints**:
- Zero mandatory external network calls or cloud API keys required to execute the POC or pass tests.
- Zero server/API dependencies (no FastAPI, no workers, no queues).
- Decoupled from live graph databases (consumes pre-materialized ECP JSON snapshots).
**Scale/Scope**: POC scope with minimum 24 explicit benchmark test fixtures.
---
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Principle / Gate | Status | Notes |
|---|---|---|
| **I. Library / Script First** | **PASS** | Standalone Python module structure, easily importable and script-callable. |
| **II. CLI Interface** | **PASS** | Clean standard CLI: `python classify.py --ecp <json> --content <md> --output <json>`, supports file paths and stdout output. |
| **III. Test-First (TDD)** | **PASS** | 24-case controlled benchmark matrix defined before code implementation. |
| **IV. Simplicity & YAGNI** | **PASS** | No premature REST API, no DB, no worker queues, no heavy frameworks. |
---
## Project Structure
### Documentation (this feature)
```text
specs/001-multilingual-entity-classifier/
├── spec.md # Feature specification
├── plan.md # Implementation plan (this file)
├── research.md # Technical research & decisions
├── data-model.md # Schemas & data contracts
├── quickstart.md # Validation & usage guide
├── checklists/
│ └── requirements.md # Quality checklist
├── contracts/
│ └── cli-contract.md # CLI input/output contract
└── tasks.md # Implementation tasks (/speckit-tasks command)
```
### Source Code (repository root)
```text
classify.py # Main CLI entrypoint script
src/
├── __init__.py
├── models.py # Dataclasses & schema validators (ECPSnapshot, Result, Error)
├── language.py # Lightweight multilingual detector & normalizer (6 languages)
├── parser.py # Markdown content parser & excerpt extractor
├── classifier.py # Core classification engine & decision logic (Tier 1 core)
└── adapters/
├── __init__.py
├── base.py # Base abstract adapter interfaces
├── embeddings.py # Optional Tier 2 embeddings adapter (disabled by default)
└── llm.py # Optional Tier 3 LLM fallback adapter (disabled by default)
examples/
├── ecp_petrobras.json # Example ECP Snapshot (PT)
├── ecp_volkswagen.json # Example ECP Snapshot (DE)
├── ecp_apple.json # Example ECP Snapshot (ES/EN)
├── content_presal_pt.md # Example Markdown Content (PT)
├── content_northvolt_de.md # Example Markdown Content (DE)
└── content_tangential_es.md # Example Markdown Content (ES)
tests/
├── __init__.py
├── test_cli.py # CLI argument parsing, flags, file I/O, stdout emission, exit codes
├── test_models.py # ECP snapshot parsing and structured error handling
├── test_language.py # Language detection and text normalization tests
├── test_classifier.py # Unit tests for decision rules (DIRECT, CONTEXTUAL, TANGENTIAL, NOT_RELATED)
├── test_benchmark_24.py # Controlled 24-case benchmark runner (6 languages x 4 decisions)
└── fixtures/
└── benchmark_24/ # 24 paired test cases (ecp_*.json + content_*.md + expected_*.json)
├── pt/
├── en/
├── es/
├── de/
├── it/
└── fr/
```
**Structure Decision**: Single modular project with root CLI `classify.py` and clear separation of models, language normalization, deterministic classification, and optional adapter stubs under `src/`.
---
## Complexity Tracking
> No constitution violations detected. Design enforces absolute simplicity and strict POC constraints.
| Component | Choice | Simpler Alternative Rejected Because |
|---|---|---|
| CLI vs REST API | Standalone CLI | REST API adds unnecessary network latency, FastAPI dependencies, and server management for a POC. |
| Deterministic Core vs Pure LLM | Deterministic Tier 1 Core | Pure LLM is expensive, non-deterministic, slow, and requires mandatory external API keys. |
| In-Memory Snapshot vs Live Graph Query | Materialized JSON Snapshot | Live Neo4j queries couple classifier runtime to external infrastructure and break test reproducibility. |