# Technical Research & Architecture Decisions (POC) **Feature**: `001-multilingual-entity-classifier` **Status**: Completed --- ## 1. Technical Decisions & Tradeoffs ### Decision 1: Execution Engine & CLI Architecture - **Decision**: Standalone Python 3 script (`classify.py`) with clean modular components under `src/`. - **Rationale**: Keeps the POC lightweight, zero-boilerplate, directly testable via standard command line and `pytest`, perfectly aligned with Ponytail and SpecKit principles. - **Alternatives Considered**: - *FastAPI REST Service*: Rejected (violates POC simplicity, introduces server overhead, unnecessary network latency for batch/CLI evaluation). - *Publishable Package / Setuptools*: Rejected (premature abstraction before core classification logic is validated). ### Decision 2: Multilingual Language Detection & Normalization (Tier 1 Core) - **Decision**: Lightweight regex + heuristic n-gram / stopword profile detection for the 6 core languages (PT, EN, ES, DE, IT, FR), with fallback to standard library / optional `langdetect` or `lingua` if installed. Text normalization converts diacritics/casing for robust deterministic matching while preserving original excerpt offsets for `evidence`. - **Rationale**: Guarantees zero heavy mandatory dependencies for basic Tier 1 execution, sub-millisecond detection latency, and 100% offline capability. - **Alternatives Considered**: - *Heavy Transformer-based language ID*: Rejected (large model download, excessive latency for short/medium texts). ### Decision 3: Materialized ECP Snapshot Contract & Matching Logic - **Decision**: Parse self-contained ECP Snapshot JSON containing: - `target_entity_id`, `target_name`, `aliases`, `domain`, `anchors`, `negative_anchors`, `graph_version`, `related_entities`. - Matching pipeline evaluates: 1. Direct entity match: `aliases` + `target_name` in Markdown. 2. Negative anchor presence: if negative terms dominate context → downgrade or mark `NOT_RELATED`. 3. Graph snapshot match: `related_entities` found in Markdown with associated `relation_type`, `weight`, `confidence`. 4. Domain/Anchor density: computes normalized score (0.0 to 1.0) and assigns one of the 4 decision categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`). - **Rationale**: Completely decouples inference from live graph database queries, enabling deterministic, fast, reproducible, and audit-friendly decisions. ### Decision 4: Tier 2 (Embeddings) & Tier 3 (LLM) Optional Adapters - **Decision**: Implement Tier 2 and Tier 3 as decoupled adapter interfaces (`EmbeddingProviderInterface`, `LLMProviderInterface`). - By default, `Tier2` and `Tier3` are disabled (`--enable-embeddings` / `--enable-llm` CLI flags). - If Tier 1 has high confidence or unambiguous negative match, Tier 2/Tier 3 are skipped entirely. - Zero mandatory external API keys are required to execute or test the POC. - **Rationale**: Satisfies the 3-tier hybrid requirement without imposing heavy dependencies (e.g. PyTorch/HuggingFace) or external API keys onto the baseline test suite. ### Decision 5: Controlled 24-Case POC Benchmark Suite - **Decision**: Construct a fixture suite of 24 controlled test cases: - 6 Languages: `pt`, `en`, `es`, `de`, `it`, `fr` - 4 Decision Types per language: `DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED` - Total: 6 × 4 = 24 benchmark pairs of `(ecp_snapshot.json, content.md, expected_result.json)`. - **Rationale**: Provides an unambiguous, reproducible test bed to measure and verify the ≥ 90% precision requirement. --- ## 2. Standardized Error Handling Strategy Standardized JSON error envelope emitted with non-zero exit code (1) when execution cannot complete: ```json { "error_code": "invalid_ecp_json | invalid_markdown | unsupported_language | empty_content | missing_required_field", "message": "Human readable description", "details": { "field": "aliases", "reason": "Field 'aliases' must be an array of strings" } } ```