4.0 KiB
4.0 KiB
Technical Research & Architecture Decisions (POC)
Feature: 001-multilingual-entity-classifier
Status: Completed
1. Technical Decisions & Tradeoffs
Decision 1: Execution Engine & CLI Architecture
- Decision: Standalone Python 3 script (
classify.py) with clean modular components undersrc/. - Rationale: Keeps the POC lightweight, zero-boilerplate, directly testable via standard command line and
pytest, perfectly aligned with Ponytail and SpecKit principles. - Alternatives Considered:
- FastAPI REST Service: Rejected (violates POC simplicity, introduces server overhead, unnecessary network latency for batch/CLI evaluation).
- Publishable Package / Setuptools: Rejected (premature abstraction before core classification logic is validated).
Decision 2: Multilingual Language Detection & Normalization (Tier 1 Core)
- Decision: Lightweight regex + heuristic n-gram / stopword profile detection for the 6 core languages (PT, EN, ES, DE, IT, FR), with fallback to standard library / optional
langdetectorlinguaif installed. Text normalization converts diacritics/casing for robust deterministic matching while preserving original excerpt offsets forevidence. - Rationale: Guarantees zero heavy mandatory dependencies for basic Tier 1 execution, sub-millisecond detection latency, and 100% offline capability.
- Alternatives Considered:
- Heavy Transformer-based language ID: Rejected (large model download, excessive latency for short/medium texts).
Decision 3: Materialized ECP Snapshot Contract & Matching Logic
- Decision: Parse self-contained ECP Snapshot JSON containing:
target_entity_id,target_name,aliases,domain,anchors,negative_anchors,graph_version,related_entities.- Matching pipeline evaluates:
- Direct entity match:
aliases+target_namein Markdown. - Negative anchor presence: if negative terms dominate context → downgrade or mark
NOT_RELATED. - Graph snapshot match:
related_entitiesfound in Markdown with associatedrelation_type,weight,confidence. - Domain/Anchor density: computes normalized score (0.0 to 1.0) and assigns one of the 4 decision categories (
DIRECT_INHERENT,CONTEXTUAL_INHERENT,TANGENTIAL,NOT_RELATED).
- Direct entity match:
- Rationale: Completely decouples inference from live graph database queries, enabling deterministic, fast, reproducible, and audit-friendly decisions.
Decision 4: Tier 2 (Embeddings) & Tier 3 (LLM) Optional Adapters
- Decision: Implement Tier 2 and Tier 3 as decoupled adapter interfaces (
EmbeddingProviderInterface,LLMProviderInterface).- By default,
Tier2andTier3are disabled (--enable-embeddings/--enable-llmCLI flags). - If Tier 1 has high confidence or unambiguous negative match, Tier 2/Tier 3 are skipped entirely.
- Zero mandatory external API keys are required to execute or test the POC.
- By default,
- Rationale: Satisfies the 3-tier hybrid requirement without imposing heavy dependencies (e.g. PyTorch/HuggingFace) or external API keys onto the baseline test suite.
Decision 5: Controlled 24-Case POC Benchmark Suite
- Decision: Construct a fixture suite of 24 controlled test cases:
- 6 Languages:
pt,en,es,de,it,fr - 4 Decision Types per language:
DIRECT_INHERENT,CONTEXTUAL_INHERENT,TANGENTIAL,NOT_RELATED - Total: 6 × 4 = 24 benchmark pairs of
(ecp_snapshot.json, content.md, expected_result.json).
- 6 Languages:
- Rationale: Provides an unambiguous, reproducible test bed to measure and verify the ≥ 90% precision requirement.
2. Standardized Error Handling Strategy
Standardized JSON error envelope emitted with non-zero exit code (1) when execution cannot complete:
{
"error_code": "invalid_ecp_json | invalid_markdown | unsupported_language | empty_content | missing_required_field",
"message": "Human readable description",
"details": {
"field": "aliases",
"reason": "Field 'aliases' must be an array of strings"
}
}