Files
TextNLPClassifierApp/specs/001-multilingual-entity-classifier/research.md
T

4.0 KiB
Raw Blame History

Technical Research & Architecture Decisions (POC)

Feature: 001-multilingual-entity-classifier Status: Completed


1. Technical Decisions & Tradeoffs

Decision 1: Execution Engine & CLI Architecture

  • Decision: Standalone Python 3 script (classify.py) with clean modular components under src/.
  • Rationale: Keeps the POC lightweight, zero-boilerplate, directly testable via standard command line and pytest, perfectly aligned with Ponytail and SpecKit principles.
  • Alternatives Considered:
    • FastAPI REST Service: Rejected (violates POC simplicity, introduces server overhead, unnecessary network latency for batch/CLI evaluation).
    • Publishable Package / Setuptools: Rejected (premature abstraction before core classification logic is validated).

Decision 2: Multilingual Language Detection & Normalization (Tier 1 Core)

  • Decision: Lightweight regex + heuristic n-gram / stopword profile detection for the 6 core languages (PT, EN, ES, DE, IT, FR), with fallback to standard library / optional langdetect or lingua if installed. Text normalization converts diacritics/casing for robust deterministic matching while preserving original excerpt offsets for evidence.
  • Rationale: Guarantees zero heavy mandatory dependencies for basic Tier 1 execution, sub-millisecond detection latency, and 100% offline capability.
  • Alternatives Considered:
    • Heavy Transformer-based language ID: Rejected (large model download, excessive latency for short/medium texts).

Decision 3: Materialized ECP Snapshot Contract & Matching Logic

  • Decision: Parse self-contained ECP Snapshot JSON containing:
    • target_entity_id, target_name, aliases, domain, anchors, negative_anchors, graph_version, related_entities.
    • Matching pipeline evaluates:
      1. Direct entity match: aliases + target_name in Markdown.
      2. Negative anchor presence: if negative terms dominate context → downgrade or mark NOT_RELATED.
      3. Graph snapshot match: related_entities found in Markdown with associated relation_type, weight, confidence.
      4. Domain/Anchor density: computes normalized score (0.0 to 1.0) and assigns one of the 4 decision categories (DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED).
  • Rationale: Completely decouples inference from live graph database queries, enabling deterministic, fast, reproducible, and audit-friendly decisions.

Decision 4: Tier 2 (Embeddings) & Tier 3 (LLM) Optional Adapters

  • Decision: Implement Tier 2 and Tier 3 as decoupled adapter interfaces (EmbeddingProviderInterface, LLMProviderInterface).
    • By default, Tier2 and Tier3 are disabled (--enable-embeddings / --enable-llm CLI flags).
    • If Tier 1 has high confidence or unambiguous negative match, Tier 2/Tier 3 are skipped entirely.
    • Zero mandatory external API keys are required to execute or test the POC.
  • Rationale: Satisfies the 3-tier hybrid requirement without imposing heavy dependencies (e.g. PyTorch/HuggingFace) or external API keys onto the baseline test suite.

Decision 5: Controlled 24-Case POC Benchmark Suite

  • Decision: Construct a fixture suite of 24 controlled test cases:
    • 6 Languages: pt, en, es, de, it, fr
    • 4 Decision Types per language: DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED
    • Total: 6 × 4 = 24 benchmark pairs of (ecp_snapshot.json, content.md, expected_result.json).
  • Rationale: Provides an unambiguous, reproducible test bed to measure and verify the ≥ 90% precision requirement.

2. Standardized Error Handling Strategy

Standardized JSON error envelope emitted with non-zero exit code (1) when execution cannot complete:

{
  "error_code": "invalid_ecp_json | invalid_markdown | unsupported_language | empty_content | missing_required_field",
  "message": "Human readable description",
  "details": {
    "field": "aliases",
    "reason": "Field 'aliases' must be an array of strings"
  }
}