Files
TextNLPClassifierApp/specs/006-article-consolidation-runtime/research.md
T

12 KiB
Raw Blame History

Research & Technical Decisions: Article Consolidation and Hygiene Runtime

Feature Branch: 006-article-consolidation-runtime
Date: 2026-08-23
Status: Complete


1. Executive Summary & Architectural Scope

The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot.

The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002):

  • Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines).
  • SQLite in WAL mode with short transactions and configurable lock timeout.
  • Local filesystem for atomic staged file writes and renames.
  • Agnostic Model Gateway managing two logical roles (runtime_primary and runtime_fallback) configured with certified low-cost models (defaults: Groq with openai/gpt-oss-20b and DeepSeek with deepseek-v4-flash).
  • 100% LLM extractive content hygiene with strict 10-step validation harness.
  • Grounded candidate selection and controlled micro-repairs without regular expressions (re) or manual keyword dictionaries.
  • Input size limit validation (FR-056) failing in a controlled manner with INVALID_ARTICLE_SCHEMA before provider calls if exceeded.
  • Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (src.classifier.classify_text).
  • Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval.
  • Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (pending_telemetry) during network outages.

2. Definitive Dependency Selections & Qualitative Impact Analysis

In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria:

Component / Task Chosen Solution Standard Library Alternative Security Impact Maintenance Impact License Size Impact Startup Impact
HTML DOM Parsing beautifulsoup4 (with html.parser) html.parser directly Low. Pure Python, robust against malformed HTML. Low. Stable and mature. MIT Minimal Negligible
Markdown AST Parsing marko None in stdlib for CommonMark AST Low. Standard CommonMark AST compliance. Low. Pure Python parser. MIT Minimal Negligible
JSON Schema Validation jsonschema + referencing Manual validation code Low. Standard Draft 2020-12 validator with immutable local schema registry. Low. Reference standard. MIT Minimal Negligible
HTTP Client / Gateway httpx urllib.request Low. Modern HTTP client with connection pooling. Low. High adoption. BSD-3-Clause Minimal Negligible
Date Parsing & ISO 8601 python-dateutil datetime.fromisoformat Low. Handles diverse timezone & date formats. Low. Industry standard. Apache 2.0 / BSD Minimal Negligible
YAML Serialization pyyaml None in stdlib Low. Safe dump (yaml.safe_dump) prevents code execution. Low. Standard YAML library. MIT Minimal Negligible
Language Detection src.language (existing monorepo) N/A None. Reuses existing repository module. Zero new dependency. Monorepo Zero Zero
Diffing without Regex difflib.SequenceMatcher Stdlib difflib None. Standard library. None. Standard library. Python Zero Zero
Unicode Normalization unicodedata Stdlib unicodedata None. Standard library. None. Standard library. Python Zero Zero
URL Parsing urllib.parse Stdlib urllib.parse None. Standard library. None. Standard library. Python Zero Zero
State Store & Queue sqlite3 Stdlib sqlite3 None. Standard library. None. Standard library. Python Zero Zero
Observability SDK langfuse Direct HTTP calls Low. Official SDK (>=4.7). Low. Active upstream support. MIT Minimal Negligible

Note on Durable Queuing: Durable local queuing during observability outages is provided directly by the SQLite table pending_telemetry, not delegated to SDK memory buffers.


3. Concrete Architectural & Technical Decisions

Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046)

  • Decision: Orchestration is implemented as an explicit Python class StateMachine managing state transitions in SQLite:
    • received → validated → content_cleaned
    • content_cleaned → ecp_approved → enriched → completed_text
    • content_cleaned → ecp_rejected (terminal state, zero Markdown files generated)
    • Valid terminal failures → failed
  • Transition Recording: Every state transition records start_time, end_time, duration_ms, and result into SQLite table state_transitions.

Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044)

  • Decision: Agnostic Model Gateway (src.gateway.client) supporting two certified logical roles:
    • runtime_primary: Default configured as Groq with openai/gpt-oss-20b (or certified low-cost equivalent).
    • runtime_fallback: Default configured as DeepSeek with deepseek-v4-flash (or certified low-cost equivalent).
  • Release Metadata Certification: Build/packaging produces an immutable src/core/release-metadata.json packaged with the release containing:
    • release_version
    • runtime_config_sha256 (SHA-256 of the approved functional config)
    • prompts_hashes (SHA-256 of each prompt file)
    • schemas_versions (contract version identifiers)
    • certified_models (logical role to approved provider/model mappings)
    • ecp_classifier_config_hash (SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters) Preflight compares the loaded configuration against src/core/release-metadata.json. Any mismatch aborts with exit code 2.
  • Retry & Fallback Policy:
    • Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response).
    • Semantic failures (invalid schema, grounding violation) transition immediately to runtime_fallback without retrying on the same model.
  • Pricing & Operational Parameters: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint.

Decision 3: Input Size Limit Enforcement (FR-056)

  • Decision: Prior to candidate extraction or remote provider calls, input size is checked against limits.max_input_bytes. If exceeded, execution terminates immediately with controlled error INVALID_ARTICLE_SCHEMA (with structured detail "INPUT_EXCEEDS_SIZE_LIMIT"). Strategy is strictly fail_before_provider. Automatic unapproved truncation is strictly prohibited.

Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020)

  • Decision:
    • The runtime preserves all structurally valid candidate elements from all extractors.
    • difflib.SequenceMatcher calculates cross-extractor sequence similarity during candidate preparation to populate the equivalences list, providing consensus evidence to the LLM without deleting or merging candidates.
    • The selected_extractor provides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection.
    • Candidate IDs are opaque, stable within execution, and carry no quality judgment.

Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030)

  • Decision:
    • LLM returns ONLY candidate IDs and micro-repair operations in article_content_hygiene.
    • Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content.
    • Micro-repairs are validated across 5 closed categories (encoding, unicode, spacing, punctuation_corruption, obvious_typo) using unicodedata and difflib without regex.
    • Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted.
    • Ungrounded candidate IDs trigger GROUNDING_VIOLATION; schema failures trigger semantic fallback; exhausted options trigger HYGIENE_FAILED.

Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036)

  • Decision:
    • The runtime receives the integral ECP Snapshot.
    • The local resolver loads the canonical schema file declared in runtime-config.ecp.canonical_schema_reference (src/adapters/ecp/schemas/ecp-profile.schema.json), verifies that its $id matches the $ref (https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json) in ecp-snapshot.schema.json, and registers it locally via referencing.Registry. Network HTTP retrieval is strictly disabled.
    • Invocations call the existing monorepo module src.classifier.InherenceClassifier directly through the runtime adapter. The adapter verifies classifier configuration matches ecp_classifier_config_hash.
    • Output validation verifies category, is_inherent, confidence, rationale, and textual evidences (grounded substrings in intermediate Markdown).
    • DIRECT_INHERENT / CONTEXTUAL_INHERENT → ecp_approved.
    • TANGENTIAL / NOT_RELATED → ecp_rejected (manifest persisted with status rejected_ecp, zero Markdown files generated).

Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081)

  • Decision:
    • SQLite in WAL mode (PRAGMA journal_mode=WAL;, PRAGMA busy_timeout=<configured_ms>;).
    • Native backup and restore implemented in src/storage/sqlite_store.py via sqlite3.Connection.backup.
    • Output files (.md and .result.json) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (os.replace).
    • Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to completed_text.
    • Signal Handling: src/cli/consolidate.py traps SIGTERM/SIGINT to safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit.

Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069)

  • Decision:
    • Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with LANGFUSE_BASE_URL and LANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY).
    • Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067:
      • input_validation_failure_total: reason, schema_version
      • llm_request_total: logical_call, provider, model, status
      • llm_retry_total: reason, provider, model
      • llm_fallback_total: logical_call, reason
      • llm_output_validation_failure_total: logical_call, reason, prompt_version
      • prompt_review_signal_total: dimensions defined in FR-068 (All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).
    • Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs).
    • Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table pending_telemetry and flushed via operational CLI src.cli.telemetry_flush.
    • Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (GROQ_API_KEY, DEEPSEEK_API_KEY, LANGFUSE_SECRET_KEY), without semantic text manipulation.

Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077)

  • Decision:
    • Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.