# Research & Technical Decisions: Article Consolidation and Hygiene Runtime **Feature Branch**: `006-article-consolidation-runtime` **Date**: 2026-08-23 **Status**: Complete --- ## 1. Executive Summary & Architectural Scope The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot. The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002): - Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines). - SQLite in WAL mode with short transactions and configurable lock timeout. - Local filesystem for atomic staged file writes and renames. - Agnostic Model Gateway managing two logical roles (`runtime_primary` and `runtime_fallback`) configured with certified low-cost models (defaults: Groq with `openai/gpt-oss-20b` and DeepSeek with `deepseek-v4-flash`). - 100% LLM extractive content hygiene with strict 10-step validation harness. - Grounded candidate selection and controlled micro-repairs without regular expressions (`re`) or manual keyword dictionaries. - Input size limit validation (FR-056) failing in a controlled manner with `INVALID_ARTICLE_SCHEMA` before provider calls if exceeded. - Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (`src.classifier.classify_text`). - Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval. - Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (`pending_telemetry`) during network outages. --- ## 2. Definitive Dependency Selections & Qualitative Impact Analysis In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria: | Component / Task | Chosen Solution | Standard Library Alternative | Security Impact | Maintenance Impact | License | Size Impact | Startup Impact | |:---|:---|:---|:---|:---|:---|:---|:---| | **HTML DOM Parsing** | `beautifulsoup4` (with `html.parser`) | `html.parser` directly | Low. Pure Python, robust against malformed HTML. | Low. Stable and mature. | MIT | Minimal | Negligible | | **Markdown AST Parsing** | `marko` | None in stdlib for CommonMark AST | Low. Standard CommonMark AST compliance. | Low. Pure Python parser. | MIT | Minimal | Negligible | | **JSON Schema Validation** | `jsonschema` + `referencing` | Manual validation code | Low. Standard Draft 2020-12 validator with immutable local schema registry. | Low. Reference standard. | MIT | Minimal | Negligible | | **HTTP Client / Gateway** | `httpx` | `urllib.request` | Low. Modern HTTP client with connection pooling. | Low. High adoption. | BSD-3-Clause | Minimal | Negligible | | **Date Parsing & ISO 8601** | `python-dateutil` | `datetime.fromisoformat` | Low. Handles diverse timezone & date formats. | Low. Industry standard. | Apache 2.0 / BSD | Minimal | Negligible | | **YAML Serialization** | `pyyaml` | None in stdlib | Low. Safe dump (`yaml.safe_dump`) prevents code execution. | Low. Standard YAML library. | MIT | Minimal | Negligible | | **Language Detection** | `src.language` (existing monorepo) | N/A | None. Reuses existing repository module. | Zero new dependency. | Monorepo | Zero | Zero | | **Diffing without Regex** | `difflib.SequenceMatcher` | Stdlib `difflib` | None. Standard library. | None. Standard library. | Python | Zero | Zero | | **Unicode Normalization** | `unicodedata` | Stdlib `unicodedata` | None. Standard library. | None. Standard library. | Python | Zero | Zero | | **URL Parsing** | `urllib.parse` | Stdlib `urllib.parse` | None. Standard library. | None. Standard library. | Python | Zero | Zero | | **State Store & Queue** | `sqlite3` | Stdlib `sqlite3` | None. Standard library. | None. Standard library. | Python | Zero | Zero | | **Observability SDK** | `langfuse` | Direct HTTP calls | Low. Official SDK (>=4.7). | Low. Active upstream support. | MIT | Minimal | Negligible | *Note on Durable Queuing*: Durable local queuing during observability outages is provided directly by the SQLite table `pending_telemetry`, not delegated to SDK memory buffers. --- ## 3. Concrete Architectural & Technical Decisions ### Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046) - **Decision**: Orchestration is implemented as an explicit Python class `StateMachine` managing state transitions in SQLite: - `received → validated → content_cleaned` - `content_cleaned → ecp_approved → enriched → completed_text` - `content_cleaned → ecp_rejected` (terminal state, zero Markdown files generated) - Valid terminal failures → `failed` - **Transition Recording**: Every state transition records `start_time`, `end_time`, `duration_ms`, and `result` into SQLite table `state_transitions`. ### Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044) - **Decision**: Agnostic Model Gateway (`src.gateway.client`) supporting two certified logical roles: - `runtime_primary`: Default configured as Groq with `openai/gpt-oss-20b` (or certified low-cost equivalent). - `runtime_fallback`: Default configured as DeepSeek with `deepseek-v4-flash` (or certified low-cost equivalent). - **Release Metadata Certification**: Build/packaging produces an immutable `src/core/release-metadata.json` packaged with the release containing: - `release_version` - `runtime_config_sha256` (SHA-256 of the approved functional config) - `prompts_hashes` (SHA-256 of each prompt file) - `schemas_versions` (contract version identifiers) - `certified_models` (logical role to approved provider/model mappings) - `ecp_classifier_config_hash` (SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters) Preflight compares the loaded configuration against `src/core/release-metadata.json`. Any mismatch aborts with exit code `2`. - **Retry & Fallback Policy**: - Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response). - Semantic failures (invalid schema, grounding violation) transition immediately to `runtime_fallback` without retrying on the same model. - **Pricing & Operational Parameters**: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint. ### Decision 3: Input Size Limit Enforcement (FR-056) - **Decision**: Prior to candidate extraction or remote provider calls, input size is checked against `limits.max_input_bytes`. If exceeded, execution terminates immediately with controlled error `INVALID_ARTICLE_SCHEMA` (with structured detail `"INPUT_EXCEEDS_SIZE_LIMIT"`). Strategy is strictly `fail_before_provider`. Automatic unapproved truncation is strictly prohibited. ### Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020) - **Decision**: - The runtime preserves all structurally valid candidate elements from all extractors. - `difflib.SequenceMatcher` calculates cross-extractor sequence similarity during candidate preparation to populate the `equivalences` list, providing consensus evidence to the LLM without deleting or merging candidates. - The `selected_extractor` provides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection. - Candidate IDs are opaque, stable within execution, and carry no quality judgment. ### Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030) - **Decision**: - LLM returns ONLY candidate IDs and micro-repair operations in `article_content_hygiene`. - Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content. - Micro-repairs are validated across 5 closed categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`) using `unicodedata` and `difflib` without regex. - Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted. - Ungrounded candidate IDs trigger `GROUNDING_VIOLATION`; schema failures trigger semantic fallback; exhausted options trigger `HYGIENE_FAILED`. ### Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036) - **Decision**: - The runtime receives the integral ECP Snapshot. - The local resolver loads the canonical schema file declared in `runtime-config.ecp.canonical_schema_reference` (`src/adapters/ecp/schemas/ecp-profile.schema.json`), verifies that its `$id` matches the `$ref` (`https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json`) in `ecp-snapshot.schema.json`, and registers it locally via `referencing.Registry`. Network HTTP retrieval is strictly disabled. - Invocations call the existing monorepo module `src.classifier.InherenceClassifier` directly through the runtime adapter. The adapter verifies classifier configuration matches `ecp_classifier_config_hash`. - Output validation verifies `category`, `is_inherent`, `confidence`, `rationale`, and textual `evidences` (grounded substrings in intermediate Markdown). - `DIRECT_INHERENT` / `CONTEXTUAL_INHERENT` → `ecp_approved`. - `TANGENTIAL` / `NOT_RELATED` → `ecp_rejected` (manifest persisted with status `rejected_ecp`, zero Markdown files generated). ### Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081) - **Decision**: - SQLite in WAL mode (`PRAGMA journal_mode=WAL;`, `PRAGMA busy_timeout=;`). - Native backup and restore implemented in `src/storage/sqlite_store.py` via `sqlite3.Connection.backup`. - Output files (`.md` and `.result.json`) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (`os.replace`). - Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to `completed_text`. - Signal Handling: `src/cli/consolidate.py` traps `SIGTERM`/`SIGINT` to safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit. ### Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069) - **Decision**: - Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with `LANGFUSE_BASE_URL` and `LANGFUSE_PUBLIC_KEY`/`LANGFUSE_SECRET_KEY`). - Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067: - `input_validation_failure_total`: `reason`, `schema_version` - `llm_request_total`: `logical_call`, `provider`, `model`, `status` - `llm_retry_total`: `reason`, `provider`, `model` - `llm_fallback_total`: `logical_call`, `reason` - `llm_output_validation_failure_total`: `logical_call`, `reason`, `prompt_version` - `prompt_review_signal_total`: dimensions defined in FR-068 *(All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).* - Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs). - Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table `pending_telemetry` and flushed via operational CLI `src.cli.telemetry_flush`. - Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (`GROQ_API_KEY`, `DEEPSEEK_API_KEY`, `LANGFUSE_SECRET_KEY`), without semantic text manipulation. ### Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077) - **Decision**: - Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.