12 KiB
12 KiB
Research & Technical Decisions: Article Consolidation and Hygiene Runtime
Feature Branch: 006-article-consolidation-runtime
Date: 2026-08-23
Status: Complete
1. Executive Summary & Architectural Scope
The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot.
The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002):
- Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines).
- SQLite in WAL mode with short transactions and configurable lock timeout.
- Local filesystem for atomic staged file writes and renames.
- Agnostic Model Gateway managing two logical roles (
runtime_primaryandruntime_fallback) configured with certified low-cost models (defaults: Groq withopenai/gpt-oss-20band DeepSeek withdeepseek-v4-flash). - 100% LLM extractive content hygiene with strict 10-step validation harness.
- Grounded candidate selection and controlled micro-repairs without regular expressions (
re) or manual keyword dictionaries. - Input size limit validation (FR-056) failing in a controlled manner with
INVALID_ARTICLE_SCHEMAbefore provider calls if exceeded. - Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (
src.classifier.classify_text). - Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval.
- Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (
pending_telemetry) during network outages.
2. Definitive Dependency Selections & Qualitative Impact Analysis
In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria:
| Component / Task | Chosen Solution | Standard Library Alternative | Security Impact | Maintenance Impact | License | Size Impact | Startup Impact |
|---|---|---|---|---|---|---|---|
| HTML DOM Parsing | beautifulsoup4 (with html.parser) |
html.parser directly |
Low. Pure Python, robust against malformed HTML. | Low. Stable and mature. | MIT | Minimal | Negligible |
| Markdown AST Parsing | marko |
None in stdlib for CommonMark AST | Low. Standard CommonMark AST compliance. | Low. Pure Python parser. | MIT | Minimal | Negligible |
| JSON Schema Validation | jsonschema + referencing |
Manual validation code | Low. Standard Draft 2020-12 validator with immutable local schema registry. | Low. Reference standard. | MIT | Minimal | Negligible |
| HTTP Client / Gateway | httpx |
urllib.request |
Low. Modern HTTP client with connection pooling. | Low. High adoption. | BSD-3-Clause | Minimal | Negligible |
| Date Parsing & ISO 8601 | python-dateutil |
datetime.fromisoformat |
Low. Handles diverse timezone & date formats. | Low. Industry standard. | Apache 2.0 / BSD | Minimal | Negligible |
| YAML Serialization | pyyaml |
None in stdlib | Low. Safe dump (yaml.safe_dump) prevents code execution. |
Low. Standard YAML library. | MIT | Minimal | Negligible |
| Language Detection | src.language (existing monorepo) |
N/A | None. Reuses existing repository module. | Zero new dependency. | Monorepo | Zero | Zero |
| Diffing without Regex | difflib.SequenceMatcher |
Stdlib difflib |
None. Standard library. | None. Standard library. | Python | Zero | Zero |
| Unicode Normalization | unicodedata |
Stdlib unicodedata |
None. Standard library. | None. Standard library. | Python | Zero | Zero |
| URL Parsing | urllib.parse |
Stdlib urllib.parse |
None. Standard library. | None. Standard library. | Python | Zero | Zero |
| State Store & Queue | sqlite3 |
Stdlib sqlite3 |
None. Standard library. | None. Standard library. | Python | Zero | Zero |
| Observability SDK | langfuse |
Direct HTTP calls | Low. Official SDK (>=4.7). | Low. Active upstream support. | MIT | Minimal | Negligible |
Note on Durable Queuing: Durable local queuing during observability outages is provided directly by the SQLite table pending_telemetry, not delegated to SDK memory buffers.
3. Concrete Architectural & Technical Decisions
Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046)
- Decision: Orchestration is implemented as an explicit Python class
StateMachinemanaging state transitions in SQLite:received → validated → content_cleanedcontent_cleaned → ecp_approved → enriched → completed_textcontent_cleaned → ecp_rejected(terminal state, zero Markdown files generated)- Valid terminal failures →
failed
- Transition Recording: Every state transition records
start_time,end_time,duration_ms, andresultinto SQLite tablestate_transitions.
Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044)
- Decision: Agnostic Model Gateway (
src.gateway.client) supporting two certified logical roles:runtime_primary: Default configured as Groq withopenai/gpt-oss-20b(or certified low-cost equivalent).runtime_fallback: Default configured as DeepSeek withdeepseek-v4-flash(or certified low-cost equivalent).
- Release Metadata Certification: Build/packaging produces an immutable
src/core/release-metadata.jsonpackaged with the release containing:release_versionruntime_config_sha256(SHA-256 of the approved functional config)prompts_hashes(SHA-256 of each prompt file)schemas_versions(contract version identifiers)certified_models(logical role to approved provider/model mappings)ecp_classifier_config_hash(SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters) Preflight compares the loaded configuration againstsrc/core/release-metadata.json. Any mismatch aborts with exit code2.
- Retry & Fallback Policy:
- Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response).
- Semantic failures (invalid schema, grounding violation) transition immediately to
runtime_fallbackwithout retrying on the same model.
- Pricing & Operational Parameters: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint.
Decision 3: Input Size Limit Enforcement (FR-056)
- Decision: Prior to candidate extraction or remote provider calls, input size is checked against
limits.max_input_bytes. If exceeded, execution terminates immediately with controlled errorINVALID_ARTICLE_SCHEMA(with structured detail"INPUT_EXCEEDS_SIZE_LIMIT"). Strategy is strictlyfail_before_provider. Automatic unapproved truncation is strictly prohibited.
Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020)
- Decision:
- The runtime preserves all structurally valid candidate elements from all extractors.
difflib.SequenceMatchercalculates cross-extractor sequence similarity during candidate preparation to populate theequivalenceslist, providing consensus evidence to the LLM without deleting or merging candidates.- The
selected_extractorprovides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection. - Candidate IDs are opaque, stable within execution, and carry no quality judgment.
Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030)
- Decision:
- LLM returns ONLY candidate IDs and micro-repair operations in
article_content_hygiene. - Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content.
- Micro-repairs are validated across 5 closed categories (
encoding,unicode,spacing,punctuation_corruption,obvious_typo) usingunicodedataanddifflibwithout regex. - Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted.
- Ungrounded candidate IDs trigger
GROUNDING_VIOLATION; schema failures trigger semantic fallback; exhausted options triggerHYGIENE_FAILED.
- LLM returns ONLY candidate IDs and micro-repair operations in
Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036)
- Decision:
- The runtime receives the integral ECP Snapshot.
- The local resolver loads the canonical schema file declared in
runtime-config.ecp.canonical_schema_reference(src/adapters/ecp/schemas/ecp-profile.schema.json), verifies that its$idmatches the$ref(https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json) inecp-snapshot.schema.json, and registers it locally viareferencing.Registry. Network HTTP retrieval is strictly disabled. - Invocations call the existing monorepo module
src.classifier.InherenceClassifierdirectly through the runtime adapter. The adapter verifies classifier configuration matchesecp_classifier_config_hash. - Output validation verifies
category,is_inherent,confidence,rationale, and textualevidences(grounded substrings in intermediate Markdown). DIRECT_INHERENT/CONTEXTUAL_INHERENT→ecp_approved.TANGENTIAL/NOT_RELATED→ecp_rejected(manifest persisted with statusrejected_ecp, zero Markdown files generated).
Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081)
- Decision:
- SQLite in WAL mode (
PRAGMA journal_mode=WAL;,PRAGMA busy_timeout=<configured_ms>;). - Native backup and restore implemented in
src/storage/sqlite_store.pyviasqlite3.Connection.backup. - Output files (
.mdand.result.json) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (os.replace). - Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to
completed_text. - Signal Handling:
src/cli/consolidate.pytrapsSIGTERM/SIGINTto safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit.
- SQLite in WAL mode (
Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069)
- Decision:
- Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with
LANGFUSE_BASE_URLandLANGFUSE_PUBLIC_KEY/LANGFUSE_SECRET_KEY). - Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067:
input_validation_failure_total:reason,schema_versionllm_request_total:logical_call,provider,model,statusllm_retry_total:reason,provider,modelllm_fallback_total:logical_call,reasonllm_output_validation_failure_total:logical_call,reason,prompt_versionprompt_review_signal_total: dimensions defined in FR-068 (All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).
- Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs).
- Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table
pending_telemetryand flushed via operational CLIsrc.cli.telemetry_flush. - Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (
GROQ_API_KEY,DEEPSEEK_API_KEY,LANGFUSE_SECRET_KEY), without semantic text manipulation.
- Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with
Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077)
- Decision:
- Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.