130 lines
12 KiB
Markdown
130 lines
12 KiB
Markdown
# Research & Technical Decisions: Article Consolidation and Hygiene Runtime
|
||
|
||
**Feature Branch**: `006-article-consolidation-runtime`
|
||
**Date**: 2026-08-23
|
||
**Status**: Complete
|
||
|
||
---
|
||
|
||
## 1. Executive Summary & Architectural Scope
|
||
|
||
The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot.
|
||
|
||
The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002):
|
||
- Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines).
|
||
- SQLite in WAL mode with short transactions and configurable lock timeout.
|
||
- Local filesystem for atomic staged file writes and renames.
|
||
- Agnostic Model Gateway managing two logical roles (`runtime_primary` and `runtime_fallback`) configured with certified low-cost models (defaults: Groq with `openai/gpt-oss-20b` and DeepSeek with `deepseek-v4-flash`).
|
||
- 100% LLM extractive content hygiene with strict 10-step validation harness.
|
||
- Grounded candidate selection and controlled micro-repairs without regular expressions (`re`) or manual keyword dictionaries.
|
||
- Input size limit validation (FR-056) failing in a controlled manner with `INVALID_ARTICLE_SCHEMA` before provider calls if exceeded.
|
||
- Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (`src.classifier.classify_text`).
|
||
- Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval.
|
||
- Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (`pending_telemetry`) during network outages.
|
||
|
||
---
|
||
|
||
## 2. Definitive Dependency Selections & Qualitative Impact Analysis
|
||
|
||
In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria:
|
||
|
||
| Component / Task | Chosen Solution | Standard Library Alternative | Security Impact | Maintenance Impact | License | Size Impact | Startup Impact |
|
||
|:---|:---|:---|:---|:---|:---|:---|:---|
|
||
| **HTML DOM Parsing** | `beautifulsoup4` (with `html.parser`) | `html.parser` directly | Low. Pure Python, robust against malformed HTML. | Low. Stable and mature. | MIT | Minimal | Negligible |
|
||
| **Markdown AST Parsing** | `marko` | None in stdlib for CommonMark AST | Low. Standard CommonMark AST compliance. | Low. Pure Python parser. | MIT | Minimal | Negligible |
|
||
| **JSON Schema Validation** | `jsonschema` + `referencing` | Manual validation code | Low. Standard Draft 2020-12 validator with immutable local schema registry. | Low. Reference standard. | MIT | Minimal | Negligible |
|
||
| **HTTP Client / Gateway** | `httpx` | `urllib.request` | Low. Modern HTTP client with connection pooling. | Low. High adoption. | BSD-3-Clause | Minimal | Negligible |
|
||
| **Date Parsing & ISO 8601** | `python-dateutil` | `datetime.fromisoformat` | Low. Handles diverse timezone & date formats. | Low. Industry standard. | Apache 2.0 / BSD | Minimal | Negligible |
|
||
| **YAML Serialization** | `pyyaml` | None in stdlib | Low. Safe dump (`yaml.safe_dump`) prevents code execution. | Low. Standard YAML library. | MIT | Minimal | Negligible |
|
||
| **Language Detection** | `src.language` (existing monorepo) | N/A | None. Reuses existing repository module. | Zero new dependency. | Monorepo | Zero | Zero |
|
||
| **Diffing without Regex** | `difflib.SequenceMatcher` | Stdlib `difflib` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||
| **Unicode Normalization** | `unicodedata` | Stdlib `unicodedata` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||
| **URL Parsing** | `urllib.parse` | Stdlib `urllib.parse` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||
| **State Store & Queue** | `sqlite3` | Stdlib `sqlite3` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||
| **Observability SDK** | `langfuse` | Direct HTTP calls | Low. Official SDK (>=4.7). | Low. Active upstream support. | MIT | Minimal | Negligible |
|
||
|
||
*Note on Durable Queuing*: Durable local queuing during observability outages is provided directly by the SQLite table `pending_telemetry`, not delegated to SDK memory buffers.
|
||
|
||
---
|
||
|
||
## 3. Concrete Architectural & Technical Decisions
|
||
|
||
### Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046)
|
||
- **Decision**: Orchestration is implemented as an explicit Python class `StateMachine` managing state transitions in SQLite:
|
||
- `received → validated → content_cleaned`
|
||
- `content_cleaned → ecp_approved → enriched → completed_text`
|
||
- `content_cleaned → ecp_rejected` (terminal state, zero Markdown files generated)
|
||
- Valid terminal failures → `failed`
|
||
- **Transition Recording**: Every state transition records `start_time`, `end_time`, `duration_ms`, and `result` into SQLite table `state_transitions`.
|
||
|
||
### Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044)
|
||
- **Decision**: Agnostic Model Gateway (`src.gateway.client`) supporting two certified logical roles:
|
||
- `runtime_primary`: Default configured as Groq with `openai/gpt-oss-20b` (or certified low-cost equivalent).
|
||
- `runtime_fallback`: Default configured as DeepSeek with `deepseek-v4-flash` (or certified low-cost equivalent).
|
||
- **Release Metadata Certification**: Build/packaging produces an immutable `src/core/release-metadata.json` packaged with the release containing:
|
||
- `release_version`
|
||
- `runtime_config_sha256` (SHA-256 of the approved functional config)
|
||
- `prompts_hashes` (SHA-256 of each prompt file)
|
||
- `schemas_versions` (contract version identifiers)
|
||
- `certified_models` (logical role to approved provider/model mappings)
|
||
- `ecp_classifier_config_hash` (SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters)
|
||
Preflight compares the loaded configuration against `src/core/release-metadata.json`. Any mismatch aborts with exit code `2`.
|
||
- **Retry & Fallback Policy**:
|
||
- Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response).
|
||
- Semantic failures (invalid schema, grounding violation) transition immediately to `runtime_fallback` without retrying on the same model.
|
||
- **Pricing & Operational Parameters**: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint.
|
||
|
||
### Decision 3: Input Size Limit Enforcement (FR-056)
|
||
- **Decision**: Prior to candidate extraction or remote provider calls, input size is checked against `limits.max_input_bytes`. If exceeded, execution terminates immediately with controlled error `INVALID_ARTICLE_SCHEMA` (with structured detail `"INPUT_EXCEEDS_SIZE_LIMIT"`). Strategy is strictly `fail_before_provider`. Automatic unapproved truncation is strictly prohibited.
|
||
|
||
### Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020)
|
||
- **Decision**:
|
||
- The runtime preserves all structurally valid candidate elements from all extractors.
|
||
- `difflib.SequenceMatcher` calculates cross-extractor sequence similarity during candidate preparation to populate the `equivalences` list, providing consensus evidence to the LLM without deleting or merging candidates.
|
||
- The `selected_extractor` provides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection.
|
||
- Candidate IDs are opaque, stable within execution, and carry no quality judgment.
|
||
|
||
### Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030)
|
||
- **Decision**:
|
||
- LLM returns ONLY candidate IDs and micro-repair operations in `article_content_hygiene`.
|
||
- Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content.
|
||
- Micro-repairs are validated across 5 closed categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`) using `unicodedata` and `difflib` without regex.
|
||
- Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted.
|
||
- Ungrounded candidate IDs trigger `GROUNDING_VIOLATION`; schema failures trigger semantic fallback; exhausted options trigger `HYGIENE_FAILED`.
|
||
|
||
### Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036)
|
||
- **Decision**:
|
||
- The runtime receives the integral ECP Snapshot.
|
||
- The local resolver loads the canonical schema file declared in `runtime-config.ecp.canonical_schema_reference` (`src/adapters/ecp/schemas/ecp-profile.schema.json`), verifies that its `$id` matches the `$ref` (`https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json`) in `ecp-snapshot.schema.json`, and registers it locally via `referencing.Registry`. Network HTTP retrieval is strictly disabled.
|
||
- Invocations call the existing monorepo module `src.classifier.InherenceClassifier` directly through the runtime adapter. The adapter verifies classifier configuration matches `ecp_classifier_config_hash`.
|
||
- Output validation verifies `category`, `is_inherent`, `confidence`, `rationale`, and textual `evidences` (grounded substrings in intermediate Markdown).
|
||
- `DIRECT_INHERENT` / `CONTEXTUAL_INHERENT` → `ecp_approved`.
|
||
- `TANGENTIAL` / `NOT_RELATED` → `ecp_rejected` (manifest persisted with status `rejected_ecp`, zero Markdown files generated).
|
||
|
||
### Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081)
|
||
- **Decision**:
|
||
- SQLite in WAL mode (`PRAGMA journal_mode=WAL;`, `PRAGMA busy_timeout=<configured_ms>;`).
|
||
- Native backup and restore implemented in `src/storage/sqlite_store.py` via `sqlite3.Connection.backup`.
|
||
- Output files (`.md` and `.result.json`) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (`os.replace`).
|
||
- Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to `completed_text`.
|
||
- Signal Handling: `src/cli/consolidate.py` traps `SIGTERM`/`SIGINT` to safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit.
|
||
|
||
### Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069)
|
||
- **Decision**:
|
||
- Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with `LANGFUSE_BASE_URL` and `LANGFUSE_PUBLIC_KEY`/`LANGFUSE_SECRET_KEY`).
|
||
- Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067:
|
||
- `input_validation_failure_total`: `reason`, `schema_version`
|
||
- `llm_request_total`: `logical_call`, `provider`, `model`, `status`
|
||
- `llm_retry_total`: `reason`, `provider`, `model`
|
||
- `llm_fallback_total`: `logical_call`, `reason`
|
||
- `llm_output_validation_failure_total`: `logical_call`, `reason`, `prompt_version`
|
||
- `prompt_review_signal_total`: dimensions defined in FR-068
|
||
*(All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).*
|
||
- Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs).
|
||
- Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table `pending_telemetry` and flushed via operational CLI `src.cli.telemetry_flush`.
|
||
- Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (`GROQ_API_KEY`, `DEEPSEEK_API_KEY`, `LANGFUSE_SECRET_KEY`), without semantic text manipulation.
|
||
|
||
### Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077)
|
||
- **Decision**:
|
||
- Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.
|