Files
TextNLPClassifierApp/specs/006-article-consolidation-runtime/research.md
T

130 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Research & Technical Decisions: Article Consolidation and Hygiene Runtime
**Feature Branch**: `006-article-consolidation-runtime`
**Date**: 2026-08-23
**Status**: Complete
---
## 1. Executive Summary & Architectural Scope
The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot.
The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002):
- Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines).
- SQLite in WAL mode with short transactions and configurable lock timeout.
- Local filesystem for atomic staged file writes and renames.
- Agnostic Model Gateway managing two logical roles (`runtime_primary` and `runtime_fallback`) configured with certified low-cost models (defaults: Groq with `openai/gpt-oss-20b` and DeepSeek with `deepseek-v4-flash`).
- 100% LLM extractive content hygiene with strict 10-step validation harness.
- Grounded candidate selection and controlled micro-repairs without regular expressions (`re`) or manual keyword dictionaries.
- Input size limit validation (FR-056) failing in a controlled manner with `INVALID_ARTICLE_SCHEMA` before provider calls if exceeded.
- Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (`src.classifier.classify_text`).
- Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval.
- Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (`pending_telemetry`) during network outages.
---
## 2. Definitive Dependency Selections & Qualitative Impact Analysis
In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria:
| Component / Task | Chosen Solution | Standard Library Alternative | Security Impact | Maintenance Impact | License | Size Impact | Startup Impact |
|:---|:---|:---|:---|:---|:---|:---|:---|
| **HTML DOM Parsing** | `beautifulsoup4` (with `html.parser`) | `html.parser` directly | Low. Pure Python, robust against malformed HTML. | Low. Stable and mature. | MIT | Minimal | Negligible |
| **Markdown AST Parsing** | `marko` | None in stdlib for CommonMark AST | Low. Standard CommonMark AST compliance. | Low. Pure Python parser. | MIT | Minimal | Negligible |
| **JSON Schema Validation** | `jsonschema` + `referencing` | Manual validation code | Low. Standard Draft 2020-12 validator with immutable local schema registry. | Low. Reference standard. | MIT | Minimal | Negligible |
| **HTTP Client / Gateway** | `httpx` | `urllib.request` | Low. Modern HTTP client with connection pooling. | Low. High adoption. | BSD-3-Clause | Minimal | Negligible |
| **Date Parsing & ISO 8601** | `python-dateutil` | `datetime.fromisoformat` | Low. Handles diverse timezone & date formats. | Low. Industry standard. | Apache 2.0 / BSD | Minimal | Negligible |
| **YAML Serialization** | `pyyaml` | None in stdlib | Low. Safe dump (`yaml.safe_dump`) prevents code execution. | Low. Standard YAML library. | MIT | Minimal | Negligible |
| **Language Detection** | `src.language` (existing monorepo) | N/A | None. Reuses existing repository module. | Zero new dependency. | Monorepo | Zero | Zero |
| **Diffing without Regex** | `difflib.SequenceMatcher` | Stdlib `difflib` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
| **Unicode Normalization** | `unicodedata` | Stdlib `unicodedata` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
| **URL Parsing** | `urllib.parse` | Stdlib `urllib.parse` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
| **State Store & Queue** | `sqlite3` | Stdlib `sqlite3` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
| **Observability SDK** | `langfuse` | Direct HTTP calls | Low. Official SDK (>=4.7). | Low. Active upstream support. | MIT | Minimal | Negligible |
*Note on Durable Queuing*: Durable local queuing during observability outages is provided directly by the SQLite table `pending_telemetry`, not delegated to SDK memory buffers.
---
## 3. Concrete Architectural & Technical Decisions
### Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046)
- **Decision**: Orchestration is implemented as an explicit Python class `StateMachine` managing state transitions in SQLite:
- `received → validated → content_cleaned`
- `content_cleaned → ecp_approved → enriched → completed_text`
- `content_cleaned → ecp_rejected` (terminal state, zero Markdown files generated)
- Valid terminal failures → `failed`
- **Transition Recording**: Every state transition records `start_time`, `end_time`, `duration_ms`, and `result` into SQLite table `state_transitions`.
### Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044)
- **Decision**: Agnostic Model Gateway (`src.gateway.client`) supporting two certified logical roles:
- `runtime_primary`: Default configured as Groq with `openai/gpt-oss-20b` (or certified low-cost equivalent).
- `runtime_fallback`: Default configured as DeepSeek with `deepseek-v4-flash` (or certified low-cost equivalent).
- **Release Metadata Certification**: Build/packaging produces an immutable `src/core/release-metadata.json` packaged with the release containing:
- `release_version`
- `runtime_config_sha256` (SHA-256 of the approved functional config)
- `prompts_hashes` (SHA-256 of each prompt file)
- `schemas_versions` (contract version identifiers)
- `certified_models` (logical role to approved provider/model mappings)
- `ecp_classifier_config_hash` (SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters)
Preflight compares the loaded configuration against `src/core/release-metadata.json`. Any mismatch aborts with exit code `2`.
- **Retry & Fallback Policy**:
- Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response).
- Semantic failures (invalid schema, grounding violation) transition immediately to `runtime_fallback` without retrying on the same model.
- **Pricing & Operational Parameters**: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint.
### Decision 3: Input Size Limit Enforcement (FR-056)
- **Decision**: Prior to candidate extraction or remote provider calls, input size is checked against `limits.max_input_bytes`. If exceeded, execution terminates immediately with controlled error `INVALID_ARTICLE_SCHEMA` (with structured detail `"INPUT_EXCEEDS_SIZE_LIMIT"`). Strategy is strictly `fail_before_provider`. Automatic unapproved truncation is strictly prohibited.
### Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020)
- **Decision**:
- The runtime preserves all structurally valid candidate elements from all extractors.
- `difflib.SequenceMatcher` calculates cross-extractor sequence similarity during candidate preparation to populate the `equivalences` list, providing consensus evidence to the LLM without deleting or merging candidates.
- The `selected_extractor` provides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection.
- Candidate IDs are opaque, stable within execution, and carry no quality judgment.
### Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030)
- **Decision**:
- LLM returns ONLY candidate IDs and micro-repair operations in `article_content_hygiene`.
- Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content.
- Micro-repairs are validated across 5 closed categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`) using `unicodedata` and `difflib` without regex.
- Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted.
- Ungrounded candidate IDs trigger `GROUNDING_VIOLATION`; schema failures trigger semantic fallback; exhausted options trigger `HYGIENE_FAILED`.
### Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036)
- **Decision**:
- The runtime receives the integral ECP Snapshot.
- The local resolver loads the canonical schema file declared in `runtime-config.ecp.canonical_schema_reference` (`src/adapters/ecp/schemas/ecp-profile.schema.json`), verifies that its `$id` matches the `$ref` (`https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json`) in `ecp-snapshot.schema.json`, and registers it locally via `referencing.Registry`. Network HTTP retrieval is strictly disabled.
- Invocations call the existing monorepo module `src.classifier.InherenceClassifier` directly through the runtime adapter. The adapter verifies classifier configuration matches `ecp_classifier_config_hash`.
- Output validation verifies `category`, `is_inherent`, `confidence`, `rationale`, and textual `evidences` (grounded substrings in intermediate Markdown).
- `DIRECT_INHERENT` / `CONTEXTUAL_INHERENT` → `ecp_approved`.
- `TANGENTIAL` / `NOT_RELATED` → `ecp_rejected` (manifest persisted with status `rejected_ecp`, zero Markdown files generated).
### Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081)
- **Decision**:
- SQLite in WAL mode (`PRAGMA journal_mode=WAL;`, `PRAGMA busy_timeout=<configured_ms>;`).
- Native backup and restore implemented in `src/storage/sqlite_store.py` via `sqlite3.Connection.backup`.
- Output files (`.md` and `.result.json`) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (`os.replace`).
- Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to `completed_text`.
- Signal Handling: `src/cli/consolidate.py` traps `SIGTERM`/`SIGINT` to safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit.
### Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069)
- **Decision**:
- Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with `LANGFUSE_BASE_URL` and `LANGFUSE_PUBLIC_KEY`/`LANGFUSE_SECRET_KEY`).
- Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067:
- `input_validation_failure_total`: `reason`, `schema_version`
- `llm_request_total`: `logical_call`, `provider`, `model`, `status`
- `llm_retry_total`: `reason`, `provider`, `model`
- `llm_fallback_total`: `logical_call`, `reason`
- `llm_output_validation_failure_total`: `logical_call`, `reason`, `prompt_version`
- `prompt_review_signal_total`: dimensions defined in FR-068
*(All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).*
- Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs).
- Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table `pending_telemetry` and flushed via operational CLI `src.cli.telemetry_flush`.
- Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (`GROQ_API_KEY`, `DEEPSEEK_API_KEY`, `LANGFUSE_SECRET_KEY`), without semantic text manipulation.
### Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077)
- **Decision**:
- Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.