Files

13 KiB

Implementation Plan: Article Consolidation and Hygiene Runtime

Branch: 006-article-consolidation-runtime | Date: 2026-08-23 | Spec: specs/006-article-consolidation-runtime/spec.md

Input: Feature specification from specs/006-article-consolidation-runtime/spec.md


Summary

The Article Consolidation and Hygiene Runtime is an ephemeral, deterministic Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot. It coordinates extraction payloads from three extractors (trafilatura, newspaper4k, readability), performs 100% LLM extractive hygiene via certified low-cost models, validates grounded candidate selections and micro-repairs without regular expressions or manual keyword dictionaries, applies an ECP relevance gate via src.classifier.InherenceClassifier, adds entity-relative sentiment and native-language tags, renders canonical Markdown with YAML front matter, persists machine-readable manifests and state atomically in SQLite (WAL), and transmits sanitized observability telemetry directly to Langfuse.


Technical Context

Language/Version: Python >= 3.10 (strictly matching requires-python in pyproject.toml)
Primary Dependencies:

  • Standard library (json, urllib.parse, unicodedata, difflib, sqlite3, pathlib, hashlib, typing, signal)
  • beautifulsoup4 (DOM parsing)
  • marko (CommonMark AST parsing)
  • jsonschema + referencing (JSON Schema Draft 2020-12 validation with immutable local schema registry)
  • python-dateutil (ISO 8601 date parsing)
  • pyyaml (Safe YAML front matter serialization)
  • httpx (HTTP client for Model Gateway)
  • langfuse (Observability SDK >= 4.7)
  • Existing monorepo modules (src.language, src.classifier.InherenceClassifier)

Storage: SQLite 3 (WAL mode, configurable busy timeout, short transactions, native backup API) + Local filesystem (atomic temporary files and renames)
Testing: pytest (unit, contract, mock integration, security, load, fault injection, operations resilience), static zero-regex multi-parser analyzer (Python AST + JSON pattern check + Promptfoo YAML check), promptfoo (offline prompt evaluations in dev/CI)
Target Platform: Ambiente suportado pelo repositório e pelo deployment pipeline
Project Type: Python CLI Tool (Ephemeral process, no API, no internal worker pool)
Performance Goals: Sustained throughput >= 100 articles/hour under external orchestrator concurrency; p50, p95, and p99 latency and cost empirically measured and approved in staging before go-live
Constraints: Zero regular expressions (re) in text processing, schemas, and assertions; zero manual keyword dictionaries; zero expensive/powerful models in runtime roles (or internal ECP classifier); 11 critical release invariants (all 0); secret exposure = 0; strict 10-step hygiene harness; input size limit enforcement (INVALID_ARTICLE_SCHEMA on overflow); no generic unused abstractions
Scale/Scope: 1 article per CLI invocation, 9 versioned contracts, 10 fault injection scenarios, 8 security scenarios (SEC-001 to SEC-008), operational resilience testing (FR-081), complete golden set with holdout, contract tests over all 20 real reference units


Constitution Check

GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.

Principle / Rule Compliance Status Description
I. Simplicity Mandate (FR-001) PASS Minimum sufficient code, no generic unused abstractions, small responsibilities share modules, zero redundant local metrics system.
II. Strict Dependency Policy (FR-002) PASS Standard library prioritized; 7-factor qualitative evaluation completed; no agent frameworks, no trivial libraries, no duplicate ECP schema, no custom regex parsers.
III. CLI Ephemeral Interface (FR-003) PASS Python CLI processing 1 article per invocation, no internal batch loops, no internal worker pools, external concurrency.
IV. Regex Prohibition (FR-015, FR-070) PASS Multi-parser static check in CI enforces zero re calls in Python AST, zero pattern keys in JSON schemas, and zero regex assertions in Promptfoo YAML.
V. Agnostic Gateway & Cheap Models (FR-041, FR-042) PASS Logical roles runtime_primary and runtime_fallback limited to certified cheap models; zero powerful models in runtime or internal classifier.
VI. Zero Hallucination Grounding (FR-023, FR-026) PASS LLM returns only candidate IDs and bounded diffs; harness strictly enforces candidate grounding and reverses ungrounded micro-repairs.
VII. Atomic Persistence & Idempotency (FR-009, FR-050) PASS SQLite WAL mode + file atomic renames; deterministic fingerprinting; hash-based crash reconciliation.

Project Structure

Documentation & 9 Versioned Contracts (this feature)

specs/006-article-consolidation-runtime/
├── spec.md                       # Feature specification (v1.0.0)
├── plan.md                       # Implementation plan (this file)
├── research.md                   # Technical research & decisions (Phase 0)
├── data-model.md                 # Entity definitions, SQLite tables & formats (Phase 1)
├── quickstart.md                 # Runnable verification guide (Phase 1)
├── contracts/                    # All 9 independent versioned contracts (Phase 1)
│   ├── article-input.schema.json         # Contract 1: Article Input Unit (v1.0.0)
│   ├── ecp-snapshot.schema.json          # Contract 2: ECP Snapshot Canonical Reference (v1.0.0)
│   ├── runtime-config.schema.json        # Contract 3: Runtime Configuration (v1.0.0)
│   ├── candidates-payload.schema.json    # Contract 4: Candidate Payload to LLM (v1.0.0)
│   ├── hygiene-response.schema.json      # Contract 5: Hygiene LLM Response (v1.0.0)
│   ├── repair-operations.schema.json     # Contract 6: Repair Operations Schema (v1.0.0)
│   ├── enrichment-response.schema.json   # Contract 7: Enrichment LLM Response (v1.0.0)
│   ├── manifest-output.schema.json       # Contract 8: Output Manifest (v1.0.0)
│   ├── prompts-contract.md               # Contract 9: Versioned Prompts Contract (v1.0.0)
│   └── cli-interface.md                  # Interface: CLI Command Interface (v1.0.0)
└── checklists/
    └── requirements.md           # Quality checklist

Source Code & Fixtures (repository layout)

src/
├── __init__.py
├── cli/
│   ├── __init__.py
│   ├── consolidate.py            # Main single-article CLI entry point (--input-article, --ecp-snapshot, --config) with SIGTERM/SIGINT graceful shutdown
│   ├── preflight.py              # Preflight validation CLI (--config verified against src/core/release-metadata.json)
│   ├── smoke.py                  # Smoke test CLI (--config, --fixture)
│   ├── reconcile.py              # State & artifact reconciliation CLI (--config)
│   └── telemetry_flush.py        # Deferred telemetry flush CLI (--config)
├── core/
│   ├── __init__.py
│   ├── config.py                 # Runtime configuration loading & SHA-256 release metadata verification (exact byte hash)
│   ├── limits.py                 # Input size limit validation (FR-056)
│   ├── fingerprint.py            # Deterministic fingerprint calculator
│   ├── state_machine.py          # Python explicit state machine & SQLite logger
│   └── release-metadata.json     # Packaged release metadata (SHA-256 hashes of config, prompts, schemas, ECP config)
├── candidate/
│   ├── __init__.py
│   ├── parser.py                 # DOM, CommonMark AST, JSON-LD structural parsing
│   ├── equivalence.py            # Backbone ordering & equivalence mapping (no deletion)
│   └── models.py                 # CandidateObject definitions
├── hygiene/
│   ├── __init__.py
│   ├── harness.py                # 10-step validation harness
│   ├── repairs.py                # Controlled micro-repair validator (difflib/unicodedata)
│   └── assembler.py              # Grounded intermediate Markdown assembler
├── ecp/
│   ├── __init__.py
│   └── adapter.py                # ECP adapter invoking src.classifier.InherenceClassifier & referencing.Registry local schema
├── enrichment/
│   ├── __init__.py
│   └── harness.py                # Entity sentiment & native tags validator
├── gateway/
│   ├── __init__.py
│   ├── client.py                 # Agnostic Model Gateway (primary/fallback)
│   └── adapters.py               # Provider-specific minimal HTTP adapters (Groq, DeepSeek)
├── storage/
│   ├── __init__.py
│   ├── sqlite_store.py           # SQLite WAL state store, transition logs, and native backup/restore API
│   ├── file_store.py             # Atomic filesystem writer (temp + rename)
│   └── markdown_renderer.py      # Canonical YAML front matter & body renderer
└── observability/
    ├── __init__.py
    ├── langfuse_tracer.py        # Spans, generations, metrics, scores & secret redaction
    └── structured_logger.py      # Sanitized JSON logger

prompts/
├── article_content_hygiene.v1.txt
└── article_sentiment_tags.v1.txt

evals/
├── promptfoo.config.yaml         # Promptfoo test suite configuration
├── golden_set/                   # Full golden dataset with reference truths
├── holdout/                      # Holdout dataset (never used in few-shot)
└── reference_20/                 # 20 reference regression cases

examples/
├── sample_article_valid.json     # Executable fixture: valid inherent article unit
├── sample_article_tangential.json # Executable fixture: non-inherent article unit
└── sample_ecp_snapshot.json      # Executable fixture: canonical ECP snapshot

runtime_config.local.json         # Executable local validation configuration fixture

tests/
├── contract/                     # Contract tests for all 9 versioned contracts (evaluated over 20 real reference units)
│   ├── test_article_input_contract.py
│   ├── test_ecp_snapshot_contract.py
│   ├── test_runtime_config_contract.py
│   ├── test_candidates_payload_contract.py
│   ├── test_hygiene_response_contract.py
│   ├── test_repair_operations_contract.py
│   ├── test_enrichment_response_contract.py
│   ├── test_manifest_output_contract.py
│   └── test_prompts_contract.py
├── unit/
│   ├── test_fingerprint.py
│   ├── test_input_limits.py
│   ├── test_candidate_parser.py
│   ├── test_equivalence_mapping.py
│   ├── test_hygiene_harness.py
│   ├── test_repairs_validator.py
│   ├── test_ecp_adapter.py
│   ├── test_enrichment_harness.py
│   ├── test_model_gateway.py
│   ├── test_sqlite_store.py
│   └── test_file_store.py
├── integration/
│   ├── test_e2e_pipeline_mock.py
│   ├── test_idempotency_concurrency.py
│   └── test_operations_resilience.py    # Backup/restore, graceful shutdown, signals, pending telemetry, rollback, credential rotation, certified model rotation, and uncertified rotation rejection (FR-081)
├── security/
│   └── test_security_scenarios.py       # Parametrized tests for SEC-001 to SEC-008
├── load/
│   └── test_load_100_art_per_hour.py    # 100 articles/hour staging load benchmark
├── fault_injection/
│   └── test_fault_injection_scenarios.py # 10 normative fault scenarios (FLT-001 to FLT-010)
└── scripts/
    └── check_zero_regex.py              # Multi-parser static check (Python AST + JSON pattern + Promptfoo YAML)

Complexity Tracking

No architectural violations detected. All modules adhere strictly to simplicity and dependency rules.

Component Design Choice Simplicity Justification
Orchestration Explicit Python state machine Eliminates LangGraph / LangChain overhead while providing full deterministic lifecycle tracking.
Model Gateway Direct 2-adapter primary/fallback Eliminates smart router complexity while guaranteeing cheap model enforcement and fallback.
Candidate Model Unified CandidateObject dictionary/dataclass Avoids class hierarchies for candidate types while capturing all required structural flags.
Persistence SQLite (WAL) + Local Filesystem Eliminates distributed database/queue dependencies; provides ACID state within local process boundaries.
ECP Integration Direct invocation of src.classifier.InherenceClassifier Leverages existing monorepo module without creating redundant remote microservices.
Observability Native Langfuse dashboards + SQLite outage queue Eliminates redundant local metric systems while fulfilling all reporting requirements.
Operations Native SQLite backup API + Signal handling in CLI Fulfills FR-081 requirements using Python standard library without external daemons.
Regex Prohibition unicodedata, difflib, NLP tokenizers Multi-parser validation ensures zero regex across code, schemas, and evals without runtime overhead.