# Implementation Tasks: Article Consolidation and Hygiene Runtime **Feature**: Article Consolidation and Hygiene Runtime (`specs/006-article-consolidation-runtime/spec.md`) **Branch**: `006-article-consolidation-runtime` | **Date**: 2026-08-23 | **Plan**: [`plan.md`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/plan.md) **Status**: Ready for Execution --- ## Phase 1: Setup (Shared Infrastructure & Tooling) **Purpose**: Project initialization, dependency management, and quality verification tooling. - [x] T001 Initialize the package structure and lockfile using only dependencies approved by the implementation plan in `pyproject.toml`, recording for every new dependency: requirement served, standard-library alternative, security impact, maintenance impact, license, size impact, and startup impact - [x] T002 [P] Implement multi-parser static policy verification script in `tests/scripts/check_zero_regex.py`: checking Python AST for imports and direct calls of `re` or any regular-expression engine/API in the scoped text-processing modules, including aliases, without inspecting internals of transitive dependencies, rejecting `pattern` keys in JSON schemas via JSON parser, and validating Promptfoo YAML configurations via YAML parser (failing on regex assertions, semantic `contains`/`not-contains` assertions, LLM-as-a-judge for grounding, powerful models as judge, and configurations relying solely on global averages without per-case and per-slice gates) - [x] T003 [P] Configure Promptfoo test environment and suite settings in `evals/promptfoo.config.yaml` strictly following policy constraints (no `contains`/`not-contains` semantic decisions, no LLM-as-a-judge for grounding, no powerful models as judge, and per-case and per-slice assertion gates) - [x] T004 [P] Create initial 20-case reference regression dataset in `evals/reference_20/`, converting each of the 20 reference articles into an individual unit file associated with a valid, versioned canonical ECP snapshot per Test Plan §4.1 - [x] T005 [P] Create local validation configuration fixture in `runtime_config.local.json` --- ## Phase 2: Foundational (Blocking Prerequisites & Shared Core) **Purpose**: Core infrastructure, base models, SQLite WAL store, atomic file writer, manifest generator, and configuration engine that MUST be complete before pipeline execution. > **CRITICAL**: No user story implementation can begin until this foundational phase is complete. - [x] T006 Implement configuration loading, validation, and exact-byte SHA-256 hash verification in `src/core/config.py` - [x] T007 [P] Implement input byte size limiter with fail-before-provider policy in `src/core/limits.py` - [x] T008 [P] Implement deterministic canonical SHA-256 execution fingerprint calculator in `src/core/fingerprint.py` - [x] T009 [P] Implement `CandidateObject` and text repair dataclasses in `src/candidate/models.py` - [x] T010 Implement SQLite WAL store in `src/storage/sqlite_store.py` with short transactions, busy timeout, native backup/restore API, and atomic fingerprint claim logic (checking completed fingerprint before remote calls, returning existing result, and safely resuming/reusing concurrent executions) - [x] T011 Implement Python explicit state machine and SQLite transition logger in `src/core/state_machine.py` - [x] T012 [P] Implement structured JSON logging with all normative fields, structural authorization-header sanitization, and exact replacement of known environment-secret values in `src/observability/structured_logger.py`, strictly omitting full ECP, full HTML, and full article text - [x] T013 [P] Implement the atomic filesystem writer and shared manifest generator in `src/storage/file_store.py`, writing temporary files in the same destination filesystem, flushing, closing, verifying exact SHA-256 hashes, and performing atomic rename (`os.replace`) without cross-filesystem moves, complying with `manifest-output.schema.json` and the 16 normative error codes (shared across `completed_text`, `rejected_ecp`, `failed_validation`, and `failed_processing`) - [x] T014 [P] Implement contract tests for runtime configuration in `tests/contract/test_runtime_config_contract.py` - [x] T015 [P] Implement contract tests for manifest output schema in `tests/contract/test_manifest_output_contract.py` **Checkpoint**: Foundation ready — Model Gateway and Pipeline components can now proceed. --- ## Phase 3: User Story 6 - Model Gateway Infrastructure & Cheap Model Enforcement **Goal**: Agnostic Model Gateway, provider adapters, technical retries, and preflight cheap model enforcement (MUST exist before any LLM hygiene or enrichment call). ### Tests for Model Gateway - [x] T016 [P] [US6] Implement unit tests for Model Gateway client and adapters in `tests/unit/test_model_gateway.py` covering normative scenarios `LLM-001` to `LLM-012` (logical roles, pricing, token tracking, timeouts) - [x] T017 [P] [US6] Implement fault injection tests for gateway transient errors and failovers in `tests/fault_injection/test_gateway_faults.py` covering provider fault scenarios ### Implementation for Model Gateway - [x] T018 [US6] Implement agnostic Model Gateway client managing logical roles (`runtime_primary`, `runtime_fallback`) and token pricing calculations in `src/gateway/client.py` - [x] T019 [US6] Implement minimal HTTP adapters for Groq and DeepSeek using `httpx` in `src/gateway/adapters.py` - [x] T020 [US6] Implement limited technical retries for timeout, connection interruption/reset, HTTP 429 with configured backoff up to limit, HTTP 5xx, and empty technical responses in `src/gateway/client.py` - [x] T021 [US6] Implement immediate semantic fallback from `runtime_primary` to `runtime_fallback` on schema or grounding failure without retrying on the same model in `src/gateway/client.py` - [x] T022 [US6] Implement a single, reusable certified-configuration validation in `src/core/config.py` rejecting any uncertified or powerful models across all runtime roles (including any internal ECP LLM) **Checkpoint**: Model Gateway implementation and configuration enforcement are ready for pipeline integration; production certification occurs only after Phase 12 gates. --- ## Phase 4: User Story 1 - Single Article Ingestion, Contract Validation, and Candidate Preparation (Priority: P1) **Goal**: Ingest single article units, validate contracts (Article and ECP) locally before any remote call, compute deterministic fingerprint, reject batch wrappers, and extract structured candidates without regular expressions. **Independent Test**: Provide single article JSON objects (valid, corrupt, batch wrapper) and ECP snapshots, verifying schema validation, SQLite state initialization (`received`, `validated`), deterministic candidate ID generation, and immediate pre-remote termination with exact error codes. ### Tests for User Story 1 - [x] T023 [P] [US1] Implement contract test for Article Input schema against all 20 real reference units in `tests/contract/test_article_input_contract.py` - [x] T024 [P] [US1] Implement contract test for ECP Snapshot schema and local `referencing.Registry` resolution in `tests/contract/test_ecp_snapshot_contract.py` - [x] T025 [P] [US1] Implement contract test for Candidates Payload schema in `tests/contract/test_candidates_payload_contract.py` - [x] T026 [P] [US1] Implement unit tests for input limits, validation, and error code mapping in `tests/unit/test_input_limits.py` covering scenarios `IN-001` to `IN-015` and proving zero remote provider, remote Langfuse, or classifier calls on local failure - [x] T027 [P] [US1] Implement unit tests for deterministic fingerprint calculation and idempotency claims in `tests/unit/test_fingerprint.py` covering scenarios `ID-001` to `ID-010` - [x] T028 [P] [US1] Implement unit tests for candidate extraction without regex in `tests/unit/test_candidate_parser.py` covering scenarios `PAR-001` to `PAR-010` - [x] T029 [P] [US1] Implement unit tests for cross-extractor sequence equivalence mapping in `tests/unit/test_equivalence_mapping.py` covering scenarios `CAN-001` to `CAN-010` ### Implementation for User Story 1 - [x] T030 [P] [US1] Create executable article and ECP fixtures in `examples/sample_article_valid.json`, `examples/sample_article_tangential.json`, and `examples/sample_ecp_snapshot.json` - [x] T031 [US1] Implement local canonical ECP schema resolution and registration via `referencing.Registry` (disabling HTTP network fetching) in `src/ecp/adapter.py` - [x] T032 [US1] Implement local pre-call input validation and batch wrapper rejection (`"articles": false`) in `src/core/config.py`, preserving unknown fields in the recorded original input while ignoring them during processing - [x] T033 [US1] Implement structural candidate parsing in `src/candidate/parser.py` using DOM for HTML, CommonMark AST for Markdown, JSON parsing for JSON-LD, URL parsing, Unicode normalization, and an appropriate multilingual tokenizer/segmenter and language detector, handling malformed HTML safely and invalid JSON-LD through a controlled warning - [x] T034 [US1] Implement non-destructive candidate equivalence mapping using `difflib.SequenceMatcher` in `src/candidate/equivalence.py` - [x] T035 [US1] Implement deterministic source URL and publication date resolution using the exact normative priorities and date-consensus rule, and prepare title, subtitle, and author candidates using their normative source priorities in `src/candidate/parser.py`, omitting invalid dates and strictly forbidding delimiter-based author splitting - [x] T036 [US1] Connect initial validation to state machine `received → validated` in `src/core/state_machine.py`, ensuring transition only occurs after article, ECP, config, `selected_extractor`, size limit, and minimum content checks pass - [x] T037 [US1] Implement CLI ingestion entrypoint in `src/cli/consolidate.py` with full idempotency checks (querying fingerprint before LLM, returning existing result on match, resuming incomplete runs, claiming atomic execution), emitting structured JSON, complete manifest on stdout when fingerprint exists, technical envelope on unparseable JSON, persisting manifest, and enforcing exact exit codes (`0`: completed/rejected_ecp, `1`: invalid article/ECP, `2`: config/preflight error, `3`: failed processing, `4`: persistence failure) **Checkpoint**: User Story 1 is independently functional, validating contracts and preparing candidates locally. --- ## Phase 5: User Story 2 - Mandatory LLM Extractive Hygiene & Controlled Text Repairs (Priority: P1) **Goal**: Execute 100% LLM extractive hygiene over candidate payloads, enforce the 10-step validation harness, permit only 5 closed micro-repair categories, and reject ungrounded edits without regex. **Independent Test**: Feed candidate payloads with consensus, divergence, and noise into the hygiene harness, verifying that the LLM returns only candidate IDs and repairs, ungrounded IDs trigger `GROUNDING_VIOLATION`, invalid repairs are discarded with originals preserved, and valid intermediate Markdown is assembled. ### Tests for User Story 2 - [x] T038 [P] [US2] Implement contract test for Hygiene Response schema in `tests/contract/test_hygiene_response_contract.py` - [x] T039 [P] [US2] Implement contract test for Repair Operations schema in `tests/contract/test_repair_operations_contract.py` - [x] T040 [P] [US2] Implement unit tests for 10-step hygiene validation harness in `tests/unit/test_hygiene_harness.py` covering scenarios `HYG-001` to `HYG-021` (testing context exclusions, grounding enforcement, and candidate ID validation) - [x] T041 [P] [US2] Implement unit tests for controlled text repairs without regex in `tests/unit/test_repairs_validator.py` covering scenarios `REP-001` to `REP-016` (5 closed categories, sensitive entity protection, exact fragment targeting, and Unicode/NLP-based diff validation without uncalibrated numeric thresholds) - [x] T042 [P] [US2] Implement Promptfoo evaluation suite for `article_content_hygiene` prompt in `evals/promptfoo.config.yaml` validating context exclusions (no raw JSON, no full HTML, no logs, no secrets, no self-healing, no semantic `contains` assertions) ### Implementation for User Story 2 - [x] T043 [P] [US2] Author normative versioned prompt in `prompts/article_content_hygiene.v1.txt` following the exact 6-block ordering (Doc 07 §5.3) - [x] T044 [US2] Implement the minimal candidate/context projection builder in `src/hygiene/harness.py`, excluding full raw JSON, full HTML, other-article data, logs, secrets, full ECP when minimal identity is sufficient, rejected prior responses except required technical fallback metadata, self-healing instructions, and language-specific semantic keyword examples, while delimiting article content strictly as untrusted data - [x] T045 [US2] Implement the micro-repair validator in `src/hygiene/repairs.py` enforcing the 5 closed categories, exact-fragment targeting, Unicode/NLP-based comparison, sensitive-entity preservation, and audit decisions without regex (no quantitative similarity threshold unless one is later approved through the golden-set evaluation) - [x] T046 [US2] Implement 10-step hygiene harness in `src/hygiene/harness.py` validating candidate IDs, ordering, links/images, minimum content, and grounding - [x] T047 [US2] Implement grounded intermediate Markdown assembler in `src/hygiene/assembler.py` - [x] T048 [US2] Implement decoupling between schema failures (semantic fallback) and grounding violations (immediate invalidation) in `src/hygiene/harness.py` - [x] T049 [US2] Implement the conservative deterministic hygiene fallback in `src/hygiene/harness.py` using only the `selected_extractor` structural backbone, removing only structurally invalid elements, without regex, keyword dictionaries, semantic advertisement filtering, or content-quality inference; use it only when grounding and minimum-content requirements are satisfied, otherwise terminate with `HYGIENE_FAILED` - [x] T050 [US2] Connect hygiene stage to state machine `validated → content_cleaned` in `src/core/state_machine.py` - [x] T051 [US2] Integrate hygiene harness execution and error handling into `src/cli/consolidate.py` **Checkpoint**: User Stories 1 and 2 operate together, performing grounded extractive hygiene and controlled repairs. --- ## Phase 6: User Story 3 - Mandatory ECP Gate & Relevance Enforcement (Priority: P1) **Goal**: Evaluate intermediate sanitized Markdown against canonical ECP Snapshot using `src.classifier.InherenceClassifier`, transitioning inherent articles to `ecp_approved` and non-inherent articles to `ecp_rejected` with zero Markdown generated. **Independent Test**: Submit intermediate Markdown to ECP adapter with profiles across all 4 categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`), verifying that only inherent articles proceed to enrichment, while non-inherent articles persist `.result.json` with status `rejected_ecp` and produce no `.md` file. ### Tests for User Story 3 - [x] T052 [P] [US3] Implement unit tests for ECP adapter invoking `InherenceClassifier` in `tests/unit/test_ecp_adapter.py` covering scenarios `ECP-001` to `ECP-009` (full output validation, grounded evidence check, tier tracking, cheap model enforcement) - [x] T053 [P] [US3] Implement integration test for ECP rejection producing zero Markdown files in `tests/integration/test_ecp_rejection_flow.py` covering scenario `OUT-008` ### Implementation for User Story 3 - [x] T054 [US3] Implement the ECP classification adapter in `src/ecp/adapter.py` invoking `src.classifier.InherenceClassifier` through its public contract and certified configuration, validating `category`, `is_inherent`, `confidence`, `rationale`, and `evidences`, asserting that all evidence fragments belong to the intermediate Markdown, and recording any classifier tier or LLM generation exposed by the classifier - [x] T055 [US3] Make the ECP adapter consume the shared certified-configuration validation from `src/core/config.py`, verifying the ECP classifier configuration against packaged release metadata without duplicating hash or certification logic - [x] T056 [US3] Connect ECP inherence gate to state machine `content_cleaned → ecp_approved | ecp_rejected` in `src/core/state_machine.py` - [x] T057 [US3] Implement `rejected_ecp` terminal flow writing manifest with status `rejected_ecp` (`ECP_REJECTED`) and strictly omitting Markdown output in `src/storage/file_store.py` - [x] T058 [US3] Integrate ECP gate execution and error handling into `src/cli/consolidate.py` **Checkpoint**: Core pipeline evaluates inherence and enforces the strict ECP publishing gate. --- ## Phase 7: User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2) **Goal**: Enrich ECP-approved articles with entity-relative sentiment and 3 to 8 native language tags supported by textual evidence IDs, strictly decoupled from body text. **Independent Test**: Submit approved intermediate Markdown and minimal ECP identity (`qid`, `canonical_name`) to enrichment harness, verifying sentiment extraction, tag bounding (3–8), NLP uniqueness without regex, evidence grounding, and failure handling without body modification. ### Tests for User Story 4 - [x] T059 [P] [US4] Implement contract test for Enrichment Response schema in `tests/contract/test_enrichment_response_contract.py` - [x] T060 [P] [US4] Implement unit tests for entity sentiment and native tags validator in `tests/unit/test_enrichment_harness.py` covering scenarios `ENR-001` to `ENR-009` (sentiment relative to entity, tag bounding, evidence IDs, context exclusions) - [x] T061 [P] [US4] Implement Promptfoo evaluation suite for `article_sentiment_tags` prompt in `evals/promptfoo.config.yaml` verifying minimal ECP identity context and prohibiting semantic `contains` assertions - [x] T062 [P] [US4] Implement contract tests for both versioned prompts in `tests/contract/test_prompts_contract.py` verifying 6-block sequence, semver parsing without regex, SHA-256 calculation, and Promptfoo parity ### Implementation for User Story 4 - [x] T063 [P] [US4] Author normative versioned prompt in `prompts/article_sentiment_tags.v1.txt` following the exact 6-block ordering (Doc 07 §5.3) - [x] T064 [US4] Implement enrichment harness in `src/enrichment/harness.py` validating sentiment enum, 3–8 unique tags via NLP/Unicode, and evidence candidate IDs, strictly limiting ECP context to `qid` and `canonical_name` (no raw ECP snapshot, no keyword lists) - [x] T065 [US4] Connect enrichment stage to state machine `ecp_approved → enriched` in `src/core/state_machine.py` - [x] T066 [US4] Implement fallback routing and terminal `ENRICHMENT_FAILED` handling (blocking Markdown generation on failure) in `src/enrichment/harness.py` - [x] T067 [US4] Integrate enrichment stage execution and error handling into `src/cli/consolidate.py` **Checkpoint**: User Story 4 delivers structured sentiment and native tags metadata for inherent articles. --- ## Phase 8: User Story 5 - Canonical Markdown Rendering & Atomic Persistence (Priority: P2) **Goal**: Render canonical Markdown with YAML front matter, generate machine-readable `.result.json` manifests, execute atomic filesystem writes (temp + rename), and maintain strict SQLite state consistency. **Independent Test**: Verify generated `.md` and `.result.json` files, validating YAML front matter structure, 64-character SHA-256 content hashes, atomic rename lifecycle, and hash-based reconciliation of interrupted writes. ### Tests for User Story 5 - [x] T068 [P] [US5] Implement unit tests for canonical YAML front matter and Markdown body renderer in `tests/unit/test_markdown_renderer.py` covering scenarios `OUT-002` to `OUT-007`, `OUT-009`, and `OUT-010` (grounding, formatting, front matter structure) - [x] T069 [P] [US5] Implement unit tests for atomic file writes, permissions, and 64-character hash verification in `tests/unit/test_file_store.py` covering scenarios `OUT-001`, `OUT-011`, and `OUT-012` - [x] T070 [P] [US5] Implement unit tests for SQLite WAL state persistence and crash reconciliation in `tests/unit/test_sqlite_store.py` covering crash recovery and multi-terminal state consistency ### Implementation for User Story 5 - [x] T071 [US5] Implement canonical YAML front matter and Markdown body renderer in `src/storage/markdown_renderer.py` (H1 title, italic subtitle when present with no extra blank line when absent, canonical body order, grounded links/images, omitting author/date/sentiment/tags/ECP from body) - [x] T072 [US5] Integrate the atomic filesystem writer from `src/storage/file_store.py` with SQLite completion state in `src/storage/sqlite_store.py` within the same logical completion unit (without distributed transactions), resolving crash divergence via hash-based reconciliation for ID-009, ensuring no terminal state is exposed as completed while files and hashes disagree, and preserving the last safe state without exposing partial final artifact pairs on `PERSISTENCE_FAILED` (covering `completed_text`, `rejected_ecp`, `failed_validation`, and `failed_processing`) - [x] T073 [US5] Connect final persistence to state machine `enriched → completed_text` in `src/core/state_machine.py` - [x] T074 [US5] Integrate final persistence and reconciliation into `src/cli/consolidate.py` and `src/cli/reconcile.py` **Checkpoint**: End-to-end pipeline produces atomic published Markdown and manifests with SQLite consistency. --- ## Phase 9: User Story 7 - Direct Langfuse Observability, Log Sanitization & Telemetry Queue (Priority: P3) **Goal**: Transmit traces, spans, generations, and metrics directly to Langfuse, enforce secret redaction, degrade gracefully to SQLite `pending_telemetry` on network outage, and provide operational telemetry flush. **Independent Test**: Process articles with Langfuse available and blocked, checking trace structure (8 stable spans), secret redaction in stderr logs, SQLite queue insertion on outage, and flush execution via `telemetry_flush` CLI. ### Tests for User Story 7 - [x] T075 [P] [US7] Implement unit tests for Langfuse tracer, secret redaction, and offline queue in `tests/unit/test_langfuse_tracer.py` covering scenarios `OBS-001` to `OBS-011` (8 stable spans, generation attributes, metric dimensions, cardinality guards) - [x] T076 [P] [US7] Implement integration tests for telemetry degradation, deduplication, and atomic replay in `tests/integration/test_telemetry_degradation.py` - [x] T077 [P] [US7] Implement specialized security tests for authorization-header and secret redaction in logs and SDK exceptions (SEC-006) in `tests/security/test_secret_redaction.py` ### Implementation for User Story 7 - [x] T078 [US7] Implement Langfuse observability integration in `src/observability/langfuse_tracer.py` managing traces with 8 stable spans (`validation`, `candidate_preparation`, `hygiene`, `grounding_validation`, `ecp_gate`, `enrichment`, `rendering`, `persistence`) and generations per LLM attempt; the validation span MUST be buffered or materialized only after successful local validation (no remote Langfuse traffic may occur while terminating local validations are running) - [x] T079 [US7] Implement local SQLite queue insertion for telemetry events during Langfuse network outages and atomic flush procedure in `src/observability/langfuse_tracer.py` - [x] T080 [US7] Implement runtime-observable metric emission in `src/observability/langfuse_tracer.py` using exactly each metric and its dimensions from Doc 05 / FR-067, enforcing cardinality restrictions and emitting `prompt_review_signal_total` without self-healing or a parallel metrics store (release-wide aggregation of the 11 critical invariants remains the responsibility of T090) - [x] T081 [US7] Implement operational telemetry flush command in `src/cli/telemetry_flush.py`, ensuring events are marked flushed only upon confirmed delivery and `telemetry_pending_total` returns to zero - [x] T082 [US7] Configure the 3 mandatory Langfuse dashboards (Runtime Health, Quality, Future Review Signals) and preserve reproducible setup evidence without creating a parallel metrics system - [x] T083 [US7] Integrate observability lifecycle and secret redaction into `src/cli/consolidate.py` **Checkpoint**: Observability is complete, compliant with the metric catalog, and resilient against outages. --- ## Phase 10: User Story 8 - Quality Gates, Zero Regex Verification & Promptfoo Evaluation (Priority: P3) **Goal**: Execute comprehensive automated quality gates, Promptfoo offline evaluations, multi-extractor golden regression across 20 reference cases, and zero-regex verification. **Independent Test**: Run `pytest tests/quality/`, `pytest evals/`, and `tests/scripts/check_zero_regex.py`, asserting that all 11 critical quality assertions pass without manual intervention. ### Tests for User Story 8 - [x] T084 [P] [US8] Implement automated test in `tests/quality/test_zero_regex_enforcement.py` executing `tests/scripts/check_zero_regex.py` across codebase, schemas, and Promptfoo YAML - [x] T085 [P] [US8] Implement automated test in `tests/quality/test_no_powerful_models.py` verifying no runtime module, config, or internal ECP classifier references powerful models - [x] T086 [P] [US8] Implement contract parity test in `tests/contract/test_contract_parity.py` checking all schema versions match 1.0.0 - [x] T087 [P] [US8] Implement multi-extractor golden-set quality tests across the 20 reference cases in `tests/quality/test_golden_reference_20.py` - [x] T088 [P] [US8] Implement adversarial prompt injection evaluation in `tests/quality/test_prompt_injection_guard.py` - [x] T089 [P] [US8] Implement automated cost budget verification in `tests/quality/test_cost_budget.py` asserting median per-article cost <= $0.0006 ### Implementation for User Story 8 - [x] T091 [US8] Implement CI multi-tier trigger runner in `scripts/ci_check.py` distinguishing PR gates, prompt/schema/model change gates, and pre-promotion gates per Doc 04 (including reproducible lockfile package build and metadata verification) - [x] T092 [US8] Implement a minimal holdout and slice evaluation aggregator in `evals/eval_runner.py` aggregating Promptfoo output without duplicating prompt execution, computing slice pass rates (≥95%), block precision/recall/F1, metadata accuracy, correct vs unauthorized repairs, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, tag acceptance, and schema validity - [x] T093 [US8] Implement empirical latency and cost calibration recorder for staging gates in `tests/load/test_load_100_art_per_hour.py` - [x] T094 [US8] Implement the release packaging script in `scripts/build_release_metadata.py` generating `src/core/release-metadata.json` with `release_version`, `runtime_config_sha256`, real prompt hashes, schema versions, certified provider/model mappings for both logical roles, and the certified ECP classifier configuration hash bundled inside the distributable package **Checkpoint**: Quality gates, security test suites, and CI evaluation infrastructure are verified. --- ## Phase 11: User Story 9 - Production Runbook Operations & Lifecycle Management (Priority: P3) **Goal**: Implement operational commands (`preflight`, `smoke`, `reconcile`, `telemetry_flush`), SQLite native backup/restore, signal handling (`SIGTERM`/`SIGINT`), rollback, and credential/model rotations. **Independent Test**: Run preflight checks against `src/core/release-metadata.json`, execute smoke tests with fixtures, perform native SQLite backup and restore, simulate `SIGTERM` graceful shutdown, and verify credential and model rotation procedures. ### Tests for User Story 9 - [x] T095 [P] [US9] Implement integration tests for operational resilience in `tests/integration/test_operations_resilience.py` covering native backup/restore, graceful shutdown signals, rollback, credential rotation, certified model rotation, and uncertified rotation rejection - [x] T096 [P] [US9] Implement unit tests for preflight verification against release metadata in `tests/unit/test_preflight_certification.py` - [x] T097 [P] [US9] Implement unit tests for smoke test execution in `tests/unit/test_smoke_cli.py` ### Implementation for User Story 9 - [x] T098 [US9] Implement preflight validation CLI in `src/cli/preflight.py` checking clock sync, release metadata hashes, schemas, SQLite access, filesystem permissions / atomic rename, minimum disk space, valid/active credentials, certified cheap models (including ECP), trace content policy, and ECP classifier/schema access (remote Langfuse outage does not block preflight) - [x] T099 [US9] Implement smoke test CLI in `src/cli/smoke.py` processing fixture and verifying end-to-end pipeline health (fingerprint, state transitions, LLM call, ECP gate, manifest, Markdown, Langfuse trace, cost/latency within approved staging baseline, and idempotent re-execution) - [x] T100 [US9] Implement state and artifact reconciliation CLI in `src/cli/reconcile.py` (completed states vs files/hashes, final files without state, orphan temp files, pending telemetry, duplicate fingerprints, reconciliation report, safe cleanup without altering editorial content) - [x] T101 [US9] Implement full graceful shutdown signal handling (`SIGTERM`, `SIGINT`) in `src/cli/consolidate.py` (stop accepting new units, complete or safely preserve active unit state, close transactions, flush files, attempt telemetry flush, preserve unsent events in SQLite, and exit with coherent exit code) - [x] T102 [US9] Validate and update operational procedures and structure in the normative runtime runbook (`docs/structured_extraction/06_Runbook_Producao_Runtime.md`) for the 11 deployment steps, preflight, smoke, backup/restore, retention, reconciliation, rollback, and credential/model rotations (ready for staging limit incorporation) **Checkpoint**: All operational procedures and lifecycle commands are testable and functional. --- ## Phase 12: Polish, Verification Gates & Release Sign-Off **Purpose**: Final end-to-end execution, evaluation runs, release evidence report generation, and formal sign-off. - [x] T103 [P] Execute multi-parser static policy verification across text-processing runtime modules, content tests/assertions, JSON schemas, and Promptfoo YAML configurations of this feature via `python -m tests.scripts.check_zero_regex` - [x] T104 [P] Execute all 9 contract test suites; the Article Input contract MUST validate all 20 real reference units via `pytest tests/contract -v` - [x] T105 Execute the complete unit and mock integration suites, including idempotency, concurrent replay, and full CLI contract validation (`completed_text`, `rejected_ecp`, `failed_validation`, `failed_processing`, idempotent result, config/preflight errors with stdout/stderr and exit codes) via `pytest tests/unit tests/integration -v` - [x] T106 Execute all 8 security scenario tests (`SEC-001` to `SEC-008`) via `pytest tests/security -v` - [x] T107 Execute all 10 fault injection scenario tests (`FLT-001` to `FLT-010`) via `pytest tests/fault_injection -v` - [x] T108 Run end-to-end quickstart validation scenarios A, B, C, D per `quickstart.md` - [x] T109 Execute Promptfoo over the 20-case regression set, production golden set, and protected holdout using the exact production prompts and schemas, preserving per-case and per-slice results - [x] T110 Execute the production-equivalent 100 articles/hour staging run using the installed lockfile-built package, generate the complete normative staging and release-evidence report, verify all 11 zero-tolerance invariants, verify formal absence of forbidden architectural patterns across this feature's runtime codebase, dependencies/lockfile, prompts, schemas, functional configs, and packaging artifacts (verifying no LangChain, LangGraph, agents, API, internal batch/worker pool, Postgres, external queue, object storage, keyword dictionaries, self-healing, powerful models in runtime, online Promptfoo), and obtain approval for cost/latency/storage/fallback limits - [x] T111 Incorporate the approved staging limits into the normative runbook (`docs/structured_extraction/06_Runbook_Producao_Runtime.md`), update README/CLI documentation, and record formal release sign-off --- ## Dependencies & Execution Order ### Phase Dependencies - **Setup (Phase 1)**: No dependencies — starts immediately. - **Foundational (Phase 2)**: Depends on Setup completion — **BLOCKS all user stories**. - **Model Gateway (Phase 3, US6)**: Depends on Foundational — **BLOCKS Hygiene & Enrichment**. - **User Story 1 (Phase 4, P1)**: Depends on Foundational — delivers initial contract validation (Article + ECP), candidate extraction & idempotency. - **User Story 2 (Phase 5, P1)**: Depends on US1 and US6 (Model Gateway) — delivers 10-step LLM extractive hygiene & repairs. - **User Story 3 (Phase 6, P1)**: Depends on US2 — delivers mandatory ECP inherence gate and zero-Markdown rejection. - **User Story 4 (Phase 7, P2)**: Depends on US3 and US6 (Model Gateway) — delivers post-ECP sentiment & tag enrichment and prompts contract testing. - **User Story 5 (Phase 8, P2)**: Depends on US4 — delivers canonical Markdown & atomic persistence across all terminal outcomes. - **User Story 7 (Phase 9, P3)**: Depends on US5 & US6 — delivers Langfuse observability & telemetry queue. - **User Story 8 (Phase 10, P3)**: Depends on US6 and US7 — delivers quality gates, CI multi-tier evals, golden set & load benchmark. - **User Story 9 (Phase 11, P3)**: Depends on US5, US7, and US8 — delivers operational runbooks & resilience commands. - **Polish & Gates (Phase 12)**: Depends on all user stories being complete (T110 executes staging and calibrates limits; T111 incorporates limits into documentation and signs off release). --- ## Parallel Execution Opportunities ```bash # Launch Foundational independent tasks in parallel: Task: "T007 Implement input byte size limiter in src/core/limits.py" Task: "T008 Implement deterministic fingerprint calculator in src/core/fingerprint.py" Task: "T009 Implement CandidateObject models in src/candidate/models.py" Task: "T012 Implement structured JSON logging in src/observability/structured_logger.py" Task: "T013 Implement atomic writer and manifest generator in src/storage/file_store.py" # Launch User Story 1 test tasks in parallel: Task: "T023 Contract test for Article Input in tests/contract/test_article_input_contract.py" Task: "T024 Contract test for ECP Snapshot in tests/contract/test_ecp_snapshot_contract.py" Task: "T025 Contract test for Candidates Payload in tests/contract/test_candidates_payload_contract.py" Task: "T026 Unit tests for input limits in tests/unit/test_input_limits.py" Task: "T027 Unit tests for fingerprint in tests/unit/test_fingerprint.py" Task: "T028 Unit tests for candidate extraction in tests/unit/test_candidate_parser.py" Task: "T029 Unit tests for equivalence mapping in tests/unit/test_equivalence_mapping.py" ``` --- ## Implementation Strategy ### Foundation & Incremental Pipeline Flow 1. Complete Phase 1: Setup 2. Complete Phase 2: Foundational (blocking prerequisites & shared stores) 3. Complete Phase 3: User Story 6 (Model Gateway infrastructure & cheap model enforcement) 4. Complete Phase 4: User Story 1 (Ingestion, Article/ECP contract validation, candidates, idempotency) 5. Complete Phase 5: User Story 2 (Extractive hygiene & micro-repairs) 6. Complete Phase 6: User Story 3 (Mandatory ECP gate & relevance enforcement) 7. Complete Phase 7: User Story 4 (Post-ECP sentiment & native tags enrichment, prompt contracts) 8. Complete Phase 8: User Story 5 (Canonical Markdown & atomic persistence across all terminal outcomes) 9. Complete Phase 9: User Story 7 (Direct Langfuse observability & telemetry queue) 10. Complete Phase 10: User Story 8 (Automated quality evaluation, golden set & CI gates) 11. Complete Phase 11: User Story 9 (Production runbook operations & lifecycle management) 12. Complete Phase 12: Polish, Verification Gates, Staging Calibration & Release Sign-Off