35 KiB
Implementation Tasks: Article Consolidation and Hygiene Runtime
Feature: Article Consolidation and Hygiene Runtime (specs/006-article-consolidation-runtime/spec.md)
Branch: 006-article-consolidation-runtime | Date: 2026-08-23 | Plan: plan.md
Status: Ready for Execution
Phase 1: Setup (Shared Infrastructure & Tooling)
Purpose: Project initialization, dependency management, and quality verification tooling.
- T001 Initialize the package structure and lockfile using only dependencies approved by the implementation plan in
pyproject.toml, recording for every new dependency: requirement served, standard-library alternative, security impact, maintenance impact, license, size impact, and startup impact - T002 [P] Implement multi-parser static policy verification script in
tests/scripts/check_zero_regex.py: checking Python AST for imports and direct calls ofreor any regular-expression engine/API in the scoped text-processing modules, including aliases, without inspecting internals of transitive dependencies, rejectingpatternkeys in JSON schemas via JSON parser, and validating Promptfoo YAML configurations via YAML parser (failing on regex assertions, semanticcontains/not-containsassertions, LLM-as-a-judge for grounding, powerful models as judge, and configurations relying solely on global averages without per-case and per-slice gates) - T003 [P] Configure Promptfoo test environment and suite settings in
evals/promptfoo.config.yamlstrictly following policy constraints (nocontains/not-containssemantic decisions, no LLM-as-a-judge for grounding, no powerful models as judge, and per-case and per-slice assertion gates) - T004 [P] Create initial 20-case reference regression dataset in
evals/reference_20/, converting each of the 20 reference articles into an individual unit file associated with a valid, versioned canonical ECP snapshot per Test Plan §4.1 - T005 [P] Create local validation configuration fixture in
runtime_config.local.json
Phase 2: Foundational (Blocking Prerequisites & Shared Core)
Purpose: Core infrastructure, base models, SQLite WAL store, atomic file writer, manifest generator, and configuration engine that MUST be complete before pipeline execution.
CRITICAL: No user story implementation can begin until this foundational phase is complete.
- T006 Implement configuration loading, validation, and exact-byte SHA-256 hash verification in
src/core/config.py - T007 [P] Implement input byte size limiter with fail-before-provider policy in
src/core/limits.py - T008 [P] Implement deterministic canonical SHA-256 execution fingerprint calculator in
src/core/fingerprint.py - T009 [P] Implement
CandidateObjectand text repair dataclasses insrc/candidate/models.py - T010 Implement SQLite WAL store in
src/storage/sqlite_store.pywith short transactions, busy timeout, native backup/restore API, and atomic fingerprint claim logic (checking completed fingerprint before remote calls, returning existing result, and safely resuming/reusing concurrent executions) - T011 Implement Python explicit state machine and SQLite transition logger in
src/core/state_machine.py - T012 [P] Implement structured JSON logging with all normative fields, structural authorization-header sanitization, and exact replacement of known environment-secret values in
src/observability/structured_logger.py, strictly omitting full ECP, full HTML, and full article text - T013 [P] Implement the atomic filesystem writer and shared manifest generator in
src/storage/file_store.py, writing temporary files in the same destination filesystem, flushing, closing, verifying exact SHA-256 hashes, and performing atomic rename (os.replace) without cross-filesystem moves, complying withmanifest-output.schema.jsonand the 16 normative error codes (shared acrosscompleted_text,rejected_ecp,failed_validation, andfailed_processing) - T014 [P] Implement contract tests for runtime configuration in
tests/contract/test_runtime_config_contract.py - T015 [P] Implement contract tests for manifest output schema in
tests/contract/test_manifest_output_contract.py
Checkpoint: Foundation ready — Model Gateway and Pipeline components can now proceed.
Phase 3: User Story 6 - Model Gateway Infrastructure & Cheap Model Enforcement
Goal: Agnostic Model Gateway, provider adapters, technical retries, and preflight cheap model enforcement (MUST exist before any LLM hygiene or enrichment call).
Tests for Model Gateway
- T016 [P] [US6] Implement unit tests for Model Gateway client and adapters in
tests/unit/test_model_gateway.pycovering normative scenariosLLM-001toLLM-012(logical roles, pricing, token tracking, timeouts) - T017 [P] [US6] Implement fault injection tests for gateway transient errors and failovers in
tests/fault_injection/test_gateway_faults.pycovering provider fault scenarios
Implementation for Model Gateway
- T018 [US6] Implement agnostic Model Gateway client managing logical roles (
runtime_primary,runtime_fallback) and token pricing calculations insrc/gateway/client.py - T019 [US6] Implement minimal HTTP adapters for Groq and DeepSeek using
httpxinsrc/gateway/adapters.py - T020 [US6] Implement limited technical retries for timeout, connection interruption/reset, HTTP 429 with configured backoff up to limit, HTTP 5xx, and empty technical responses in
src/gateway/client.py - T021 [US6] Implement immediate semantic fallback from
runtime_primarytoruntime_fallbackon schema or grounding failure without retrying on the same model insrc/gateway/client.py - T022 [US6] Implement a single, reusable certified-configuration validation in
src/core/config.pyrejecting any uncertified or powerful models across all runtime roles (including any internal ECP LLM)
Checkpoint: Model Gateway implementation and configuration enforcement are ready for pipeline integration; production certification occurs only after Phase 12 gates.
Phase 4: User Story 1 - Single Article Ingestion, Contract Validation, and Candidate Preparation (Priority: P1)
Goal: Ingest single article units, validate contracts (Article and ECP) locally before any remote call, compute deterministic fingerprint, reject batch wrappers, and extract structured candidates without regular expressions.
Independent Test: Provide single article JSON objects (valid, corrupt, batch wrapper) and ECP snapshots, verifying schema validation, SQLite state initialization (received, validated), deterministic candidate ID generation, and immediate pre-remote termination with exact error codes.
Tests for User Story 1
- T023 [P] [US1] Implement contract test for Article Input schema against all 20 real reference units in
tests/contract/test_article_input_contract.py - T024 [P] [US1] Implement contract test for ECP Snapshot schema and local
referencing.Registryresolution intests/contract/test_ecp_snapshot_contract.py - T025 [P] [US1] Implement contract test for Candidates Payload schema in
tests/contract/test_candidates_payload_contract.py - T026 [P] [US1] Implement unit tests for input limits, validation, and error code mapping in
tests/unit/test_input_limits.pycovering scenariosIN-001toIN-015and proving zero remote provider, remote Langfuse, or classifier calls on local failure - T027 [P] [US1] Implement unit tests for deterministic fingerprint calculation and idempotency claims in
tests/unit/test_fingerprint.pycovering scenariosID-001toID-010 - T028 [P] [US1] Implement unit tests for candidate extraction without regex in
tests/unit/test_candidate_parser.pycovering scenariosPAR-001toPAR-010 - T029 [P] [US1] Implement unit tests for cross-extractor sequence equivalence mapping in
tests/unit/test_equivalence_mapping.pycovering scenariosCAN-001toCAN-010
Implementation for User Story 1
- T030 [P] [US1] Create executable article and ECP fixtures in
examples/sample_article_valid.json,examples/sample_article_tangential.json, andexamples/sample_ecp_snapshot.json - T031 [US1] Implement local canonical ECP schema resolution and registration via
referencing.Registry(disabling HTTP network fetching) insrc/ecp/adapter.py - T032 [US1] Implement local pre-call input validation and batch wrapper rejection (
"articles": false) insrc/core/config.py, preserving unknown fields in the recorded original input while ignoring them during processing - T033 [US1] Implement structural candidate parsing in
src/candidate/parser.pyusing DOM for HTML, CommonMark AST for Markdown, JSON parsing for JSON-LD, URL parsing, Unicode normalization, and an appropriate multilingual tokenizer/segmenter and language detector, handling malformed HTML safely and invalid JSON-LD through a controlled warning - T034 [US1] Implement non-destructive candidate equivalence mapping using
difflib.SequenceMatcherinsrc/candidate/equivalence.py - T035 [US1] Implement deterministic source URL and publication date resolution using the exact normative priorities and date-consensus rule, and prepare title, subtitle, and author candidates using their normative source priorities in
src/candidate/parser.py, omitting invalid dates and strictly forbidding delimiter-based author splitting - T036 [US1] Connect initial validation to state machine
received → validatedinsrc/core/state_machine.py, ensuring transition only occurs after article, ECP, config,selected_extractor, size limit, and minimum content checks pass - T037 [US1] Implement CLI ingestion entrypoint in
src/cli/consolidate.pywith full idempotency checks (querying fingerprint before LLM, returning existing result on match, resuming incomplete runs, claiming atomic execution), emitting structured JSON, complete manifest on stdout when fingerprint exists, technical envelope on unparseable JSON, persisting manifest, and enforcing exact exit codes (0: completed/rejected_ecp,1: invalid article/ECP,2: config/preflight error,3: failed processing,4: persistence failure)
Checkpoint: User Story 1 is independently functional, validating contracts and preparing candidates locally.
Phase 5: User Story 2 - Mandatory LLM Extractive Hygiene & Controlled Text Repairs (Priority: P1)
Goal: Execute 100% LLM extractive hygiene over candidate payloads, enforce the 10-step validation harness, permit only 5 closed micro-repair categories, and reject ungrounded edits without regex.
Independent Test: Feed candidate payloads with consensus, divergence, and noise into the hygiene harness, verifying that the LLM returns only candidate IDs and repairs, ungrounded IDs trigger GROUNDING_VIOLATION, invalid repairs are discarded with originals preserved, and valid intermediate Markdown is assembled.
Tests for User Story 2
- T038 [P] [US2] Implement contract test for Hygiene Response schema in
tests/contract/test_hygiene_response_contract.py - T039 [P] [US2] Implement contract test for Repair Operations schema in
tests/contract/test_repair_operations_contract.py - T040 [P] [US2] Implement unit tests for 10-step hygiene validation harness in
tests/unit/test_hygiene_harness.pycovering scenariosHYG-001toHYG-021(testing context exclusions, grounding enforcement, and candidate ID validation) - T041 [P] [US2] Implement unit tests for controlled text repairs without regex in
tests/unit/test_repairs_validator.pycovering scenariosREP-001toREP-016(5 closed categories, sensitive entity protection, exact fragment targeting, and Unicode/NLP-based diff validation without uncalibrated numeric thresholds) - T042 [P] [US2] Implement Promptfoo evaluation suite for
article_content_hygieneprompt inevals/promptfoo.config.yamlvalidating context exclusions (no raw JSON, no full HTML, no logs, no secrets, no self-healing, no semanticcontainsassertions)
Implementation for User Story 2
- T043 [P] [US2] Author normative versioned prompt in
prompts/article_content_hygiene.v1.txtfollowing the exact 6-block ordering (Doc 07 §5.3) - T044 [US2] Implement the minimal candidate/context projection builder in
src/hygiene/harness.py, excluding full raw JSON, full HTML, other-article data, logs, secrets, full ECP when minimal identity is sufficient, rejected prior responses except required technical fallback metadata, self-healing instructions, and language-specific semantic keyword examples, while delimiting article content strictly as untrusted data - T045 [US2] Implement the micro-repair validator in
src/hygiene/repairs.pyenforcing the 5 closed categories, exact-fragment targeting, Unicode/NLP-based comparison, sensitive-entity preservation, and audit decisions without regex (no quantitative similarity threshold unless one is later approved through the golden-set evaluation) - T046 [US2] Implement 10-step hygiene harness in
src/hygiene/harness.pyvalidating candidate IDs, ordering, links/images, minimum content, and grounding - T047 [US2] Implement grounded intermediate Markdown assembler in
src/hygiene/assembler.py - T048 [US2] Implement decoupling between schema failures (semantic fallback) and grounding violations (immediate invalidation) in
src/hygiene/harness.py - T049 [US2] Implement the conservative deterministic hygiene fallback in
src/hygiene/harness.pyusing only theselected_extractorstructural backbone, removing only structurally invalid elements, without regex, keyword dictionaries, semantic advertisement filtering, or content-quality inference; use it only when grounding and minimum-content requirements are satisfied, otherwise terminate withHYGIENE_FAILED - T050 [US2] Connect hygiene stage to state machine
validated → content_cleanedinsrc/core/state_machine.py - T051 [US2] Integrate hygiene harness execution and error handling into
src/cli/consolidate.py
Checkpoint: User Stories 1 and 2 operate together, performing grounded extractive hygiene and controlled repairs.
Phase 6: User Story 3 - Mandatory ECP Gate & Relevance Enforcement (Priority: P1)
Goal: Evaluate intermediate sanitized Markdown against canonical ECP Snapshot using src.classifier.InherenceClassifier, transitioning inherent articles to ecp_approved and non-inherent articles to ecp_rejected with zero Markdown generated.
Independent Test: Submit intermediate Markdown to ECP adapter with profiles across all 4 categories (DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED), verifying that only inherent articles proceed to enrichment, while non-inherent articles persist <fingerprint>.result.json with status rejected_ecp and produce no .md file.
Tests for User Story 3
- T052 [P] [US3] Implement unit tests for ECP adapter invoking
InherenceClassifierintests/unit/test_ecp_adapter.pycovering scenariosECP-001toECP-009(full output validation, grounded evidence check, tier tracking, cheap model enforcement) - T053 [P] [US3] Implement integration test for ECP rejection producing zero Markdown files in
tests/integration/test_ecp_rejection_flow.pycovering scenarioOUT-008
Implementation for User Story 3
- T054 [US3] Implement the ECP classification adapter in
src/ecp/adapter.pyinvokingsrc.classifier.InherenceClassifierthrough its public contract and certified configuration, validatingcategory,is_inherent,confidence,rationale, andevidences, asserting that all evidence fragments belong to the intermediate Markdown, and recording any classifier tier or LLM generation exposed by the classifier - T055 [US3] Make the ECP adapter consume the shared certified-configuration validation from
src/core/config.py, verifying the ECP classifier configuration against packaged release metadata without duplicating hash or certification logic - T056 [US3] Connect ECP inherence gate to state machine
content_cleaned → ecp_approved | ecp_rejectedinsrc/core/state_machine.py - T057 [US3] Implement
rejected_ecpterminal flow writing manifest with statusrejected_ecp(ECP_REJECTED) and strictly omitting Markdown output insrc/storage/file_store.py - T058 [US3] Integrate ECP gate execution and error handling into
src/cli/consolidate.py
Checkpoint: Core pipeline evaluates inherence and enforces the strict ECP publishing gate.
Phase 7: User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2)
Goal: Enrich ECP-approved articles with entity-relative sentiment and 3 to 8 native language tags supported by textual evidence IDs, strictly decoupled from body text.
Independent Test: Submit approved intermediate Markdown and minimal ECP identity (qid, canonical_name) to enrichment harness, verifying sentiment extraction, tag bounding (3–8), NLP uniqueness without regex, evidence grounding, and failure handling without body modification.
Tests for User Story 4
- T059 [P] [US4] Implement contract test for Enrichment Response schema in
tests/contract/test_enrichment_response_contract.py - T060 [P] [US4] Implement unit tests for entity sentiment and native tags validator in
tests/unit/test_enrichment_harness.pycovering scenariosENR-001toENR-009(sentiment relative to entity, tag bounding, evidence IDs, context exclusions) - T061 [P] [US4] Implement Promptfoo evaluation suite for
article_sentiment_tagsprompt inevals/promptfoo.config.yamlverifying minimal ECP identity context and prohibiting semanticcontainsassertions - T062 [P] [US4] Implement contract tests for both versioned prompts in
tests/contract/test_prompts_contract.pyverifying 6-block sequence, semver parsing without regex, SHA-256 calculation, and Promptfoo parity
Implementation for User Story 4
- T063 [P] [US4] Author normative versioned prompt in
prompts/article_sentiment_tags.v1.txtfollowing the exact 6-block ordering (Doc 07 §5.3) - T064 [US4] Implement enrichment harness in
src/enrichment/harness.pyvalidating sentiment enum, 3–8 unique tags via NLP/Unicode, and evidence candidate IDs, strictly limiting ECP context toqidandcanonical_name(no raw ECP snapshot, no keyword lists) - T065 [US4] Connect enrichment stage to state machine
ecp_approved → enrichedinsrc/core/state_machine.py - T066 [US4] Implement fallback routing and terminal
ENRICHMENT_FAILEDhandling (blocking Markdown generation on failure) insrc/enrichment/harness.py - T067 [US4] Integrate enrichment stage execution and error handling into
src/cli/consolidate.py
Checkpoint: User Story 4 delivers structured sentiment and native tags metadata for inherent articles.
Phase 8: User Story 5 - Canonical Markdown Rendering & Atomic Persistence (Priority: P2)
Goal: Render canonical Markdown with YAML front matter, generate machine-readable .result.json manifests, execute atomic filesystem writes (temp + rename), and maintain strict SQLite state consistency.
Independent Test: Verify generated .md and .result.json files, validating YAML front matter structure, 64-character SHA-256 content hashes, atomic rename lifecycle, and hash-based reconciliation of interrupted writes.
Tests for User Story 5
- T068 [P] [US5] Implement unit tests for canonical YAML front matter and Markdown body renderer in
tests/unit/test_markdown_renderer.pycovering scenariosOUT-002toOUT-007,OUT-009, andOUT-010(grounding, formatting, front matter structure) - T069 [P] [US5] Implement unit tests for atomic file writes, permissions, and 64-character hash verification in
tests/unit/test_file_store.pycovering scenariosOUT-001,OUT-011, andOUT-012 - T070 [P] [US5] Implement unit tests for SQLite WAL state persistence and crash reconciliation in
tests/unit/test_sqlite_store.pycovering crash recovery and multi-terminal state consistency
Implementation for User Story 5
- T071 [US5] Implement canonical YAML front matter and Markdown body renderer in
src/storage/markdown_renderer.py(H1 title, italic subtitle when present with no extra blank line when absent, canonical body order, grounded links/images, omitting author/date/sentiment/tags/ECP from body) - T072 [US5] Integrate the atomic filesystem writer from
src/storage/file_store.pywith SQLite completion state insrc/storage/sqlite_store.pywithin the same logical completion unit (without distributed transactions), resolving crash divergence via hash-based reconciliation for ID-009, ensuring no terminal state is exposed as completed while files and hashes disagree, and preserving the last safe state without exposing partial final artifact pairs onPERSISTENCE_FAILED(coveringcompleted_text,rejected_ecp,failed_validation, andfailed_processing) - T073 [US5] Connect final persistence to state machine
enriched → completed_textinsrc/core/state_machine.py - T074 [US5] Integrate final persistence and reconciliation into
src/cli/consolidate.pyandsrc/cli/reconcile.py
Checkpoint: End-to-end pipeline produces atomic published Markdown and manifests with SQLite consistency.
Phase 9: User Story 7 - Direct Langfuse Observability, Log Sanitization & Telemetry Queue (Priority: P3)
Goal: Transmit traces, spans, generations, and metrics directly to Langfuse, enforce secret redaction, degrade gracefully to SQLite pending_telemetry on network outage, and provide operational telemetry flush.
Independent Test: Process articles with Langfuse available and blocked, checking trace structure (8 stable spans), secret redaction in stderr logs, SQLite queue insertion on outage, and flush execution via telemetry_flush CLI.
Tests for User Story 7
- T075 [P] [US7] Implement unit tests for Langfuse tracer, secret redaction, and offline queue in
tests/unit/test_langfuse_tracer.pycovering scenariosOBS-001toOBS-011(8 stable spans, generation attributes, metric dimensions, cardinality guards) - T076 [P] [US7] Implement integration tests for telemetry degradation, deduplication, and atomic replay in
tests/integration/test_telemetry_degradation.py - T077 [P] [US7] Implement specialized security tests for authorization-header and secret redaction in logs and SDK exceptions (SEC-006) in
tests/security/test_secret_redaction.py
Implementation for User Story 7
- T078 [US7] Implement Langfuse observability integration in
src/observability/langfuse_tracer.pymanaging traces with 8 stable spans (validation,candidate_preparation,hygiene,grounding_validation,ecp_gate,enrichment,rendering,persistence) and generations per LLM attempt; the validation span MUST be buffered or materialized only after successful local validation (no remote Langfuse traffic may occur while terminating local validations are running) - T079 [US7] Implement local SQLite queue insertion for telemetry events during Langfuse network outages and atomic flush procedure in
src/observability/langfuse_tracer.py - T080 [US7] Implement runtime-observable metric emission in
src/observability/langfuse_tracer.pyusing exactly each metric and its dimensions from Doc 05 / FR-067, enforcing cardinality restrictions and emittingprompt_review_signal_totalwithout self-healing or a parallel metrics store (release-wide aggregation of the 11 critical invariants remains the responsibility of T090) - T081 [US7] Implement operational telemetry flush command in
src/cli/telemetry_flush.py, ensuring events are marked flushed only upon confirmed delivery andtelemetry_pending_totalreturns to zero - T082 [US7] Configure the 3 mandatory Langfuse dashboards (Runtime Health, Quality, Future Review Signals) and preserve reproducible setup evidence without creating a parallel metrics system
- T083 [US7] Integrate observability lifecycle and secret redaction into
src/cli/consolidate.py
Checkpoint: Observability is complete, compliant with the metric catalog, and resilient against outages.
Phase 10: User Story 8 - Quality Gates, Zero Regex Verification & Promptfoo Evaluation (Priority: P3)
Goal: Execute comprehensive automated quality gates, Promptfoo offline evaluations, multi-extractor golden regression across 20 reference cases, and zero-regex verification.
Independent Test: Run pytest tests/quality/, pytest evals/, and tests/scripts/check_zero_regex.py, asserting that all 11 critical quality assertions pass without manual intervention.
Tests for User Story 8
- T084 [P] [US8] Implement automated test in
tests/quality/test_zero_regex_enforcement.pyexecutingtests/scripts/check_zero_regex.pyacross codebase, schemas, and Promptfoo YAML - T085 [P] [US8] Implement automated test in
tests/quality/test_no_powerful_models.pyverifying no runtime module, config, or internal ECP classifier references powerful models - T086 [P] [US8] Implement contract parity test in
tests/contract/test_contract_parity.pychecking all schema versions match 1.0.0 - T087 [P] [US8] Implement multi-extractor golden-set quality tests across the 20 reference cases in
tests/quality/test_golden_reference_20.py - T088 [P] [US8] Implement adversarial prompt injection evaluation in
tests/quality/test_prompt_injection_guard.py - T089 [P] [US8] Implement automated cost budget verification in
tests/quality/test_cost_budget.pyasserting median per-article cost <= $0.0006
Implementation for User Story 8
- T091 [US8] Implement CI multi-tier trigger runner in
scripts/ci_check.pydistinguishing PR gates, prompt/schema/model change gates, and pre-promotion gates per Doc 04 (including reproducible lockfile package build and metadata verification) - T092 [US8] Implement a minimal holdout and slice evaluation aggregator in
evals/eval_runner.pyaggregating Promptfoo output without duplicating prompt execution, computing slice pass rates (≥95%), block precision/recall/F1, metadata accuracy, correct vs unauthorized repairs, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, tag acceptance, and schema validity - T093 [US8] Implement empirical latency and cost calibration recorder for staging gates in
tests/load/test_load_100_art_per_hour.py - T094 [US8] Implement the release packaging script in
scripts/build_release_metadata.pygeneratingsrc/core/release-metadata.jsonwithrelease_version,runtime_config_sha256, real prompt hashes, schema versions, certified provider/model mappings for both logical roles, and the certified ECP classifier configuration hash bundled inside the distributable package
Checkpoint: Quality gates, security test suites, and CI evaluation infrastructure are verified.
Phase 11: User Story 9 - Production Runbook Operations & Lifecycle Management (Priority: P3)
Goal: Implement operational commands (preflight, smoke, reconcile, telemetry_flush), SQLite native backup/restore, signal handling (SIGTERM/SIGINT), rollback, and credential/model rotations.
Independent Test: Run preflight checks against src/core/release-metadata.json, execute smoke tests with fixtures, perform native SQLite backup and restore, simulate SIGTERM graceful shutdown, and verify credential and model rotation procedures.
Tests for User Story 9
- T095 [P] [US9] Implement integration tests for operational resilience in
tests/integration/test_operations_resilience.pycovering native backup/restore, graceful shutdown signals, rollback, credential rotation, certified model rotation, and uncertified rotation rejection - T096 [P] [US9] Implement unit tests for preflight verification against release metadata in
tests/unit/test_preflight_certification.py - T097 [P] [US9] Implement unit tests for smoke test execution in
tests/unit/test_smoke_cli.py
Implementation for User Story 9
- T098 [US9] Implement preflight validation CLI in
src/cli/preflight.pychecking clock sync, release metadata hashes, schemas, SQLite access, filesystem permissions / atomic rename, minimum disk space, valid/active credentials, certified cheap models (including ECP), trace content policy, and ECP classifier/schema access (remote Langfuse outage does not block preflight) - T099 [US9] Implement smoke test CLI in
src/cli/smoke.pyprocessing fixture and verifying end-to-end pipeline health (fingerprint, state transitions, LLM call, ECP gate, manifest, Markdown, Langfuse trace, cost/latency within approved staging baseline, and idempotent re-execution) - T100 [US9] Implement state and artifact reconciliation CLI in
src/cli/reconcile.py(completed states vs files/hashes, final files without state, orphan temp files, pending telemetry, duplicate fingerprints, reconciliation report, safe cleanup without altering editorial content) - T101 [US9] Implement full graceful shutdown signal handling (
SIGTERM,SIGINT) insrc/cli/consolidate.py(stop accepting new units, complete or safely preserve active unit state, close transactions, flush files, attempt telemetry flush, preserve unsent events in SQLite, and exit with coherent exit code) - T102 [US9] Validate and update operational procedures and structure in the normative runtime runbook (
docs/structured_extraction/06_Runbook_Producao_Runtime.md) for the 11 deployment steps, preflight, smoke, backup/restore, retention, reconciliation, rollback, and credential/model rotations (ready for staging limit incorporation)
Checkpoint: All operational procedures and lifecycle commands are testable and functional.
Phase 12: Polish, Verification Gates & Release Sign-Off
Purpose: Final end-to-end execution, evaluation runs, release evidence report generation, and formal sign-off.
- T103 [P] Execute multi-parser static policy verification across text-processing runtime modules, content tests/assertions, JSON schemas, and Promptfoo YAML configurations of this feature via
python -m tests.scripts.check_zero_regex - T104 [P] Execute all 9 contract test suites; the Article Input contract MUST validate all 20 real reference units via
pytest tests/contract -v - T105 Execute the complete unit and mock integration suites, including idempotency, concurrent replay, and full CLI contract validation (
completed_text,rejected_ecp,failed_validation,failed_processing, idempotent result, config/preflight errors with stdout/stderr and exit codes) viapytest tests/unit tests/integration -v - T106 Execute all 8 security scenario tests (
SEC-001toSEC-008) viapytest tests/security -v - T107 Execute all 10 fault injection scenario tests (
FLT-001toFLT-010) viapytest tests/fault_injection -v - T108 Run end-to-end quickstart validation scenarios A, B, C, D per
quickstart.md - T109 Execute Promptfoo over the 20-case regression set, production golden set, and protected holdout using the exact production prompts and schemas, preserving per-case and per-slice results
- T110 Execute the production-equivalent 100 articles/hour staging run using the installed lockfile-built package, generate the complete normative staging and release-evidence report, verify all 11 zero-tolerance invariants, verify formal absence of forbidden architectural patterns across this feature's runtime codebase, dependencies/lockfile, prompts, schemas, functional configs, and packaging artifacts (verifying no LangChain, LangGraph, agents, API, internal batch/worker pool, Postgres, external queue, object storage, keyword dictionaries, self-healing, powerful models in runtime, online Promptfoo), and obtain approval for cost/latency/storage/fallback limits
- T111 Incorporate the approved staging limits into the normative runbook (
docs/structured_extraction/06_Runbook_Producao_Runtime.md), update README/CLI documentation, and record formal release sign-off
Dependencies & Execution Order
Phase Dependencies
- Setup (Phase 1): No dependencies — starts immediately.
- Foundational (Phase 2): Depends on Setup completion — BLOCKS all user stories.
- Model Gateway (Phase 3, US6): Depends on Foundational — BLOCKS Hygiene & Enrichment.
- User Story 1 (Phase 4, P1): Depends on Foundational — delivers initial contract validation (Article + ECP), candidate extraction & idempotency.
- User Story 2 (Phase 5, P1): Depends on US1 and US6 (Model Gateway) — delivers 10-step LLM extractive hygiene & repairs.
- User Story 3 (Phase 6, P1): Depends on US2 — delivers mandatory ECP inherence gate and zero-Markdown rejection.
- User Story 4 (Phase 7, P2): Depends on US3 and US6 (Model Gateway) — delivers post-ECP sentiment & tag enrichment and prompts contract testing.
- User Story 5 (Phase 8, P2): Depends on US4 — delivers canonical Markdown & atomic persistence across all terminal outcomes.
- User Story 7 (Phase 9, P3): Depends on US5 & US6 — delivers Langfuse observability & telemetry queue.
- User Story 8 (Phase 10, P3): Depends on US6 and US7 — delivers quality gates, CI multi-tier evals, golden set & load benchmark.
- User Story 9 (Phase 11, P3): Depends on US5, US7, and US8 — delivers operational runbooks & resilience commands.
- Polish & Gates (Phase 12): Depends on all user stories being complete (T110 executes staging and calibrates limits; T111 incorporates limits into documentation and signs off release).
Parallel Execution Opportunities
# Launch Foundational independent tasks in parallel:
Task: "T007 Implement input byte size limiter in src/core/limits.py"
Task: "T008 Implement deterministic fingerprint calculator in src/core/fingerprint.py"
Task: "T009 Implement CandidateObject models in src/candidate/models.py"
Task: "T012 Implement structured JSON logging in src/observability/structured_logger.py"
Task: "T013 Implement atomic writer and manifest generator in src/storage/file_store.py"
# Launch User Story 1 test tasks in parallel:
Task: "T023 Contract test for Article Input in tests/contract/test_article_input_contract.py"
Task: "T024 Contract test for ECP Snapshot in tests/contract/test_ecp_snapshot_contract.py"
Task: "T025 Contract test for Candidates Payload in tests/contract/test_candidates_payload_contract.py"
Task: "T026 Unit tests for input limits in tests/unit/test_input_limits.py"
Task: "T027 Unit tests for fingerprint in tests/unit/test_fingerprint.py"
Task: "T028 Unit tests for candidate extraction in tests/unit/test_candidate_parser.py"
Task: "T029 Unit tests for equivalence mapping in tests/unit/test_equivalence_mapping.py"
Implementation Strategy
Foundation & Incremental Pipeline Flow
- Complete Phase 1: Setup
- Complete Phase 2: Foundational (blocking prerequisites & shared stores)
- Complete Phase 3: User Story 6 (Model Gateway infrastructure & cheap model enforcement)
- Complete Phase 4: User Story 1 (Ingestion, Article/ECP contract validation, candidates, idempotency)
- Complete Phase 5: User Story 2 (Extractive hygiene & micro-repairs)
- Complete Phase 6: User Story 3 (Mandatory ECP gate & relevance enforcement)
- Complete Phase 7: User Story 4 (Post-ECP sentiment & native tags enrichment, prompt contracts)
- Complete Phase 8: User Story 5 (Canonical Markdown & atomic persistence across all terminal outcomes)
- Complete Phase 9: User Story 7 (Direct Langfuse observability & telemetry queue)
- Complete Phase 10: User Story 8 (Automated quality evaluation, golden set & CI gates)
- Complete Phase 11: User Story 9 (Production runbook operations & lifecycle management)
- Complete Phase 12: Polish, Verification Gates, Staging Calibration & Release Sign-Off