Files
TextNLPClassifierApp/specs/006-article-consolidation-runtime/tasks.md
T

35 KiB
Raw Blame History

Implementation Tasks: Article Consolidation and Hygiene Runtime

Feature: Article Consolidation and Hygiene Runtime (specs/006-article-consolidation-runtime/spec.md)
Branch: 006-article-consolidation-runtime | Date: 2026-08-23 | Plan: plan.md
Status: Ready for Execution


Phase 1: Setup (Shared Infrastructure & Tooling)

Purpose: Project initialization, dependency management, and quality verification tooling.

  • T001 Initialize the package structure and lockfile using only dependencies approved by the implementation plan in pyproject.toml, recording for every new dependency: requirement served, standard-library alternative, security impact, maintenance impact, license, size impact, and startup impact
  • T002 [P] Implement multi-parser static policy verification script in tests/scripts/check_zero_regex.py: checking Python AST for imports and direct calls of re or any regular-expression engine/API in the scoped text-processing modules, including aliases, without inspecting internals of transitive dependencies, rejecting pattern keys in JSON schemas via JSON parser, and validating Promptfoo YAML configurations via YAML parser (failing on regex assertions, semantic contains/not-contains assertions, LLM-as-a-judge for grounding, powerful models as judge, and configurations relying solely on global averages without per-case and per-slice gates)
  • T003 [P] Configure Promptfoo test environment and suite settings in evals/promptfoo.config.yaml strictly following policy constraints (no contains/not-contains semantic decisions, no LLM-as-a-judge for grounding, no powerful models as judge, and per-case and per-slice assertion gates)
  • T004 [P] Create initial 20-case reference regression dataset in evals/reference_20/, converting each of the 20 reference articles into an individual unit file associated with a valid, versioned canonical ECP snapshot per Test Plan §4.1
  • T005 [P] Create local validation configuration fixture in runtime_config.local.json

Phase 2: Foundational (Blocking Prerequisites & Shared Core)

Purpose: Core infrastructure, base models, SQLite WAL store, atomic file writer, manifest generator, and configuration engine that MUST be complete before pipeline execution.

CRITICAL: No user story implementation can begin until this foundational phase is complete.

  • T006 Implement configuration loading, validation, and exact-byte SHA-256 hash verification in src/core/config.py
  • T007 [P] Implement input byte size limiter with fail-before-provider policy in src/core/limits.py
  • T008 [P] Implement deterministic canonical SHA-256 execution fingerprint calculator in src/core/fingerprint.py
  • T009 [P] Implement CandidateObject and text repair dataclasses in src/candidate/models.py
  • T010 Implement SQLite WAL store in src/storage/sqlite_store.py with short transactions, busy timeout, native backup/restore API, and atomic fingerprint claim logic (checking completed fingerprint before remote calls, returning existing result, and safely resuming/reusing concurrent executions)
  • T011 Implement Python explicit state machine and SQLite transition logger in src/core/state_machine.py
  • T012 [P] Implement structured JSON logging with all normative fields, structural authorization-header sanitization, and exact replacement of known environment-secret values in src/observability/structured_logger.py, strictly omitting full ECP, full HTML, and full article text
  • T013 [P] Implement the atomic filesystem writer and shared manifest generator in src/storage/file_store.py, writing temporary files in the same destination filesystem, flushing, closing, verifying exact SHA-256 hashes, and performing atomic rename (os.replace) without cross-filesystem moves, complying with manifest-output.schema.json and the 16 normative error codes (shared across completed_text, rejected_ecp, failed_validation, and failed_processing)
  • T014 [P] Implement contract tests for runtime configuration in tests/contract/test_runtime_config_contract.py
  • T015 [P] Implement contract tests for manifest output schema in tests/contract/test_manifest_output_contract.py

Checkpoint: Foundation ready — Model Gateway and Pipeline components can now proceed.


Phase 3: User Story 6 - Model Gateway Infrastructure & Cheap Model Enforcement

Goal: Agnostic Model Gateway, provider adapters, technical retries, and preflight cheap model enforcement (MUST exist before any LLM hygiene or enrichment call).

Tests for Model Gateway

  • T016 [P] [US6] Implement unit tests for Model Gateway client and adapters in tests/unit/test_model_gateway.py covering normative scenarios LLM-001 to LLM-012 (logical roles, pricing, token tracking, timeouts)
  • T017 [P] [US6] Implement fault injection tests for gateway transient errors and failovers in tests/fault_injection/test_gateway_faults.py covering provider fault scenarios

Implementation for Model Gateway

  • T018 [US6] Implement agnostic Model Gateway client managing logical roles (runtime_primary, runtime_fallback) and token pricing calculations in src/gateway/client.py
  • T019 [US6] Implement minimal HTTP adapters for Groq and DeepSeek using httpx in src/gateway/adapters.py
  • T020 [US6] Implement limited technical retries for timeout, connection interruption/reset, HTTP 429 with configured backoff up to limit, HTTP 5xx, and empty technical responses in src/gateway/client.py
  • T021 [US6] Implement immediate semantic fallback from runtime_primary to runtime_fallback on schema or grounding failure without retrying on the same model in src/gateway/client.py
  • T022 [US6] Implement a single, reusable certified-configuration validation in src/core/config.py rejecting any uncertified or powerful models across all runtime roles (including any internal ECP LLM)

Checkpoint: Model Gateway implementation and configuration enforcement are ready for pipeline integration; production certification occurs only after Phase 12 gates.


Phase 4: User Story 1 - Single Article Ingestion, Contract Validation, and Candidate Preparation (Priority: P1)

Goal: Ingest single article units, validate contracts (Article and ECP) locally before any remote call, compute deterministic fingerprint, reject batch wrappers, and extract structured candidates without regular expressions.

Independent Test: Provide single article JSON objects (valid, corrupt, batch wrapper) and ECP snapshots, verifying schema validation, SQLite state initialization (received, validated), deterministic candidate ID generation, and immediate pre-remote termination with exact error codes.

Tests for User Story 1

  • T023 [P] [US1] Implement contract test for Article Input schema against all 20 real reference units in tests/contract/test_article_input_contract.py
  • T024 [P] [US1] Implement contract test for ECP Snapshot schema and local referencing.Registry resolution in tests/contract/test_ecp_snapshot_contract.py
  • T025 [P] [US1] Implement contract test for Candidates Payload schema in tests/contract/test_candidates_payload_contract.py
  • T026 [P] [US1] Implement unit tests for input limits, validation, and error code mapping in tests/unit/test_input_limits.py covering scenarios IN-001 to IN-015 and proving zero remote provider, remote Langfuse, or classifier calls on local failure
  • T027 [P] [US1] Implement unit tests for deterministic fingerprint calculation and idempotency claims in tests/unit/test_fingerprint.py covering scenarios ID-001 to ID-010
  • T028 [P] [US1] Implement unit tests for candidate extraction without regex in tests/unit/test_candidate_parser.py covering scenarios PAR-001 to PAR-010
  • T029 [P] [US1] Implement unit tests for cross-extractor sequence equivalence mapping in tests/unit/test_equivalence_mapping.py covering scenarios CAN-001 to CAN-010

Implementation for User Story 1

  • T030 [P] [US1] Create executable article and ECP fixtures in examples/sample_article_valid.json, examples/sample_article_tangential.json, and examples/sample_ecp_snapshot.json
  • T031 [US1] Implement local canonical ECP schema resolution and registration via referencing.Registry (disabling HTTP network fetching) in src/ecp/adapter.py
  • T032 [US1] Implement local pre-call input validation and batch wrapper rejection ("articles": false) in src/core/config.py, preserving unknown fields in the recorded original input while ignoring them during processing
  • T033 [US1] Implement structural candidate parsing in src/candidate/parser.py using DOM for HTML, CommonMark AST for Markdown, JSON parsing for JSON-LD, URL parsing, Unicode normalization, and an appropriate multilingual tokenizer/segmenter and language detector, handling malformed HTML safely and invalid JSON-LD through a controlled warning
  • T034 [US1] Implement non-destructive candidate equivalence mapping using difflib.SequenceMatcher in src/candidate/equivalence.py
  • T035 [US1] Implement deterministic source URL and publication date resolution using the exact normative priorities and date-consensus rule, and prepare title, subtitle, and author candidates using their normative source priorities in src/candidate/parser.py, omitting invalid dates and strictly forbidding delimiter-based author splitting
  • T036 [US1] Connect initial validation to state machine received → validated in src/core/state_machine.py, ensuring transition only occurs after article, ECP, config, selected_extractor, size limit, and minimum content checks pass
  • T037 [US1] Implement CLI ingestion entrypoint in src/cli/consolidate.py with full idempotency checks (querying fingerprint before LLM, returning existing result on match, resuming incomplete runs, claiming atomic execution), emitting structured JSON, complete manifest on stdout when fingerprint exists, technical envelope on unparseable JSON, persisting manifest, and enforcing exact exit codes (0: completed/rejected_ecp, 1: invalid article/ECP, 2: config/preflight error, 3: failed processing, 4: persistence failure)

Checkpoint: User Story 1 is independently functional, validating contracts and preparing candidates locally.


Phase 5: User Story 2 - Mandatory LLM Extractive Hygiene & Controlled Text Repairs (Priority: P1)

Goal: Execute 100% LLM extractive hygiene over candidate payloads, enforce the 10-step validation harness, permit only 5 closed micro-repair categories, and reject ungrounded edits without regex.

Independent Test: Feed candidate payloads with consensus, divergence, and noise into the hygiene harness, verifying that the LLM returns only candidate IDs and repairs, ungrounded IDs trigger GROUNDING_VIOLATION, invalid repairs are discarded with originals preserved, and valid intermediate Markdown is assembled.

Tests for User Story 2

  • T038 [P] [US2] Implement contract test for Hygiene Response schema in tests/contract/test_hygiene_response_contract.py
  • T039 [P] [US2] Implement contract test for Repair Operations schema in tests/contract/test_repair_operations_contract.py
  • T040 [P] [US2] Implement unit tests for 10-step hygiene validation harness in tests/unit/test_hygiene_harness.py covering scenarios HYG-001 to HYG-021 (testing context exclusions, grounding enforcement, and candidate ID validation)
  • T041 [P] [US2] Implement unit tests for controlled text repairs without regex in tests/unit/test_repairs_validator.py covering scenarios REP-001 to REP-016 (5 closed categories, sensitive entity protection, exact fragment targeting, and Unicode/NLP-based diff validation without uncalibrated numeric thresholds)
  • T042 [P] [US2] Implement Promptfoo evaluation suite for article_content_hygiene prompt in evals/promptfoo.config.yaml validating context exclusions (no raw JSON, no full HTML, no logs, no secrets, no self-healing, no semantic contains assertions)

Implementation for User Story 2

  • T043 [P] [US2] Author normative versioned prompt in prompts/article_content_hygiene.v1.txt following the exact 6-block ordering (Doc 07 §5.3)
  • T044 [US2] Implement the minimal candidate/context projection builder in src/hygiene/harness.py, excluding full raw JSON, full HTML, other-article data, logs, secrets, full ECP when minimal identity is sufficient, rejected prior responses except required technical fallback metadata, self-healing instructions, and language-specific semantic keyword examples, while delimiting article content strictly as untrusted data
  • T045 [US2] Implement the micro-repair validator in src/hygiene/repairs.py enforcing the 5 closed categories, exact-fragment targeting, Unicode/NLP-based comparison, sensitive-entity preservation, and audit decisions without regex (no quantitative similarity threshold unless one is later approved through the golden-set evaluation)
  • T046 [US2] Implement 10-step hygiene harness in src/hygiene/harness.py validating candidate IDs, ordering, links/images, minimum content, and grounding
  • T047 [US2] Implement grounded intermediate Markdown assembler in src/hygiene/assembler.py
  • T048 [US2] Implement decoupling between schema failures (semantic fallback) and grounding violations (immediate invalidation) in src/hygiene/harness.py
  • T049 [US2] Implement the conservative deterministic hygiene fallback in src/hygiene/harness.py using only the selected_extractor structural backbone, removing only structurally invalid elements, without regex, keyword dictionaries, semantic advertisement filtering, or content-quality inference; use it only when grounding and minimum-content requirements are satisfied, otherwise terminate with HYGIENE_FAILED
  • T050 [US2] Connect hygiene stage to state machine validated → content_cleaned in src/core/state_machine.py
  • T051 [US2] Integrate hygiene harness execution and error handling into src/cli/consolidate.py

Checkpoint: User Stories 1 and 2 operate together, performing grounded extractive hygiene and controlled repairs.


Phase 6: User Story 3 - Mandatory ECP Gate & Relevance Enforcement (Priority: P1)

Goal: Evaluate intermediate sanitized Markdown against canonical ECP Snapshot using src.classifier.InherenceClassifier, transitioning inherent articles to ecp_approved and non-inherent articles to ecp_rejected with zero Markdown generated.

Independent Test: Submit intermediate Markdown to ECP adapter with profiles across all 4 categories (DIRECT_INHERENT, CONTEXTUAL_INHERENT, TANGENTIAL, NOT_RELATED), verifying that only inherent articles proceed to enrichment, while non-inherent articles persist <fingerprint>.result.json with status rejected_ecp and produce no .md file.

Tests for User Story 3

  • T052 [P] [US3] Implement unit tests for ECP adapter invoking InherenceClassifier in tests/unit/test_ecp_adapter.py covering scenarios ECP-001 to ECP-009 (full output validation, grounded evidence check, tier tracking, cheap model enforcement)
  • T053 [P] [US3] Implement integration test for ECP rejection producing zero Markdown files in tests/integration/test_ecp_rejection_flow.py covering scenario OUT-008

Implementation for User Story 3

  • T054 [US3] Implement the ECP classification adapter in src/ecp/adapter.py invoking src.classifier.InherenceClassifier through its public contract and certified configuration, validating category, is_inherent, confidence, rationale, and evidences, asserting that all evidence fragments belong to the intermediate Markdown, and recording any classifier tier or LLM generation exposed by the classifier
  • T055 [US3] Make the ECP adapter consume the shared certified-configuration validation from src/core/config.py, verifying the ECP classifier configuration against packaged release metadata without duplicating hash or certification logic
  • T056 [US3] Connect ECP inherence gate to state machine content_cleaned → ecp_approved | ecp_rejected in src/core/state_machine.py
  • T057 [US3] Implement rejected_ecp terminal flow writing manifest with status rejected_ecp (ECP_REJECTED) and strictly omitting Markdown output in src/storage/file_store.py
  • T058 [US3] Integrate ECP gate execution and error handling into src/cli/consolidate.py

Checkpoint: Core pipeline evaluates inherence and enforces the strict ECP publishing gate.


Phase 7: User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2)

Goal: Enrich ECP-approved articles with entity-relative sentiment and 3 to 8 native language tags supported by textual evidence IDs, strictly decoupled from body text.

Independent Test: Submit approved intermediate Markdown and minimal ECP identity (qid, canonical_name) to enrichment harness, verifying sentiment extraction, tag bounding (3–8), NLP uniqueness without regex, evidence grounding, and failure handling without body modification.

Tests for User Story 4

  • T059 [P] [US4] Implement contract test for Enrichment Response schema in tests/contract/test_enrichment_response_contract.py
  • T060 [P] [US4] Implement unit tests for entity sentiment and native tags validator in tests/unit/test_enrichment_harness.py covering scenarios ENR-001 to ENR-009 (sentiment relative to entity, tag bounding, evidence IDs, context exclusions)
  • T061 [P] [US4] Implement Promptfoo evaluation suite for article_sentiment_tags prompt in evals/promptfoo.config.yaml verifying minimal ECP identity context and prohibiting semantic contains assertions
  • T062 [P] [US4] Implement contract tests for both versioned prompts in tests/contract/test_prompts_contract.py verifying 6-block sequence, semver parsing without regex, SHA-256 calculation, and Promptfoo parity

Implementation for User Story 4

  • T063 [P] [US4] Author normative versioned prompt in prompts/article_sentiment_tags.v1.txt following the exact 6-block ordering (Doc 07 §5.3)
  • T064 [US4] Implement enrichment harness in src/enrichment/harness.py validating sentiment enum, 3–8 unique tags via NLP/Unicode, and evidence candidate IDs, strictly limiting ECP context to qid and canonical_name (no raw ECP snapshot, no keyword lists)
  • T065 [US4] Connect enrichment stage to state machine ecp_approved → enriched in src/core/state_machine.py
  • T066 [US4] Implement fallback routing and terminal ENRICHMENT_FAILED handling (blocking Markdown generation on failure) in src/enrichment/harness.py
  • T067 [US4] Integrate enrichment stage execution and error handling into src/cli/consolidate.py

Checkpoint: User Story 4 delivers structured sentiment and native tags metadata for inherent articles.


Phase 8: User Story 5 - Canonical Markdown Rendering & Atomic Persistence (Priority: P2)

Goal: Render canonical Markdown with YAML front matter, generate machine-readable .result.json manifests, execute atomic filesystem writes (temp + rename), and maintain strict SQLite state consistency.

Independent Test: Verify generated .md and .result.json files, validating YAML front matter structure, 64-character SHA-256 content hashes, atomic rename lifecycle, and hash-based reconciliation of interrupted writes.

Tests for User Story 5

  • T068 [P] [US5] Implement unit tests for canonical YAML front matter and Markdown body renderer in tests/unit/test_markdown_renderer.py covering scenarios OUT-002 to OUT-007, OUT-009, and OUT-010 (grounding, formatting, front matter structure)
  • T069 [P] [US5] Implement unit tests for atomic file writes, permissions, and 64-character hash verification in tests/unit/test_file_store.py covering scenarios OUT-001, OUT-011, and OUT-012
  • T070 [P] [US5] Implement unit tests for SQLite WAL state persistence and crash reconciliation in tests/unit/test_sqlite_store.py covering crash recovery and multi-terminal state consistency

Implementation for User Story 5

  • T071 [US5] Implement canonical YAML front matter and Markdown body renderer in src/storage/markdown_renderer.py (H1 title, italic subtitle when present with no extra blank line when absent, canonical body order, grounded links/images, omitting author/date/sentiment/tags/ECP from body)
  • T072 [US5] Integrate the atomic filesystem writer from src/storage/file_store.py with SQLite completion state in src/storage/sqlite_store.py within the same logical completion unit (without distributed transactions), resolving crash divergence via hash-based reconciliation for ID-009, ensuring no terminal state is exposed as completed while files and hashes disagree, and preserving the last safe state without exposing partial final artifact pairs on PERSISTENCE_FAILED (covering completed_text, rejected_ecp, failed_validation, and failed_processing)
  • T073 [US5] Connect final persistence to state machine enriched → completed_text in src/core/state_machine.py
  • T074 [US5] Integrate final persistence and reconciliation into src/cli/consolidate.py and src/cli/reconcile.py

Checkpoint: End-to-end pipeline produces atomic published Markdown and manifests with SQLite consistency.


Phase 9: User Story 7 - Direct Langfuse Observability, Log Sanitization & Telemetry Queue (Priority: P3)

Goal: Transmit traces, spans, generations, and metrics directly to Langfuse, enforce secret redaction, degrade gracefully to SQLite pending_telemetry on network outage, and provide operational telemetry flush.

Independent Test: Process articles with Langfuse available and blocked, checking trace structure (8 stable spans), secret redaction in stderr logs, SQLite queue insertion on outage, and flush execution via telemetry_flush CLI.

Tests for User Story 7

  • T075 [P] [US7] Implement unit tests for Langfuse tracer, secret redaction, and offline queue in tests/unit/test_langfuse_tracer.py covering scenarios OBS-001 to OBS-011 (8 stable spans, generation attributes, metric dimensions, cardinality guards)
  • T076 [P] [US7] Implement integration tests for telemetry degradation, deduplication, and atomic replay in tests/integration/test_telemetry_degradation.py
  • T077 [P] [US7] Implement specialized security tests for authorization-header and secret redaction in logs and SDK exceptions (SEC-006) in tests/security/test_secret_redaction.py

Implementation for User Story 7

  • T078 [US7] Implement Langfuse observability integration in src/observability/langfuse_tracer.py managing traces with 8 stable spans (validation, candidate_preparation, hygiene, grounding_validation, ecp_gate, enrichment, rendering, persistence) and generations per LLM attempt; the validation span MUST be buffered or materialized only after successful local validation (no remote Langfuse traffic may occur while terminating local validations are running)
  • T079 [US7] Implement local SQLite queue insertion for telemetry events during Langfuse network outages and atomic flush procedure in src/observability/langfuse_tracer.py
  • T080 [US7] Implement runtime-observable metric emission in src/observability/langfuse_tracer.py using exactly each metric and its dimensions from Doc 05 / FR-067, enforcing cardinality restrictions and emitting prompt_review_signal_total without self-healing or a parallel metrics store (release-wide aggregation of the 11 critical invariants remains the responsibility of T090)
  • T081 [US7] Implement operational telemetry flush command in src/cli/telemetry_flush.py, ensuring events are marked flushed only upon confirmed delivery and telemetry_pending_total returns to zero
  • T082 [US7] Configure the 3 mandatory Langfuse dashboards (Runtime Health, Quality, Future Review Signals) and preserve reproducible setup evidence without creating a parallel metrics system
  • T083 [US7] Integrate observability lifecycle and secret redaction into src/cli/consolidate.py

Checkpoint: Observability is complete, compliant with the metric catalog, and resilient against outages.


Phase 10: User Story 8 - Quality Gates, Zero Regex Verification & Promptfoo Evaluation (Priority: P3)

Goal: Execute comprehensive automated quality gates, Promptfoo offline evaluations, multi-extractor golden regression across 20 reference cases, and zero-regex verification.

Independent Test: Run pytest tests/quality/, pytest evals/, and tests/scripts/check_zero_regex.py, asserting that all 11 critical quality assertions pass without manual intervention.

Tests for User Story 8

  • T084 [P] [US8] Implement automated test in tests/quality/test_zero_regex_enforcement.py executing tests/scripts/check_zero_regex.py across codebase, schemas, and Promptfoo YAML
  • T085 [P] [US8] Implement automated test in tests/quality/test_no_powerful_models.py verifying no runtime module, config, or internal ECP classifier references powerful models
  • T086 [P] [US8] Implement contract parity test in tests/contract/test_contract_parity.py checking all schema versions match 1.0.0
  • T087 [P] [US8] Implement multi-extractor golden-set quality tests across the 20 reference cases in tests/quality/test_golden_reference_20.py
  • T088 [P] [US8] Implement adversarial prompt injection evaluation in tests/quality/test_prompt_injection_guard.py
  • T089 [P] [US8] Implement automated cost budget verification in tests/quality/test_cost_budget.py asserting median per-article cost <= $0.0006

Implementation for User Story 8

  • T091 [US8] Implement CI multi-tier trigger runner in scripts/ci_check.py distinguishing PR gates, prompt/schema/model change gates, and pre-promotion gates per Doc 04 (including reproducible lockfile package build and metadata verification)
  • T092 [US8] Implement a minimal holdout and slice evaluation aggregator in evals/eval_runner.py aggregating Promptfoo output without duplicating prompt execution, computing slice pass rates (≥95%), block precision/recall/F1, metadata accuracy, correct vs unauthorized repairs, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, tag acceptance, and schema validity
  • T093 [US8] Implement empirical latency and cost calibration recorder for staging gates in tests/load/test_load_100_art_per_hour.py
  • T094 [US8] Implement the release packaging script in scripts/build_release_metadata.py generating src/core/release-metadata.json with release_version, runtime_config_sha256, real prompt hashes, schema versions, certified provider/model mappings for both logical roles, and the certified ECP classifier configuration hash bundled inside the distributable package

Checkpoint: Quality gates, security test suites, and CI evaluation infrastructure are verified.


Phase 11: User Story 9 - Production Runbook Operations & Lifecycle Management (Priority: P3)

Goal: Implement operational commands (preflight, smoke, reconcile, telemetry_flush), SQLite native backup/restore, signal handling (SIGTERM/SIGINT), rollback, and credential/model rotations.

Independent Test: Run preflight checks against src/core/release-metadata.json, execute smoke tests with fixtures, perform native SQLite backup and restore, simulate SIGTERM graceful shutdown, and verify credential and model rotation procedures.

Tests for User Story 9

  • T095 [P] [US9] Implement integration tests for operational resilience in tests/integration/test_operations_resilience.py covering native backup/restore, graceful shutdown signals, rollback, credential rotation, certified model rotation, and uncertified rotation rejection
  • T096 [P] [US9] Implement unit tests for preflight verification against release metadata in tests/unit/test_preflight_certification.py
  • T097 [P] [US9] Implement unit tests for smoke test execution in tests/unit/test_smoke_cli.py

Implementation for User Story 9

  • T098 [US9] Implement preflight validation CLI in src/cli/preflight.py checking clock sync, release metadata hashes, schemas, SQLite access, filesystem permissions / atomic rename, minimum disk space, valid/active credentials, certified cheap models (including ECP), trace content policy, and ECP classifier/schema access (remote Langfuse outage does not block preflight)
  • T099 [US9] Implement smoke test CLI in src/cli/smoke.py processing fixture and verifying end-to-end pipeline health (fingerprint, state transitions, LLM call, ECP gate, manifest, Markdown, Langfuse trace, cost/latency within approved staging baseline, and idempotent re-execution)
  • T100 [US9] Implement state and artifact reconciliation CLI in src/cli/reconcile.py (completed states vs files/hashes, final files without state, orphan temp files, pending telemetry, duplicate fingerprints, reconciliation report, safe cleanup without altering editorial content)
  • T101 [US9] Implement full graceful shutdown signal handling (SIGTERM, SIGINT) in src/cli/consolidate.py (stop accepting new units, complete or safely preserve active unit state, close transactions, flush files, attempt telemetry flush, preserve unsent events in SQLite, and exit with coherent exit code)
  • T102 [US9] Validate and update operational procedures and structure in the normative runtime runbook (docs/structured_extraction/06_Runbook_Producao_Runtime.md) for the 11 deployment steps, preflight, smoke, backup/restore, retention, reconciliation, rollback, and credential/model rotations (ready for staging limit incorporation)

Checkpoint: All operational procedures and lifecycle commands are testable and functional.


Phase 12: Polish, Verification Gates & Release Sign-Off

Purpose: Final end-to-end execution, evaluation runs, release evidence report generation, and formal sign-off.

  • T103 [P] Execute multi-parser static policy verification across text-processing runtime modules, content tests/assertions, JSON schemas, and Promptfoo YAML configurations of this feature via python -m tests.scripts.check_zero_regex
  • T104 [P] Execute all 9 contract test suites; the Article Input contract MUST validate all 20 real reference units via pytest tests/contract -v
  • T105 Execute the complete unit and mock integration suites, including idempotency, concurrent replay, and full CLI contract validation (completed_text, rejected_ecp, failed_validation, failed_processing, idempotent result, config/preflight errors with stdout/stderr and exit codes) via pytest tests/unit tests/integration -v
  • T106 Execute all 8 security scenario tests (SEC-001 to SEC-008) via pytest tests/security -v
  • T107 Execute all 10 fault injection scenario tests (FLT-001 to FLT-010) via pytest tests/fault_injection -v
  • T108 Run end-to-end quickstart validation scenarios A, B, C, D per quickstart.md
  • T109 Execute Promptfoo over the 20-case regression set, production golden set, and protected holdout using the exact production prompts and schemas, preserving per-case and per-slice results
  • T110 Execute the production-equivalent 100 articles/hour staging run using the installed lockfile-built package, generate the complete normative staging and release-evidence report, verify all 11 zero-tolerance invariants, verify formal absence of forbidden architectural patterns across this feature's runtime codebase, dependencies/lockfile, prompts, schemas, functional configs, and packaging artifacts (verifying no LangChain, LangGraph, agents, API, internal batch/worker pool, Postgres, external queue, object storage, keyword dictionaries, self-healing, powerful models in runtime, online Promptfoo), and obtain approval for cost/latency/storage/fallback limits
  • T111 Incorporate the approved staging limits into the normative runbook (docs/structured_extraction/06_Runbook_Producao_Runtime.md), update README/CLI documentation, and record formal release sign-off

Dependencies & Execution Order

Phase Dependencies

  • Setup (Phase 1): No dependencies — starts immediately.
  • Foundational (Phase 2): Depends on Setup completion — BLOCKS all user stories.
  • Model Gateway (Phase 3, US6): Depends on Foundational — BLOCKS Hygiene & Enrichment.
  • User Story 1 (Phase 4, P1): Depends on Foundational — delivers initial contract validation (Article + ECP), candidate extraction & idempotency.
  • User Story 2 (Phase 5, P1): Depends on US1 and US6 (Model Gateway) — delivers 10-step LLM extractive hygiene & repairs.
  • User Story 3 (Phase 6, P1): Depends on US2 — delivers mandatory ECP inherence gate and zero-Markdown rejection.
  • User Story 4 (Phase 7, P2): Depends on US3 and US6 (Model Gateway) — delivers post-ECP sentiment & tag enrichment and prompts contract testing.
  • User Story 5 (Phase 8, P2): Depends on US4 — delivers canonical Markdown & atomic persistence across all terminal outcomes.
  • User Story 7 (Phase 9, P3): Depends on US5 & US6 — delivers Langfuse observability & telemetry queue.
  • User Story 8 (Phase 10, P3): Depends on US6 and US7 — delivers quality gates, CI multi-tier evals, golden set & load benchmark.
  • User Story 9 (Phase 11, P3): Depends on US5, US7, and US8 — delivers operational runbooks & resilience commands.
  • Polish & Gates (Phase 12): Depends on all user stories being complete (T110 executes staging and calibrates limits; T111 incorporates limits into documentation and signs off release).

Parallel Execution Opportunities

# Launch Foundational independent tasks in parallel:
Task: "T007 Implement input byte size limiter in src/core/limits.py"
Task: "T008 Implement deterministic fingerprint calculator in src/core/fingerprint.py"
Task: "T009 Implement CandidateObject models in src/candidate/models.py"
Task: "T012 Implement structured JSON logging in src/observability/structured_logger.py"
Task: "T013 Implement atomic writer and manifest generator in src/storage/file_store.py"

# Launch User Story 1 test tasks in parallel:
Task: "T023 Contract test for Article Input in tests/contract/test_article_input_contract.py"
Task: "T024 Contract test for ECP Snapshot in tests/contract/test_ecp_snapshot_contract.py"
Task: "T025 Contract test for Candidates Payload in tests/contract/test_candidates_payload_contract.py"
Task: "T026 Unit tests for input limits in tests/unit/test_input_limits.py"
Task: "T027 Unit tests for fingerprint in tests/unit/test_fingerprint.py"
Task: "T028 Unit tests for candidate extraction in tests/unit/test_candidate_parser.py"
Task: "T029 Unit tests for equivalence mapping in tests/unit/test_equivalence_mapping.py"

Implementation Strategy

Foundation & Incremental Pipeline Flow

  1. Complete Phase 1: Setup
  2. Complete Phase 2: Foundational (blocking prerequisites & shared stores)
  3. Complete Phase 3: User Story 6 (Model Gateway infrastructure & cheap model enforcement)
  4. Complete Phase 4: User Story 1 (Ingestion, Article/ECP contract validation, candidates, idempotency)
  5. Complete Phase 5: User Story 2 (Extractive hygiene & micro-repairs)
  6. Complete Phase 6: User Story 3 (Mandatory ECP gate & relevance enforcement)
  7. Complete Phase 7: User Story 4 (Post-ECP sentiment & native tags enrichment, prompt contracts)
  8. Complete Phase 8: User Story 5 (Canonical Markdown & atomic persistence across all terminal outcomes)
  9. Complete Phase 9: User Story 7 (Direct Langfuse observability & telemetry queue)
  10. Complete Phase 10: User Story 8 (Automated quality evaluation, golden set & CI gates)
  11. Complete Phase 11: User Story 9 (Production runbook operations & lifecycle management)
  12. Complete Phase 12: Polish, Verification Gates, Staging Calibration & Release Sign-Off