86 KiB
Feature Specification: Article Consolidation and Hygiene Runtime
Feature Branch: 006-article-consolidation-runtime
Created: 2026-08-23
Status: Draft
Input: User description: "baseado em todos esses arquivos docs/structured_extraction (01_PRD_Runtime_Consolidacao_Artigos.md, 02_Arquitetura_Runtime_Consolidacao_Artigos.md, 03_ADRs_Runtime_Consolidacao_Artigos.md, 04_Plano_Testes_Evals_Runtime.md, 05_Metricas_KPIs_Runtime.md, 06_Runbook_Producao_Runtime.md, 07_Especificacao_Prompt_Contexto_Harness_Runtime.md). Nao deve ser negligenciado nada! Tem que implementar 100% do que esta previsto nessas documentacoes, nao pode extrapolar em nada!"
Priority & Governance Statement
Important
Priority Definitions: In this specification, priority labels (P1, P2, P3) denote the sequential order of module implementation and testing. Every User Story (US1 through US9), functional requirement (FR-001 to FR-084), metric, operational procedure, and quality gate in this document is strictly mandatory for production go-live. No requirement is optional.
User Scenarios & Testing (mandatory)
User Story 1 - Single Article Ingestion, Contract Validation, and Deterministic Candidate Preparation (Priority: P1)
As a pipeline orchestrator, I want to submit a single extracted news article unit (containing structural crawl fields, extractions from Trafilatura, Newspaper4k, and Readability, and a pre-calculated selected_extractor) alongside a versioned canonical ECP snapshot, so that the runtime validates input contracts locally before any remote call, establishes a deterministic idempotency fingerprint over canonical serialization, preserves unknown fields, and structures candidate elements using standard parsers without regular expressions or keyword dictionaries.
Why this priority: Foundational entry point of the pipeline. Enforces strict input validation, protects against premature remote calls, and establishes candidate provenance.
Independent Test: Can be tested by providing individual article JSON objects and ECP snapshots (valid, corrupted, missing fields, or batch wrappers), verifying schema validation, deterministic fingerprint generation over canonical serialization, SQLite state initialization (received, validated), candidate ID assignment (opaque, without quality judgment, stable within execution), and immediate failure with explicit error codes before any remote call when preconditions fail.
Acceptance Scenarios:
- Given a valid single-article JSON input containing structural fields (
crawled_url,error_message,extraction_status,http_status,input_meta,page_title,selected_extractor,trafilatura,newspaper4k,readability), a usableselected_extractor, at least one valid HTTP/HTTPS source URL, at least one non-empty candidate title, at least one processable text block, and a valid ECP snapshot matching the canonical ECP schema, When ingestion executes, Then the runtime validates all contracts locally without making remote provider, remote Langfuse, or ECP classifier calls, computes a deterministic content hash over canonical serialization, recordsreceivedandvalidatedstates in SQLite, preserves unknown fields in recorded input while ignoring them in processing, and extracts structural candidates using DOM, CommonMark AST, JSON parser, URL parser, Unicode normalizers, and multilingual tokenizers/segmenters without any regular expressions. - Given an invalid input payload (a batch JSON containing a root
articlesarray, unparseable JSON, missing or unrecognizedselected_extractor,selected_extractorpointing to an extraction with no usable content, missing source URL, missing title candidate, or missing textual content), When validation executes, Then the runtime terminates immediately prior to any remote call, returns a structured JSON error result, and records the specific error code:INVALID_ARTICLE_SCHEMAfor malformed JSON or batch payload;MISSING_SELECTED_EXTRACTORwhenselected_extractoris absent;INVALID_SELECTED_EXTRACTORwhenselected_extractoris unrecognized;SELECTED_EXTRACTOR_UNAVAILABLEwhen the selected extractor has no usable content;MISSING_SOURCE_URLwhen no valid source URL can be resolved;MISSING_TITLE_CANDIDATEwhen no valid candidate title exists;MISSING_CONTENTwhen no processable text content exists. The runtime MUST NOT recalculate or silently substituteselected_extractor.
- Given an absent, unparseable, or schema-incompatible ECP snapshot, When validation executes, Then the runtime terminates immediately with
INVALID_ECP_SCHEMAbefore candidate preparation or any remote invocation. - Given an input whose deterministic fingerprint matches a previously completed execution under identical functional versions and configurations, When ingestion executes, Then the runtime recognizes the completed state and returns the existing persisted result and artifacts without redundant LLM invocations.
- Given deterministic metadata resolution, When metadata candidates are resolved, Then:
- Source URL is resolved strictly in priority order (
crawled_url→input_meta.url→ canonical URL ofselected_extractor→newspaper4k.canonical_link→trafilatura.canonical_url) and is never chosen or modified by the LLM. - Published Date is resolved via date parser and normalized to ISO 8601, prioritizing source consensus or fallback priority (
newspaper4k.publish_date→input_meta.quando_publicado→trafilatura.date), omittingpublished_atif invalid, and is never chosen or modified by the LLM. - Candidate Title sources include
input_meta.titulo,page_title,trafilatura.title,newspaper4k.title,readability.title,readability.short_title. - Candidate Subtitle sources include
input_meta.subtitulo,trafilatura.description,newspaper4k.meta_description. - Candidate Author sources include
trafilatura.author, items ofnewspaper4k.authors(preserving structured list items and order without regex/delimiter splitting),readability.author. - Candidate Types include title, subtitle, author, date, block, heading, list, quote, link, and image.
- Candidate IDs are opaque, contain no quality judgment, are unique within the execution, and duplicate IDs result in internal construction failure.
- Cross-Extractor Equivalence is evaluated via Unicode normalization, whitespace normalization by library, tokenization, and sequence similarity without regex or keyword dictionaries. Low similarity retains distinct candidate IDs.
- Language is detected using an appropriate multilingual NLP library.
- Structural Backbone:
selected_extractordefines the base ordering backbone; the LLM is not forced to select all its blocks, and blocks exclusive to secondary extractors enter only via explicit LLM selection and grounding validation. - HTML/JSON-LD Parsing: Malformed HTML is parsed tolerantly into a safe DOM or controlled failure; valid JSON-LD
Article/NewsArticleyields structural metadata; invalid JSON-LD is ignored with a warning logged.
- Source URL is resolved strictly in priority order (
User Story 2 - Mandatory LLM Extractive Content Hygiene and Controlled Text Repairs (Priority: P1)
As an editorial consumer, I want every article to undergo LLM extractive hygiene to select genuine editorial candidates and eliminate noise (advertisements, cross-promotions, player chrome, navigation, duplicate snippets, and newsletters), while permitting only auditable, bounded micro-repairs for unmistakable encoding and typographical flaws.
Why this priority: Eliminates editorial noise and guarantees 100% grounded content selection without autonomous rewriting.
Independent Test: Can be tested with single articles exhibiting consensus or divergence across extractors, containing textual ads, navigation links, and controlled encoding flaws, verifying that the LLM returns only candidate IDs and explicit repair operations, the harness validates grounding, invalid repairs are discarded with originals preserved, and ungrounded responses trigger fallback.
Acceptance Scenarios:
- Given an article where all three extractors agree 100% on content, When content hygiene executes, Then the LLM hygiene step is executed unconditionally (consensus improves evidence but never bypasses LLM hygiene).
- Given the
article_content_hygieneprompt and candidate payload, When the LLM responds, Then the response conforms strictly to the schema returningtitle_candidate_id,subtitle_candidate_id(or null),author_candidate_id(or null),kept_block_ids(in order),kept_link_ids,kept_image_ids,repairs, and categorical removal reasons (when enabled for observability), with an absolute absence of free-form body text or free Markdown fields. - Given the hygiene harness validation pipeline, When an LLM hygiene response is evaluated, Then the harness executes the strict 10-step sequence:
- Parse JSON;
- Validate schema;
- Validate metadata IDs and their expected types;
- Validate block IDs exist in candidate storage;
- Validate ordering compatibility with canonical representation;
- Validate links and images against input candidates;
- Validate proposed repairs individually;
- Assemble intermediate structure by retrieving candidate text from internal maps;
- Validate grounding of assembled Markdown;
- Validate minimum content requirements (rejecting selections that omit material editorial content).
- Given an LLM hygiene response, When the harness validates the response, Then an ungrounded candidate ID, ungrounded link/image URL, or ungrounded content triggers a
GROUNDING_VIOLATION(invalidating the entire response) and routes to fallback; an invalid schema triggers semantic fallback without being labeled as a grounding violation; and if all fallback options are exhausted, processing terminates withHYGIENE_FAILED. - Given proposed text repairs within closed allowable categories (
encoding,unicode,spacing,punctuation_corruption,obvious_typo) containing target ID, exact original fragment, replacement fragment, category, and rationale, When the harness validates the repair, Then it normalizes and tokenizes original and replacement with Unicode/NLP libraries without regex, computes diffs without regex, verifies closed category, and records original, replacement, decision, and rationale. - Given a proposed repair targeting sensitive entities (names, numbers, dates, scores, quotes, or facts), When the harness validates the repair, Then the repair is rejected (
INVALID_TEXT_REPAIR) UNLESS the difference is strictly caused by an unmistakable, verifiable encoding/Unicode defect (such as mojibake in a proper name). If in doubt, the original text is preserved. - Given a proposed repair that attempts stylistic improvement, synonym replacement, paraphrasing, tone change, or whose original fragment is ambiguous or missing, When the harness validates the repair, Then the invalid repair is discarded (
INVALID_TEXT_REPAIR), the exact original candidate text is preserved, and processing continues without invalidating the rest of the valid selection. - Given editorial assembly of intermediate Markdown, When the assembler constructs the document, Then it follows the
selected_extractorbackbone, resolves candidate content from internal maps, applies validated repairs, eliminates exact structural title/subtitle duplicates, grounds image positions structurally with alt/caption derived strictly from input text (images without editorial position are omitted), verifies link URLs and anchor text exist in input (rejecting malformed links), and formats Markdown without underline or inline HTML. - Given editorial rules, When hygiene executes, Then the system strictly enforces: no translation, no summarization, no narrative reorganization, no transition creation, no information completion, no factual correction, no linguistic variant changes, no repetition of author/date/sentiment/tags/ECP in the body. The LLM never controls the renderer or filesystem.
User Story 3 - Mandatory ECP Gate and Relevance Enforcement (Priority: P1)
As a content governance stakeholder, I want the intermediate sanitized Markdown to be evaluated by the ECP classifier adapter before any final editorial output is generated, ensuring that only articles with direct or contextual inherent relevance are permitted to produce published Markdown.
Why this priority: Enforces business relevance and editorial boundary constraints; prevents non-inherent content from being published while maintaining an audit trail.
Independent Test: Can be tested by passing intermediate sanitized Markdown documents to the ECP classifier adapter with associated ECP profiles across all four classification outcomes, verifying that only DIRECT_INHERENT and CONTEXTUAL_INHERENT progress to enrichment and Markdown output, while TANGENTIAL and NOT_RELATED generate a persisted rejected_ecp manifest and zero Markdown files.
Acceptance Scenarios:
- Given intermediate sanitized Markdown submitted to the ECP adapter, When the ECP classifier returns
DIRECT_INHERENTorCONTEXTUAL_INHERENT, Then the state machine transitions fromcontent_cleanedtoecp_approvedand proceeds to enrichment. - Given intermediate sanitized Markdown submitted to the ECP adapter, When the ECP classifier returns
TANGENTIALorNOT_RELATED, Then the state machine transitions fromcontent_cleanedtoecp_rejected(a terminal state), writes a persistent<fingerprint>.result.jsonmanifest with statusrejected_ecp(ECP_REJECTED) and classification metadata, and guarantees no Markdown file is created. - Given intermediate sanitized Markdown submitted to the ECP adapter, When the ECP adapter validates the classifier output, Then it requires category,
is_inherent, confidence, rationale, and evidences, and verifies that all evidence fragments belong to the intermediate sanitized Markdown. - Given an ECP classifier technical failure, invalid enum, or evidence fragment not present in the intermediate document, When the adapter processes the output, Then the execution terminates with
ECP_CLASSIFICATION_FAILED, transitions tofailed, and outputs no Markdown file. - Given any LLM tier utilized by the ECP classifier, When the adapter executes, Then the adapter verifies that only certified cheap models are configured, records the generation telemetry, and rejects any configuration specifying uncertified powerful models.
User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2)
As a content consumer, I want ECP-approved articles to receive structured metadata enrichment consisting of entity-relative sentiment (positive, negative, neutral) and 3 to 8 topic tags in the article's native language, strictly decoupled from the editorial body text.
Why this priority: Required for mandatory front matter metadata; must remain isolated from body text to prevent content modification.
Independent Test: Can be tested by providing approved intermediate Markdown, language, and minimal ECP entity identity to the article_sentiment_tags prompt, verifying that sentiment is evaluated strictly relative to the ECP entity, tags are in the article's language and bounded between 3 and 8 with evidence IDs, duplicate tags are rejected via NLP libraries, and the body text is not modified.
Acceptance Scenarios:
- Given an approved article and target ECP entity, When the
article_sentiment_tagsprompt executes, Then it returns sentiment (positive,negative, orneutralevaluated strictly relative to the ECP entity), between 3 and 8 unique tags in the article's language supported by textual evidence, and evidence candidate IDs, without returning or altering body text. - Given an enrichment output, When the harness validates the response, Then it verifies tag count (3 to 8), validates tag uniqueness using Unicode/NLP libraries without regex, verifies all evidence IDs against the intermediate document, and rejects responses containing duplicate tags, ungrounded tags, or attempts to output/modify body content.
- Given an invalid enrichment response on
runtime_primary, When the harness handles the failure, Then it routes toruntime_fallback. - Given a scenario where both primary and fallback enrichment calls fail, When the stage concludes, Then processing terminates with status
failed_processing(ENRICHMENT_FAILED), transitions tofailed, and no Markdown file is created (sentiment and tags are mandatory front matter fields).
User Story 5 - Canonical Markdown Rendering and Atomic Persistence (Priority: P2)
As a system integrator, I want the runtime to render a standardized Markdown document with structured YAML front matter and an accompanying machine-readable JSON result manifest, persisted atomically using fingerprint-based filenames and tracked in SQLite, so that downstream consumers never encounter partial, corrupt, or inconsistent artifacts.
Why this priority: Guarantees file-system atomicity, data integrity, and deterministic contract adherence under normal and failure conditions.
Independent Test: Can be tested by verifying generated .result.json manifests and .md files, validating YAML front matter structure, verifying atomic rename lifecycles, content hash verification, and SQLite state consistency across normal completions, concurrent executions, and interrupted writes.
Acceptance Scenarios:
- Given an approved, enriched article, When final Markdown rendering executes, Then it produces a document with:
- Mandatory YAML front matter:
title,source_url,sentiment,tags,ecp_qid,ecp_canonical_name,ecp_category,ecp_confidence. - Optional YAML front matter (omitted when empty/absent):
subtitle,author,published_at. - Body Structure: H1
# Title, italic*Subtitle*(when present, with NO artificial blank line generated if absent), and editorial content in canonical order with grounded links and images, omitting author, date, sentiment, tags, and ECP data from the body text.
- Mandatory YAML front matter:
- Given output artifacts (
<fingerprint>.result.jsonand optional<fingerprint>.md), When persistence executes, Then artifacts are written to temporary files in the same filesystem, flushed, closed, verified against content hashes, atomically renamed to their final destination, and the SQLite state is persisted in the same logical completion unit, without exposing an inconsistent completed state. Manifest and Markdown are never presented as completed while the pair is inconsistent. - Given any execution (success, validation failure, ECP rejection, or processing failure), When persistence completes, Then:
- A structured JSON result is returned.
- When a deterministic fingerprint exists,
<fingerprint>.result.jsonis persisted containing:fingerprint: deterministic hash of the execution;source_url: resolved source URL;selected_extractor: extractor received;final_status:completed_text,rejected_ecp,failed_validation, orfailed_processing;generate_markdown: boolean decision indicating if Markdown was produced;markdown_path: file path of Markdown output, or null;ecp_classification: category and confidence, or null if pre-ECP failure;provider_versions: configured provider identifiers and versions, or null if pre-LLM failure;model_versions: configured model versions, or null if pre-LLM failure;prompt_versions: prompt names, versions, and hashes, or null if pre-LLM failure;config_version: runtime functional configuration version;trace_id: Langfuse trace identifier, or null if local failure;error_codes: specific error or rejection codes (INVALID_ARTICLE_SCHEMA,INVALID_ECP_SCHEMA,MISSING_SELECTED_EXTRACTOR,INVALID_SELECTED_EXTRACTOR,SELECTED_EXTRACTOR_UNAVAILABLE,MISSING_SOURCE_URL,MISSING_TITLE_CANDIDATE,MISSING_CONTENT,HYGIENE_FAILED,GROUNDING_VIOLATION,INVALID_TEXT_REPAIR,ECP_CLASSIFICATION_FAILED,ECP_REJECTED,ENRICHMENT_FAILED,PERSISTENCE_FAILED,TELEMETRY_PENDING).
- Given a storage write failure or hash mismatch, When persistence fails, Then processing terminates with
PERSISTENCE_FAILEDand leaves no incomplete final output. - Given an interrupted execution where Markdown was written to disk but SQLite state was not updated before interruption, When recovery or replay executes, Then hash-based reconciliation identifies the existing valid output file, updates the SQLite state to
completed_text, and avoids duplicate calls or duplicate output creation. - Given two identical concurrent executions with the same fingerprint, When both run simultaneously, Then exactly one execution completes the full flow and the other execution safely resumes or reuses the persisted result, resulting in zero duplicate published outputs.
User Story 6 - Agnostic Model Gateway, Cheap Models, and Controlled Fallbacks (Priority: P2)
As a cloud operations manager, I want all runtime LLM calls to be managed by an agnostic Model Gateway configured exclusively with certified cheap models (logical roles runtime_primary and runtime_fallback), supporting limited technical retries for transient errors, immediate fallback on schema/grounding failures, and conservative deterministic fallback, so that operational costs remain low and powerful models are never invoked.
Why this priority: Enforces strict cost discipline and architectural decoupling from specific LLM providers.
Independent Test: Can be tested by configuring primary and fallback providers, simulating transient network errors (timeouts, connection resets, HTTP 429, HTTP 5xx, empty responses) and semantic failures (schema violations, grounding failures), verifying technical retries with backoff, fallback routing, deterministic hygiene fallback, and preflight rejection of uncertified/powerful models.
Acceptance Scenarios:
- Given the Model Gateway interface, When invoked by the runtime, Then it accepts logical role, messages/context, structured schema, timeout, and trace metadata, and returns structured output, effective provider, effective model, tokens (input, output, cached), cost, latency, technical status, attempt number, and fallback indication.
- Given certified role configuration, When mapped in the gateway, Then each role maps to provider, model, parameters, compatible prompt, expected schema, timeout, and version.
- Given a transient technical error (timeout, connection interruption / connection reset, HTTP 429 with backoff up to configured limit, HTTP 5xx, empty response due to technical failure), When the Model Gateway executes a call, Then it performs limited technical retries on the same provider according to certified role limits.
- Given a semantic failure (invalid JSON schema, grounding violation) or exhausted technical retries on
runtime_primary, When the gateway handles the failure, Then it transitions immediately toruntime_fallbackwithout retrying semantic errors on the same model and without entering semantic retry loops. - Given the gateway architecture, When routing calls, Then the gateway uses two adapters/configurations (
runtime_primaryandruntime_fallback) without implementing smart routers, autonomous model selectors, or dynamic model scoring. - Given a failure of both
runtime_primaryandruntime_fallbackduring hygiene, When fallback evaluates, Then the runtime applies a conservative deterministic fallback based on theselected_extractorbackbone, which eliminates only structurally invalid elements, does NOT use regex, does NOT attempt semantic ad/recommendation filtering, and fails withHYGIENE_FAILEDif grounding or minimum content requirements cannot be guaranteed. - Given any gateway configuration specifying a model uncertified for the runtime role (including any powerful model tier), When preflight validation runs, Then the runtime fails preflight checks and blocks execution.
User Story 7 - Production Observability, Log Sanitization, and Deferred Telemetry Resend (Priority: P3)
As an SRE, I want every validated execution to record structured JSON logs and end-to-end tracing in Langfuse (with stable spans covering validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, and persistence, and generations capturing token counts, latency, calculated costs, and versions), with automatic redaction of secrets and graceful degradation to local SQLite queuing (TELEMETRY_PENDING) during observability outages, so that telemetry is complete and operations remain resilient.
Why this priority: Production traceability, cost tracking, and incident diagnosis without risking article pipeline blockage.
Independent Test: Can be tested by running executions with Langfuse available and unavailable, checking structured JSON logs for sanitized fields and correct event codes, inspecting Langfuse spans and generations, and executing the operational telemetry resend routine to flush pending records from SQLite.
Acceptance Scenarios:
- Given a validated execution, When telemetry is recorded, Then a Langfuse trace is created under a stable
run_idwith stable spans covering the normative stages (validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, persistence) devoid of URLs, model names, or dynamic IDs, and each LLM attempt is recorded as a separate generation capturing prompt versions/hashes, model/provider, schema version, cached tokens (when available), token counts, calculated cost, latency, timeout status, attempts, technical status, applicable scores, schema results, fallback status, logical role, normalized context sent, structured response, grounding validation result, applied repairs, and rejected repairs (subject to trace content policy), without duplicating raw HTML or full JSON payloads. - Given log outputs, Langfuse traces, and manifest files, When payloads are generated, Then all API keys, authorization headers, and environment secrets are redacted (
trace_redaction_failure_total= 0). Full ECP, full HTML, and full article text are omitted from standard logs. Repair diffs reside in controlled traces only, never in metric labels. - Given a network failure or outage reaching Langfuse, When an article is processed, Then the article processing completes normally, a
TELEMETRY_PENDINGrecord is saved in SQLite, and execution exits cleanly without stalling. - Given pending telemetry records in SQLite, When the operational telemetry resend procedure is executed, Then pending events are delivered to Langfuse, deduplicated by event ID, and
telemetry_pending_totalreturns to zero. - Given metrics collection, When metrics are recorded, Then high-cardinality values (URLs, fingerprints, run IDs, trace IDs, titles, authors, full text, free tags) are strictly forbidden as metric labels and restricted to traces and logs.
- Given runtime operations, When prompt review signals occur (schema failures, grounding violations, rejected repairs, fallback invocations, terminal failures), Then the runtime emits
prompt_review_signal_totalmetrics capturinglogical_call,prompt_version,provider,model,language,source_domain_group, andreason, without initiating any autonomous self-healing. - Given operations dashboards, When monitored, Then the system provides:
- Runtime Health Dashboard: received, completed, rejected, failed, throughput, latency, cost, providers, fallback, persistence, pending telemetry.
- Quality Dashboard: schema, grounding, repairs, ECP, sentiment, tags, and results by prompt/model/language/domain.
- Future Review Signals Dashboard:
prompt_review_signal_total, failures after fallback, concentration by prompt version, and associated potential cost.
User Story 8 - Automated Quality Evaluation, Staging Baselines, and CI Quality Gates (Priority: P3)
As a release engineer, I want the CI pipeline and staging environments to execute AST static analysis for regex prohibition in text pipeline modules and assertions, Promptfoo evaluations executed in development/CI over the reference fixture (20 cases) and golden dataset (with holdout), fault injection, and 100 articles/hour load tests, so that cost/latency SLOs are empirically calibrated and zero-tolerance invariants prevent flawed releases.
Why this priority: Guarantees production readiness, validates multilingual performance across slices, and prevents architectural degradation.
Independent Test: Can be tested by running AST static checks on text pipeline modules and assertions, executing Promptfoo in dev/CI against prompt files without online runtime coupling, simulating fault injection (primary down, both down, Langfuse unavailable, SQLite lock, disk full, process crash during write, ECP down, truncated LLM response, orphan temp, telemetry flush failure), running 100 articles/hour load tests, and verifying critical invariant gates.
Acceptance Scenarios:
- Given text processing pipeline modules and Promptfoo assertion configurations, When AST static analysis runs in CI, Then it validates that no Python
remodule imports or regex functions are called in those modules/assertions, failing the build if any are detected. - Given Promptfoo evaluation suites executed in development/CI, When evals execute, Then Promptfoo loads the identical versioned prompt files used in production, evaluating JSON schemas, candidate ID grounding, and repair diffs without using regex assertions, semantic keyword dictionaries, or LLM-as-a-judge for grounding.
- Given critical release gates, When a build is evaluated, Then promotion is blocked if any of the 11 critical metrics is greater than zero:
ungrounded_text_total,ungrounded_url_total,ungrounded_image_total,unauthorized_rewrite_total,critical_fact_change_total,duplicate_output_total,lost_article_total,secret_exposure_total,powerful_runtime_model_call_total,online_promptfoo_call_total,text_regex_usage_total. - Given quality evaluations across golden dataset and holdout, When metrics are computed, Then pass rate is evaluated per slice (language, domain, extractor, prompt version, model version), requiring at least 95% pass rate per slice without allowing a global average to mask localized failures, and measuring precision/recall/F1 of kept blocks, metadata accuracy, correct vs unauthorized repair rates, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, and tag acceptance.
- Given staging load testing, When 100 articles per hour are processed under realistic concurrency using the identical SQLite and output directory planned for production, with a representative mix of languages/domains/extractors/structures and realistic fallback rates, Then the run completes with zero lost articles, zero duplicate outputs, zero partial files exposed, zero database corruption, stable memory/disk usage, 100% traces delivered or preserved as pending telemetry, and produces empirical p50/p95/p99 latency and cost baselines for operational approval prior to go-live.
- Given the CI/CD pipeline, When changes are integrated, Then execution follows the minimum sequence:
- Format validation;
- Lint & static analysis;
- AST regex check;
- Unit tests;
- Contract tests;
- Simulated integration tests (with simulated/mock providers);
- Reduced Promptfoo eval;
- Package build;
- Full golden set eval (pre-promotion, using authorized offline providers);
- Staging 100 art/h load test;
- Cost/latency limits approval;
- Controlled promotion.
User Story 9 - Production Runbook Operations and Lifecycle Management (Priority: P3)
As a production operations engineer, I want standardized operational routines for preflight checks, smoke tests, consistent SQLite backups/restores, graceful shutdown, reconciliation, manual rollbacks, credential rotations, and certified model/provider rotations, so that production can be reliably maintained, diagnosed, and recovered without manual file tampering.
Why this priority: Production operational readiness requirement ensuring all failure modes, deployments, and rollbacks have verified, auditable procedures.
Independent Test: Can be tested by executing preflight validation, running smoke tests against versioned fixtures, performing consistent SQLite backup/restore cycles, triggering graceful shutdown signals, running reconciliation reports, executing manual rollbacks to previous certified configurations, and testing credential rotations.
Acceptance Scenarios:
- Given a new deployment or environment startup, When preflight validation runs, Then it validates system clock synchronization, prompts and configs belonging strictly to the same release, local configuration reading, prompt file existence & hashes, schema compatibility, SQLite access, filesystem atomic write permissions & directory permissions, minimum disk space, validated active credentials (not just presence), absence of powerful models in runtime roles, Langfuse local configuration (environment, content policy), and ECP classifier/schema access before accepting live traffic. Remote Langfuse connectivity failure does NOT block preflight.
- Given a preflight pass, When smoke testing executes, Then it processes a versioned reference fixture, verifying fingerprint generation, state transitions, LLM call, ECP gate, manifest, Markdown output, Langfuse trace, cost and latency within approved staging ranges, and idempotent re-execution.
- Given deployment procedures, When a release is deployed, Then it follows the complete 11-step sequence:
- Pause new executions in orchestrator;
- Await or gracefully drain existing executions;
- Preserve consistent backup of SQLite and active configuration;
- Deploy package, prompts, and schemas of the release;
- Test state migration on a copy before applying to production;
- Run preflight checks;
- Run smoke test with approved fixture;
- Confirm manifest, Markdown, state, and trace;
- Release with reduced concurrency;
- Verify errors, fallback rate, cost, and latency;
- Release full volume.
- Given operational lifecycle procedures, When maintenance tasks execute, Then:
- Responsibilities Matrix: Operational roles reproduce the normative matrix:
- Orchestrator: provide article/ECP, control concurrency, and consume the manifest.
- Operations: deploy, monitor, recover, and execute rollback.
- Engineering: correct code, prompts, schemas, or integrations through the normal release process.
- Curator/Eval: maintain the golden set and approve quality.
- Backup: SQLite is backed up using consistent database backup mechanisms (not raw file copies during active writes).
- Shutdown: Upon receiving a shutdown signal, the runtime stops accepting new units, completes or persists the active unit safely, closes open transactions, flushes open files/telemetry, preserves pending telemetry, and exits cleanly with a coherent status.
- Reconciliation: Periodic routines detect and report discrepancies between SQLite states, manifest files, Markdown files, orphan temp files, and pending telemetry without manual editing.
- Rollback: Triggered by critical invariant violations, contract incompatibilities, or unmitigated failures, manual rollback restores the previous certified package/configuration/prompt versions and reprocesses affected units.
- Rotation: Model/provider changes follow certification via Promptfoo, golden set evals, and staging baseline calibration before promotion, maintaining the previous configuration for rollback.
- Credentials: Credential rotations create least-privilege credentials, update environment secrets, run preflight/smoke tests, verify absence of exposure, revoke old credentials, and avoid altering functional fingerprints when only operational secrets change.
- Incidents: Failures in providers, ECP, enrichment, Langfuse, SQLite, disk, grounding, cost, latency, or input validation / producer schema mismatches (checking producer version, comparing with release schema, confirming unit payload, never calling LLM manually, fixing producer or contract via standard release), preserving incident evidence and creating mandatory regression test cases after critical incidents. Production prompts, sentiments, tags, Markdown files, manifests, and SQLite MUST NEVER be edited manually. Reprocessing MUST locate state/fingerprint, verify versions, reuse existing outputs or resume incomplete states under the same configuration, generate distinct fingerprints for new configurations, avoid manual file edits, and never reprocess articles merely to recreate traces.
- Responsibilities Matrix: Operational roles reproduce the normative matrix:
Edge Cases
- Batch Array Payload: Input containing a root
articlesarray is rejected immediately with error codeINVALID_ARTICLE_SCHEMA. - Selected Extractor Handling: If
selected_extractoris absent, invalid, or lacks usable content, execution terminates withMISSING_SELECTED_EXTRACTOR,INVALID_SELECTED_EXTRACTOR, orSELECTED_EXTRACTOR_UNAVAILABLE. The runtime never recalculates or silently substitutes the extractor. - Extractor Full Consensus: LLM hygiene executes unconditionally even when all three extractors agree 100%.
- Textual Noise: The LLM hygiene step removes textual ads, cross-promotions, player chrome, navigation, duplicate snippets, and newsletters using semantic understanding, without relying on hardcoded keyword lists per language.
- Unauthorized Textual Repairs: Repairs attempting semantic paraphrasing, style improvements, synonym replacement, or alterations to facts, names, dates, numbers, scores, or quotes are rejected by the harness (
INVALID_TEXT_REPAIR); the original candidate text is preserved, UNLESS the difference is strictly an unmistakable, verifiable encoding/Unicode defect (such as mojibake in a proper name). - Ambiguous or Non-Existent Repair Target: Repairs with ungrounded target IDs or ambiguous original fragments are rejected; original text is preserved.
- ECP Rejection (
TANGENTIALorNOT_RELATED): The runtime terminates editorial processing cleanly, persists a<fingerprint>.result.jsonmanifest with statusrejected_ecp(ECP_REJECTED), and produces NO Markdown file. - ECP Outage or Invalid Output: Terminate with
ECP_CLASSIFICATION_FAILEDand generate no Markdown. - Primary & Fallback Model Outage: If both primary and fallback LLMs fail during hygiene, conservative deterministic fallback is used only if baseline structural integrity is satisfied; otherwise,
HYGIENE_FAILEDis returned. If both fail during enrichment,ENRICHMENT_FAILEDis returned and no Markdown is output. - Prompt Injection in Article Content: Article text and metadata are strictly treated as untrusted data candidates. Structural parsing, candidate ID referencing, and output schemas prevent prompt injection from executing instructions.
- Observability Outage: Langfuse network failures do not halt processing; telemetry records are stored in SQLite as
TELEMETRY_PENDINGfor deferred resending. - SQLite Concurrency & Lock Wait: Concurrent CLI executions on the same SQLite state store use WAL mode, busy timeout, and short transactions to prevent lock corruption under 100 articles/hour load.
- Process Termination During Write: Staged writing to temporary files and atomic rename prevent partial files from being exposed as final outputs.
- Interrupted Write Reconciliation: Hash-based reconciliation resolves crashes between Markdown write and SQLite final state update without duplicate processing.
- Future Self-Healing Boundary: The runtime exclusively logs telemetry, versions, and review signals (
prompt_review_signal_total). It does NOT contain prompt optimizers, judges, candidate generation, canaries, or automatic rollback mechanisms.
Requirements (mandatory)
Functional Requirements
Master Simplicity and Architecture Principles
- FR-001: System MUST satisfy all production requirements using the minimum necessary code, abstractions, dependencies, and components. No complexity MAY be added without satisfying an explicit requirement, documented risk, or proven operational need. Small responsibilities MAY share modules; the architectural component list does NOT mandate a separate class or package per item. No anticipatory implementation of future components (such as self-healing) is permitted.
- FR-002: System MUST adhere to the strict dependency preference hierarchy: 1. standard library; 2. existing monorepo dependencies; 3. consolidated libraries eliminating significant custom implementation; 4. custom code for product-specific rules only. Every new dependency MUST document its requirement, stdlib alternative, security impact, maintenance impact, license, size impact, and startup impact. Agent frameworks, libraries for trivial single functions, secondary ECP schemas, manual HTML/Markdown/URL parsers, and anticipatory self-healing dependencies are strictly prohibited.
- FR-003: System MUST execute as an ephemeral Python CLI processing exactly one article unit per invocation, without providing an API, without internal batch loops, and without internal worker pools. Concurrency is managed externally by the orchestrator, sharing the SQLite state store and filesystem safely.
- FR-004: System MUST maintain independent semantic versions for all system contracts: input article contract, ECP Snapshot reference, runtime configuration, candidates payload, hygiene response schema, repair operations schema, enrichment response schema, output manifest schema, and prompts. Any incompatible change in any contract MUST require a new version and complete evaluation.
Input, Validation, and Fingerprinting Contracts
- FR-005: System MUST require a valid ECP snapshot matching the canonical ECP schema per execution, failing immediately with
INVALID_ECP_SCHEMAbefore any remote call if absent or invalid. The runtime MUST NOT duplicate or redefine the ECP schema. - FR-006: System MUST require and validate
selected_extractorto be one oftrafilatura,newspaper4k, orreadability, and MUST NOT calculate, recalculate, or silently substitute the selected extractor. - FR-007: System MUST validate minimum editorial validity locally before any remote call (provider, remote Langfuse, or ECP classifier), failing with specific error codes:
MISSING_SOURCE_URLif no valid HTTP/HTTPS source URL exists;MISSING_TITLE_CANDIDATEif no non-empty candidate title exists;MISSING_CONTENTif no processable text block exists;SELECTED_EXTRACTOR_UNAVAILABLEif the selected extractor lacks usable content.
- FR-008: System MUST compute a deterministic hash over the canonical serialization of article identity, extraction content hashes,
selected_extractor, ECP snapshot identity/version, prompt versions, functional runtime configuration, and configured model identifiers (excluding runtime timestamps and trace IDs). - FR-009: System MUST enforce idempotency via SQLite state store, returning existing completed outputs without re-running LLMs when an identical fingerprint and configuration are submitted. When identical concurrent executions occur, exactly one completes effectively and the other safely resumes or reuses the persisted result, ensuring zero duplicate published outputs.
- FR-010: System MUST preserve unknown fields in recorded input while ignoring them during pipeline processing.
- FR-011: System MUST consume available structural fields (
crawled_url,error_message,extraction_status,http_status,input_meta,page_title,selected_extractor) and extraction fields:- Trafilatura:
title,author,date,description,text,markdown,canonical_url,image,language,sitename,categories,tags,raw_json,pagetype,error. - Newspaper4k:
title,authors,publish_date,meta_description,text,article_html,canonical_link,top_image,images,meta_data,meta_lang,meta_site_name,tags,keywords,error. - Readability:
title,short_title,author,cleaned_text,cleaned_html,error.
- Trafilatura:
- FR-012: System MUST NOT perform any remote calls (provider, remote Langfuse, or ECP classifier) before completing all local validations capable of terminating execution.
- FR-013: System MUST load environment configuration containing paths (input, output, state), provider endpoints/credentials, role configurations (
runtime_primary,runtime_fallback), timeouts, retry limits, prompt versions, Langfuse config, trace content policy, and state store concurrency limits. Versioned functional configs MUST enter the fingerprint; operational secrets MUST NOT. System MUST support reproducible packaging and builds.
Deterministic Structural Candidate Model
- FR-014: System MUST parse HTML (DOM parser), Markdown (CommonMark AST), JSON-LD (JSON parser), URLs (URL parser), and text (Unicode normalization, multilingual tokenizer/segmenter), creating identified candidate objects with opaque IDs (without embedded quality judgment, stable within execution), extractor origin, source field, original text/URL, structural type (title, subtitle, author, date, block, heading, list, quote, link, image), content hash, and cross-extractor equivalences. Duplicate candidate IDs MUST cause an internal construction failure.
- FR-015: System MUST NOT use regular expressions (
reor regex engines) anywhere in the text processing pipeline, classification, hygiene, parsing, or test assertions. - FR-016: System MUST NOT use manual keyword dictionaries or hardcoded word lists per language to make semantic decisions regarding advertising, recommendations, or editorial value.
- FR-017: System MUST deterministically resolve the source URL using URL parsers in priority order:
crawled_url→input_meta.url→ canonical URL ofselected_extractor→newspaper4k.canonical_link→trafilatura.canonical_url. The LLM MUST NOT select or modify the source URL. - FR-018: System MUST deterministically resolve the publication date normalized to ISO 8601, prioritizing consensus across sources, or fallback priority
newspaper4k.publish_date→input_meta.quando_publicado→trafilatura.date, omittingpublished_atif invalid. The LLM MUST NOT select or modify the publication date. - FR-019: System MUST resolve candidate metadata sources:
- Title:
input_meta.titulo,page_title,trafilatura.title,newspaper4k.title,readability.title,readability.short_title. - Subtitle:
input_meta.subtitulo,trafilatura.description,newspaper4k.meta_description. - Author:
trafilatura.author, items ofnewspaper4k.authors(preserving structured list items and order without regex/delimiter splitting),readability.author.
- Title:
- FR-020: System MUST use
selected_extractoras the structural backbone for ordering, using secondary extractors as consensus evidence and alternative candidate sources, without adopting similarity thresholds not calibrated on the golden set.selected_extractordoes not force the LLM to keep all its blocks; secondary blocks enter only by explicit LLM selection and grounding validation. Low similarity MUST keep candidates distinct. - FR-021: System MUST parse malformed HTML into a safe DOM or controlled failure, register structural metadata from JSON-LD
Article/NewsArticle, and ignore invalid JSON-LD with a logged warning.
LLM Extractive Content Hygiene and Controlled Repairs
- FR-022: System MUST invoke the LLM content hygiene step for every article, even when all three extractors agree 100%.
- FR-023: System MUST provide the LLM with candidate IDs, metadata candidates, block candidates, link candidates, image candidates, detected language, and strict JSON output schema, and the LLM MUST return only selected IDs and repair diffs without outputting free-form full article text or free Markdown.
- FR-024: System MUST require the hygiene output schema to return:
title_candidate_id,subtitle_candidate_id(or null),author_candidate_id(or null),kept_block_ids(in order),kept_link_ids,kept_image_ids,repairs, and categorical removal reasons when enabled for observability. - FR-025: System MUST execute the strict 10-step hygiene harness validation: 1. parse JSON; 2. validate schema; 3. validate metadata IDs and types; 4. validate block IDs; 5. validate order compatibility with canonical representation; 6. validate links and images against input candidates; 7. validate repairs individually; 8. assemble intermediate structure by retrieving candidates from internal maps; 9. validate grounding of assembled Markdown; 10. validate minimum content requirements (rejecting selections omitting material content).
- FR-026: System MUST reject any hygiene response containing ungrounded candidate IDs, ungrounded URLs, ungrounded images, or candidate text not originating from input candidates as a
GROUNDING_VIOLATION, invalidating the entire response and triggering fallback. Schema violations MUST trigger semantic fallback without being labeled as grounding violations, and if all fallback options are exhausted, processing MUST terminate withHYGIENE_FAILED. - FR-027: System MUST validate proposed text repairs against closed categories (
encoding,unicode,spacing,punctuation_corruption,obvious_typo), requiring target candidate ID, exact original fragment, replacement fragment, category, and short rationale. Harness MUST normalize/tokenize original and replacement via Unicode/NLP libraries without regex, compute diffs without regex, and record original, replacement, decision, and rationale. - FR-028: System MUST reject any repair modifying facts, names, dates, numbers, scores, quotes, tone, or style, or referencing ambiguous/missing fragments (
INVALID_TEXT_REPAIR), UNLESS the difference is strictly an unmistakable, verifiable encoding/Unicode defect (e.g. mojibake in a proper name). If in doubt, original text MUST be preserved. - FR-029: System MUST discard invalid repairs while preserving the exact original candidate text, continuing processing without authorizing unconstrained regeneration.
- FR-030: System MUST enforce editorial rules: no translation, no summarization, no narrative reorganization, no transition creation, no information completion, no factual correction, no linguistic variant changes, removal of exact structural title/subtitle duplicates, no repetition of author/date/sentiment/tags/ECP in the body, image alt/captions derived strictly from input text, structurally grounded image positioning (omitting images without editorial position), and link validation (link URLs and anchor texts must exist in input, malformed links structurally rejected). The LLM MUST NEVER control the renderer or filesystem.
ECP Gate and Relevance Enforcement
- FR-031: System MUST submit intermediate sanitized Markdown to the ECP classifier adapter prior to generating final Markdown.
- FR-032: System MUST require the ECP adapter output to provide category,
is_inherent, confidence, rationale, and evidences, and MUST verify that all evidence fragments belong to the intermediate sanitized Markdown. - FR-033: System MUST proceed to enrichment only when ECP returns
DIRECT_INHERENTorCONTEXTUAL_INHERENT. - FR-034: System MUST block Markdown output and generate a persisted
rejected_ecpmanifest when ECP returnsTANGENTIALorNOT_RELATED(ECP_REJECTED). - FR-035: System MUST terminate with
ECP_CLASSIFICATION_FAILEDand block Markdown output if the ECP classifier fails, returns invalid enums, or contains evidence not belonging to the document. - FR-036: System MUST verify that any LLM tier used by the ECP classifier uses certified cheap models and records generation telemetry.
Post-ECP Enrichment
- FR-037: System MUST invoke enrichment only after ECP approval, submitting final title, subtitle (if any), body Markdown, language, and minimal ECP identity.
- FR-038: System MUST classify entity sentiment as strictly
positive,negative, orneutralrelative specifically to the ECP entity. - FR-039: System MUST generate between 3 and 8 unique tags in the article's native language, grounded in content evidence with candidate evidence IDs. Tag uniqueness MUST be verified via Unicode/NLP libraries without regex. Responses attempting to output or alter body text MUST be rejected.
- FR-040: System MUST terminate processing with
ENRICHMENT_FAILEDand generate no Markdown if primary and fallback enrichment calls fail.
Model Gateway and Fallback Strategy
- FR-041: System MUST route runtime LLM invocations through an agnostic Model Gateway supporting logical roles
runtime_primaryandruntime_fallback. The gateway MUST accept role, messages/context, schema, timeout, and trace metadata, and return structured output, effective provider, effective model, tokens (input, output, cached), cost, latency, technical status, attempt number, and fallback indicator. - FR-042: System MUST configure each role with provider, model, parameters, compatible prompt, expected schema, timeout, and version. Only certified cheap models MAY be configured in runtime roles; powerful/expensive models MUST NOT be configured or called in the runtime.
- FR-043: System MUST perform limited technical retries on the same provider for: timeout, connection interruption (connection reset), HTTP 429 (with backoff up to configured limit), HTTP 5xx, and empty technical response. System MUST switch immediately to
runtime_fallbackupon semantic failure (invalid schema, grounding violation) or retry exhaustion, without entering semantic retry loops on the same model. - FR-044: System MUST use two adapters/configurations (
runtime_primary,runtime_fallback) without implementing smart routers, autonomous model selectors, or dynamic model scoring. - FR-045: System MUST apply a conservative deterministic fallback for hygiene only if baseline structural integrity is satisfied (using
selected_extractorbackbone, eliminating structurally invalid elements, without regex and without semantic ad filtering); otherwise, it must terminate withHYGIENE_FAILED.
State Machine, Persistence, and Rendering
- FR-046: System MUST implement orchestration as an explicit Python state machine (
received→validated→content_cleaned;content_cleaned→ecp_approved→enriched→completed_text;content_cleaned→ecp_rejected; valid terminal failures →failed) persisted in SQLite (WAL mode, short transactions, configurable lock timeout), without using LangChain, LangGraph, agents, planners, workflow frameworks, Postgres, external message queues, or object storage within the runtime. Each state transition MUST recordstart_time,end_time,duration, andresult. - FR-047: System MUST output a structured JSON result on every invocation and persist a
<fingerprint>.result.jsonmanifest whenever a deterministic fingerprint is established, including fingerprint, source URL, selected extractor, final status, generate_markdown decision, markdown_path (or null), ECP classification/confidence (or null), provider versions (or null), model versions (or null), prompt versions/hashes (or null), config_version, trace_id (or null), and error/rejection codes. - FR-048: System MUST generate
<fingerprint>.mdconditionally (only upon ECP approval and valid enrichment) containing mandatory YAML front matter (title,source_url,sentiment,tags,ecp_qid,ecp_canonical_name,ecp_category,ecp_confidence) and optional fields (subtitle,author,published_at) omitted when absent. - FR-049: System MUST render Markdown body with H1 title, italic subtitle (when present, with NO artificial blank line generated if absent), and editorial blocks in canonical order (headings, paragraphs, bold, italic, blockquotes, lists, links, images), without underline or inline HTML.
- FR-050: System MUST persist Markdown and manifest through temporary files and atomic renames, and persist SQLite state in the same logical completion unit, verifying content hashes before renaming. Manifest and Markdown MUST NOT be presented as completed while inconsistent. Hash-based reconciliation MUST resolve crashes between file write and SQLite update without duplicate processing. Persistence failures MUST terminate with
PERSISTENCE_FAILED.
Security
- FR-051: System MUST treat HTML and article text as untrusted data candidates, never executing scripts embedded in inputs.
- FR-052: System MUST treat input URLs as data, never accessing or crawling them during runtime execution.
- FR-053: System MUST construct output filenames strictly from deterministic fingerprints to prevent path traversal attacks.
- FR-054: System MUST NOT overwrite pre-existing output files from other executions without an exact idempotent fingerprint match.
- FR-055: System MUST manage credentials strictly via secure environment mechanisms (never CLI parameters, committed configs, or log outputs) and require minimal directory permissions.
- FR-056: System MUST enforce maximum input size limits, failing in a controlled manner before the provider or applying a previously approved context strategy if exceeded.
Prompts and Context Engineering
- FR-057: System MUST maintain exactly two atomic prompt responsibilities:
article_content_hygieneandarticle_sentiment_tags. Prompts MUST be versioned in repository files with semantic version, file hash, expected schemas, and associated Promptfoo test cases. Production and Promptfoo MUST load identical prompt files. Langfuse receives prompt references but is not the primary source. Promptfoo suites MUST explicitly test:- Accepted Repairs when unmistakable: mojibake, broken Unicode, accidental spacing, corrupted punctuation, small typos.
- Rejected Repairs: synonyms, paraphrasing, title improvements, entity name changes without verified encoding defects, dates/numbers/scores alterations, factual corrections, tone changes, quote rewriting.
- Mandatory Assertions: JSON Schema validation, custom Python validators without regex, candidate IDs belonging to context, expected sets and ordering, URLs belonging to input candidates, precision and recall metrics, enum and cardinality validation, diffs computed with sequence/Unicode libraries without regex, comparison with ground truth reference data, and cost/latency threshold metrics.
- FR-058: System MUST structure LLM context in strict normative order: 1. system rules; 2. call responsibility; 3. schema and enums; 4. structural context; 5. candidates and evidence; 6. final structured response request. Article content MUST be delimited as data and separated from instructions.
- FR-059: System MUST strictly exclude from LLM context: full raw JSON when selected fields suffice, full raw HTML when reduced DOM/AST suffices, data from other articles, logs, rejected prior responses (except technical fallback metadata), full ECP when minimal identity suffices, secrets, self-healing instructions, and language-specific semantic keyword examples.
Observability, Logging, and Metrics
- FR-060: System MUST record Langfuse traces under stable
run_idwith stable spans covering validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, and persistence, without embedding URLs, model names, or dynamic IDs in span names. - FR-061: System MUST record each LLM attempt as a separate generation capturing prompt versions/hashes, model/provider, schema version, cached tokens (when available), token counts, calculated cost, latency, timeout status, attempts, technical status, applicable scores, schema results, fallback status, logical role, normalized context sent, structured response, grounding validation result, applied repairs, and rejected repairs (subject to trace content policy), without duplicating raw HTML or full JSON payloads.
- FR-062: System MUST provide configuration to disable textual content in Langfuse traces while preserving hashes, metrics, and execution status.
- FR-063: System MUST redact all API keys, authorization headers, and environment secrets from logs, traces, and output files (
trace_redaction_failure_total= 0). Standard logs MUST omit full ECP, full HTML, and full article text. Repair diffs MUST reside in controlled traces only, never in metric labels. - FR-064: System MUST degrade gracefully when Langfuse is unavailable by storing pending telemetry in SQLite (
TELEMETRY_PENDING), attempting a flush on shutdown, and providing an operational resend routine that returnstelemetry_pending_totalto zero without blocking article processing. - FR-065: System MUST emit structured JSON logs capturing timestamp, severity, environment, run_id, fingerprint, state, event, error codes, logical call, provider/model, prompt version, duration, retry/fallback, trace ID, and final status. Stack traces for unexpected errors MUST be logged locally and sanitized.
- FR-066: System MUST enforce metric label cardinality, forbidding URLs, fingerprints, run IDs, trace IDs, titles, authors, full text, and free tags as metric labels. Metric labels MUST be restricted to:
environment,state,error_code,logical_call,provider,model,prompt_version,schema_version,language,extractor,ecp_category,reason, andsource_domain_group. - FR-067: System MUST instrument the normative metric groups with each metric strictly using its specific dimensions defined in the metric catalog:
- Volume/Result:
article_received_total,article_validated_total,article_duplicate_total,article_completed_text_total,article_rejected_ecp_total,article_failed_validation_total,article_failed_processing_total. - Derived Rates: text completion rate per validated article, ECP rejection rate per validated article, validation failure rate per received article, processing failure rate per validated article, duplication rate per received article.
- Input:
input_validation_failure_total(dimensions:reason,schema_version),selected_extractor_total,selected_extractor_unavailable_total,source_language_total,source_domain_group_total,ecp_version_total. - Hygiene:
hygiene_call_total,hygiene_schema_failure_total,hygiene_grounding_failure_total,hygiene_fallback_total,hygiene_deterministic_fallback_total,hygiene_terminal_failure_total,block_candidate_total,block_kept_total,block_removed_total,link_kept_total,image_kept_total. - Repairs:
text_repair_proposed_total,text_repair_applied_total,text_repair_rejected_total,text_repair_category_total,text_repair_ambiguous_target_total,text_repair_sensitive_change_total. - ECP:
ecp_classification_total,ecp_classification_failure_total,ecp_pass_total,ecp_reject_total,ecp_latency_seconds,ecp_fallback_tier_total. - Enrichment:
sentiment_total,tag_count,enrichment_schema_failure_total,enrichment_grounding_failure_total,enrichment_fallback_total,enrichment_terminal_failure_total. - LLM:
llm_request_total(dimensions:logical_call,provider,model,status),llm_input_tokens_total,llm_output_tokens_total,llm_cost_total,llm_latency_seconds,llm_retry_total(dimensions:reason,provider,model),llm_fallback_total(dimensions:logical_call,reason),llm_output_validation_failure_total. - State/Persistence:
state_transition_total,state_transition_failure_total,sqlite_lock_wait_seconds,sqlite_busy_failure_total,atomic_write_failure_total,resume_total,idempotent_hit_total,orphan_temp_file_total. - Observability:
trace_created_total,telemetry_send_failure_total,telemetry_pending_total,telemetry_flush_failure_total,trace_content_disabled_total,trace_redaction_failure_total. - Capacity: throughput, concurrency, total duration p50/p95/p99, duration per state p50/p95/p99, CPU, memory, SQLite growth, disk usage separated by outputs and temporary files, lock wait seconds, saturation failures, tokens/cost per received article, tokens/cost per approved Markdown.
- Volume/Result:
- FR-068: System MUST record objective prompt review signals (
prompt_review_signal_total) capturinglogical_call,prompt_version,provider,model,language,source_domain_group, andreasonwithout executing autonomous self-healing. - FR-069: System MUST provide minimum dashboards for Runtime Health, Quality, and Future Review Signals.
Testing, CI Quality Gates, and Staging Baselines
- FR-070: System MUST enforce static AST verification in CI to prevent regex imports or calls in text pipeline modules and test assertions.
- FR-071: System MUST execute unit tests, contract tests, simulated integration tests (with simulated/mock providers; real remote providers are permitted strictly in authorized offline evaluations and calibration suites), and basic security tests on every pull request.
- FR-072: System MUST execute Promptfoo evals (in dev/CI), regression of the 20 reference cases, and cost/latency comparisons on every change to prompts, context, schema, or models.
- FR-073: System MUST execute full golden set evaluation across stratified language/domain/extractor slices (with golden set size justified by observed stability and confidence intervals, holdout never used for few-shot examples, and ground truth covering: full raw input, ECP and version, expected status, title/subtitle/author/date expected or candidates, kept/removed blocks, expected links/images, allowed/forbidden repairs, expected ECP, expected sentiment, accepted tags or a closed evaluation rubric, expected Markdown/structure, discard reason), full security tests on every release, reprocessing tests on every release, fault injection (10 scenarios: primary down, both down, Langfuse unavailable during processing, SQLite lock timeout, disk full, process terminated during write, ECP down, truncated LLM response, orphan temp, telemetry flush failure), and 100 articles/hour load test before promotion.
- FR-074: System MUST evaluate quality metrics (precision, recall, F1 of kept blocks; metadata accuracy; correct vs unauthorized repair rates; material loss; residual noise; link/image precision; ECP accuracy; sentiment accuracy; tag acceptance) across slices (language, domain, extractor, prompt version, model version) requiring at least 95% pass rate per slice without allowing global averages to hide slice failures. 100% of accepted LLM responses MUST have valid schema.
- FR-075: System MUST enforce zero-tolerance release gates blocking promotion if any of the 11 critical invariants occurs:
ungrounded_text_total> 0,ungrounded_url_total> 0,ungrounded_image_total> 0,unauthorized_rewrite_total> 0,critical_fact_change_total> 0,duplicate_output_total> 0,lost_article_total> 0,secret_exposure_total> 0,powerful_runtime_model_call_total> 0,online_promptfoo_call_total> 0,text_regex_usage_total> 0. - FR-076: System MUST sustain 100 articles/hour load in staging using production-equivalent SQLite and filesystem configurations with zero lost articles, zero duplicate outputs, zero partial files exposed, zero database corruption, stable memory/disk, 100% traces sent or queued, establishing empirical cost and latency baselines (max cost per article, max cost per approved Markdown, p50/p95/p99 latency, provider timeouts, fallback limits, storage limits) for approval prior to go-live.
- FR-077: System MUST preserve mandatory release evidence artifacts and a complete staging report containing: code version, prompt versions/hashes, providers/models config, Promptfoo config, golden set hashes, per-case and per-slice results, critical violation reports, load test report, formal release sign-off, corpus size and composition (languages, domains, extractors, structures), code/prompt/model/ECP versions, throughput, latency p50/p95/p99 per step and total, cost p50/p95/p99 per article, cost per approved Markdown, fallback rate, resource utilization (CPU, memory, disk, SQLite growth), failures, and outliers.
Production Operations and Runbook Procedures
- FR-078: System MUST implement preflight checks validating system clock synchronization, prompts and configs belonging strictly to the same release, local configuration reading, prompt file existence & hashes, schema compatibility, SQLite access, filesystem atomic write permissions & directory permissions, minimum disk space, validated active credentials (not just presence), absence of powerful models in runtime roles, Langfuse local configuration (environment, content policy), and ECP classifier/schema access before accepting live traffic. Remote Langfuse connectivity failure MUST NOT block preflight.
- FR-079: System MUST implement smoke tests executing a versioned fixture to verify end-to-end processing, artifacts, Langfuse trace, cost and latency within approved staging ranges, and idempotent re-execution.
- FR-080: System MUST follow the complete 11-step deployment sequence (pause new runs, drain active runs, backup SQLite/config, deploy package/prompts/schemas, test migration on copy, preflight, smoke test, verify artifacts/state/trace, release with reduced concurrency, verify metrics, release full volume).
- FR-081: System MUST support consistent SQLite backup/restore mechanisms, graceful shutdown upon receiving a shutdown signal (stopping new units, completing active unit safely, closing transactions, flushing files/telemetry, preserving pending telemetry, exiting cleanly), periodic reconciliation reports, safe orphan temporary-file cleanup through approved operational/reconciliation routines, manual rollback procedures to previous certified configurations, and certified model/provider rotation.
- FR-082: System MUST enforce operational retention policies for manifests, Markdown, logs, and traces.
- FR-083: System MUST provide credential rotation procedures updating environment secrets, running preflight/smoke tests, revoking old credentials, and avoiding altering functional fingerprints when only operational secrets change.
- FR-084: System MUST provide documented incident procedures for providers, ECP, enrichment, Langfuse, SQLite, disk, grounding, cost, latency, and input validation / producer schema mismatches (checking producer version, comparing with release schema, confirming unit payload, never calling LLM manually, fixing producer or contract via standard release), preserving incident evidence and creating mandatory regression test cases after critical incidents. Production prompts, sentiments, tags, Markdown files, manifests, and SQLite MUST NEVER be edited manually. Reprocessing MUST locate state/fingerprint, verify versions, reuse existing outputs or resume incomplete states under the same configuration, generate distinct fingerprints for new configurations, avoid manual file edits, and never reprocess articles merely to recreate traces.
Key Entities
- Article Input Unit: Single news article JSON containing
crawled_url,error_message,extraction_status,http_status,input_meta,page_title,selected_extractor(trafilatura|newspaper4k|readability), individual extractor payloads (trafilatura,newspaper4k,readability), and preserved unknown fields. - Entity Context Profile (ECP) Canonical Schema Reference: Canonical versioned profile managed exclusively by the ECP module; referenced by the runtime without duplicating schema definitions.
- Candidate Object: Identifiable structural unit (title, subtitle, author, date, block, heading, list, quote, link, or image) with an opaque ID without quality judgment (stable within execution), extractor origin, source field, content hash, and cross-extractor equivalences.
- Text Repair Operation: Controlled micro-edit specifying target candidate ID, exact original fragment, replacement fragment, category (
encoding|unicode|spacing|punctuation_corruption|obvious_typo), and short rationale. - State Machine Record: SQLite-persisted lifecycle state containing
fingerprint,current_state(received→validated→content_cleaned;content_cleaned→ecp_approved→enriched→completed_text;content_cleaned→ecp_rejectedas terminal state without Markdown; valid terminal failures →failed),timestamps(start_time,end_time,duration),result,output_paths,file_hashes,functional_versions,terminal_error, andpending_telemetry. - Output Manifest (
<fingerprint>.result.json): Machine-readable summary containing fingerprint, source URL, selected extractor, final status (completed_text|rejected_ecp|failed_validation|failed_processing), generate_markdown decision, markdown_path (or null), ECP classification/confidence (or null), provider versions (or null), model versions (or null), prompt versions/hashes (or null), config_version, trace_id (or null), and error/rejection codes. - Published Markdown (
<fingerprint>.md): Markdown document with structured YAML front matter and clean, grounded editorial body. - Telemetry Event: Queued observability payload stored in SQLite (
TELEMETRY_PENDING) for deferred transmission when Langfuse is unavailable.
Success Criteria (mandatory)
Measurable Outcomes
- SC-001 (Zero Hallucination): 100% of published text, URLs, and images are traceable to input candidates or approved repairs (
ungrounded_text_total= 0,ungrounded_url_total= 0,ungrounded_image_total= 0). - SC-002 (Zero Critical Fact Corruption): 0 unauthorized rewrites or alterations to names, dates, numbers, scores, quotes, or facts across all evaluations (
unauthorized_rewrite_total= 0,critical_fact_change_total= 0). - SC-003 (Zero Duplication & Data Integrity): 0 duplicate outputs generated for identical input fingerprints and functional configurations; 0 lost articles (
duplicate_output_total= 0,lost_article_total= 0). - SC-004 (End-to-End Quality Pass Rate by Slice): At least 95% end-to-end pass rate across the reference golden dataset and across every stratified language, domain, extractor, prompt version, and model version slice without global average masking. 100% of accepted LLM responses have valid schema.
- SC-005 (Telemetry Completeness): 100% of validated executions have complete traces either delivered to Langfuse or preserved in SQLite as pending telemetry (
trace_created_total/article_validated_total= 100%);trace_redaction_failure_total= 0;telemetry_pending_totalreturns to zero after recovery. - SC-006 (Throughput and Concurrency): Sustained throughput of at least 100 articles per hour under realistic concurrency without data loss, partial file exposure, or database corruption, with controlled handling of lock waits and measured
sqlite_lock_wait_secondsandsqlite_busy_failure_total. - SC-007 (Empirical SLO Approval): Measured p50, p95, and p99 latency (per step and total), cost per received article, cost per approved Markdown, provider timeouts, acceptable fallback threshold, and storage limits established in staging and approved prior to production go-live.
- SC-008 (Strict Regex & Dependency Discipline): 0 occurrences of regular expression imports/calls within text pipeline modules and test assertions (
text_regex_usage_total= 0); all dependencies strictly justified and locked in project lockfile. - SC-009 (Model Cost Control): 0 invocations of powerful/expensive LLMs within runtime roles (
powerful_runtime_model_call_total= 0). - SC-010 (Operational Readiness): 100% completion of preflight checks, smoke tests, 11-step deployment sequence verification, backup/restore verifications, graceful shutdown handling, reconciliation routines, and manual rollback drills.
Traceability Matrix
| Documento Fonte | Cláusula / Tópico Normativo | Cobertura Específica na Spec |
|---|---|---|
| 01_PRD | 1–2: Contexto e Objetivo (Artigo único, ECP, pico 100 art/h) | US1, US6, FR-003, FR-005, FR-076, SC-006 |
| 01_PRD | 3: Princípio Mestre (Zero complexidade supérflua, orquestração Python direta) | FR-001, FR-046, SC-008, Assumptions |
| 01_PRD | 4.1: Fundamentação e Proibição de Invenção/Alucinação | US2, FR-023, FR-026, SC-001, SC-002 |
| 01_PRD | 4.2: Proibição Estrita de Regex e Listas Manuais de Palavras | US1, FR-015, FR-016, FR-070, SC-008 |
| 01_PRD | 4.3: Higienização LLM Obrigatória Mesmo em Consenso 100% | US2, FR-022, Edge Cases |
| 01_PRD | 4.4: Modelos Baratos no Runtime (Zero modelos caros) | US6, FR-042, SC-009, Assumptions |
| 01_PRD | 5: Escopo Incluído / Fora de Escopo (Vídeo/galeria upstream, sem self-healing) | FR-003, FR-068, Assumptions, Edge Cases |
| 01_PRD | 6–7: Atores e Unidade de Processamento (1 artigo + ECP snapshot) | US1, FR-003, FR-005, Key Entities |
| 01_PRD | 8: Contrato do Artigo (selected_extractor, campos dos extratores) |
US1, FR-006, FR-007, FR-010, FR-011 |
| 01_PRD | 9: Contrato do ECP (Snapshot canônico, Gate, 4 classificações) | US1, US3, FR-005, FR-031, FR-032, FR-033, FR-034 |
| 01_PRD | 10: Contrato de Saída (JSON, 4 status, Markdown YAML front matter) | US5, FR-047, FR-048, FR-049, Key Entities |
| 01_PRD | 11–13: Fluxo, Preparação Determinística, Resolução URL/Data, Candidatos | US1, FR-008, FR-014, FR-017, FR-018, FR-019, FR-020, FR-021 |
| 01_PRD | 14: Higienização Extrativa LLM (Entrada, Saída por IDs, Regras Editoriais) | US2, FR-023, FR-024, FR-025, FR-026, FR-030 |
| 01_PRD | 15: Pequenos Reparos Textuais (5 categorias, diffs, reversibilidade, exceção encoding) | US2, FR-027, FR-028, FR-029, Key Entities |
| 01_PRD | 16: Imagens e Links Editoriais (Grounding estrutural, alt/caption, links válidos) | US2, FR-026, FR-030 |
| 01_PRD | 17–18: Gate ECP e Enriquecimento (Sentimento relativo, 3-8 tags nativas) | US3, US4, FR-031, FR-032, FR-037, FR-038, FR-039, FR-040 |
| 01_PRD | 19–20: Model Gateway, Idempotência e Persistência Atômica | US5, US6, FR-008, FR-009, FR-041, FR-043, FR-050 |
| 01_PRD | 21–22: Observabilidade Langfuse e Promptfoo fora do runtime | US7, US8, FR-057, FR-060, FR-064, FR-072 |
| 01_PRD | 23–25: Requisitos Funcionais, NFRs e 16 Códigos Mínimos de Erro | FR-001 a FR-084, Acceptance Scenarios, Edge Cases |
| 01_PRD | 26–29: Critérios de Aceite do Produto, Métricas e DoD | SC-001 a SC-010, FR-067, FR-075, FR-076 |
| 02_Arquitetura | 1–6: Propósito, Direcionadores, Limites e Componentes | US1 a US9, FR-001 a FR-084 |
| 02_Arquitetura | 7: Orquestração (Máquina de estados Python, SQLite WAL, sem frameworks/agentes) | FR-046, Key Entities |
| 02_Arquitetura | 8–10: Módulos, Política de Dependências, Contratos Versionados | FR-001, FR-002, FR-004, FR-015, SC-008, Assumptions |
| 02_Arquitetura | 11–14: Validação sem chamadas remotas, Fingerprint, Parsing, Candidatos | US1, FR-007, FR-008, FR-012, FR-014, FR-019, FR-020, FR-021 |
| 02_Arquitetura | 15–19: Higienização, Reparos, Assembler, ECP Adapter, Enriquecimento | US2, US3, US4, FR-023, FR-024, FR-027, FR-031, FR-032, FR-037 |
| 02_Arquitetura | 20–22: Model Gateway (2 adapters, sem router), Prompts no repo, Escrita Atômica | US5, US6, FR-041, FR-042, FR-044, FR-048, FR-050, FR-057 |
| 02_Arquitetura | 23–26: Observabilidade, Logs JSON, Segurança, Concorrência 100 art/h | US7, FR-051 a FR-056, FR-060 a FR-066, FR-076 |
| 02_Arquitetura | 27–33: Falhas, Deploy, CI/CD, Simplicidade, Riscos | US8, US9, FR-045, FR-070 a FR-077, FR-078 a FR-084 |
| 03_ADRs | ADR-001: Separação de Runtime e Self-Healing | FR-068, Assumptions |
| 03_ADRs | ADR-002: Início após seleção do extrator (Sem recálculo) | FR-006, Edge Cases |
| 03_ADRs | ADR-003: Orquestração direta em Python (Sem LangChain/LangGraph/agentes) | FR-046 |
| 03_ADRs | ADR-004: Gateway agnóstico com apenas modelos baratos no runtime | US6, FR-041, FR-042, SC-009 |
| 03_ADRs | ADR-005: Proibição de regex e palavras-chave manuais em decisões textuais | US1, FR-015, FR-016, FR-070, SC-008 |
| 03_ADRs | ADR-006: LLM seleciona IDs e propõe reparos (Não regenera o artigo) | US2, FR-023, FR-024, FR-027, FR-030 |
| 03_ADRs | ADR-007: ECP obrigatório antes de toda saída editorial | US3, FR-005, FR-031, FR-033, FR-034 |
| 03_ADRs | ADR-008: Langfuse no runtime e Promptfoo no CI | US7, US8, FR-057, FR-060, FR-064, FR-072 |
| 03_ADRs | ADR-009: Persistência em SQLite (WAL) e saídas no filesystem | US5, FR-046, FR-050 |
| 03_ADRs | ADR-010: Resultado estruturado sempre e Markdown condicional | US5, FR-047, FR-048 |
| 03_ADRs | ADR-011: Definição de SLOs de custo e latência a partir de staging | US8, FR-076, SC-007 |
| 04_Plano_Testes | 1–7: Objetivo, Princípios, Camadas de Teste, Dados, Golden Set (Contrato 4.2 e 4.4), Holdout, Gates | US8, FR-071, FR-072, FR-073, FR-074, FR-075, SC-004 |
| 04_Plano_Testes | 8: Matriz IN (IN-001 a IN-015: Contratos de entrada e validações locais) | US1, FR-003, FR-005, FR-006, FR-007, FR-010, FR-011, FR-012 |
| 04_Plano_Testes | 9: Matriz ID (ID-001 a ID-010: Fingerprint, idempotência, concorrência e reconciliação) | US1, US5, FR-008, FR-009, FR-050 |
| 04_Plano_Testes | 10: Matriz PAR (PAR-001 a PAR-010: Parsing estrutural, malformed HTML, JSON-LD, sem regex) | US1, US8, FR-014, FR-015, FR-021, FR-070 |
| 04_Plano_Testes | 11: Matriz CAN (CAN-001 a CAN-010: Candidatos, IDs únicos, similaridade, sem imagem auto) | US1, FR-014, FR-019, FR-020, FR-030 |
| 04_Plano_Testes | 12: Matriz HYG (HYG-001 a HYG-021: 10 passos do harness, grounding, fallback determinístico) | US2, US6, FR-022 a FR-030, FR-043, FR-045 |
| 04_Plano_Testes | 13: Matriz REP (REP-001 a REP-016: Reparos permitidos, proibidos, sensíveis e reversão) | US2, FR-027, FR-028, FR-029, FR-030 |
| 04_Plano_Testes | 14: Matriz ECP (ECP-001 a ECP-009: Gate ECP, evidências no doc e cheap tiers) | US3, FR-031 a FR-036 |
| 04_Plano_Testes | 15: Matriz ENR (ENR-001 a ENR-009: Sentimento relativo, tags nativas sem duplicatas, falhas) | US4, FR-037 a FR-040 |
| 04_Plano_Testes | 16: Matriz OUT (OUT-001 a OUT-012: Markdown YAML, sem linha artificial sem subtítulo, escrita atômica) | US5, FR-047, FR-048, FR-049, FR-050 |
| 04_Plano_Testes | 17: Matriz LLM (LLM-001 a LLM-012: Retries técnicos autorizados, fallbacks e gates de modelos) | US6, FR-041, FR-042, FR-043, FR-044, FR-045 |
| 04_Plano_Testes | 18: Matriz OBS (OBS-001 a OBS-011: Traces, spans, degradação graciosa e logs) | US7, FR-060 a FR-065 |
| 04_Plano_Testes | 19: Matriz SEC (SEC-001 a SEC-008: Prompt injection, traversal, pre-existing files, secrets, limites) | FR-051 a FR-056 |
| 04_Plano_Testes | 20: Teste de Carga (100 art/h, sem perda, estabilidade de memória e disco) | US8, FR-076, SC-006 |
| 04_Plano_Testes | 21: Fault Injection (FLT-001 a FLT-010: 10 cenários completos incluindo Langfuse down) | US8, FR-073, FR-084 |
| 04_Plano_Testes | 22–25: Promptfoo em CI, Ordem CI/CD, Evidências Preservadas e Conclusão | US8, FR-057, FR-072, FR-077 |
| 05_Metricas | 1–4: KPIs do Produto e 11 Invariantes Críticas (Meta zero) | US8, SC-001 a SC-010, FR-075 |
| 05_Metricas | 5–11: Métricas de Volume, Entrada, Higienização, Reparos, ECP, Enriquecimento, LLM | US7, FR-067 |
| 05_Metricas | 12: Sinais para Revisão Futura de Prompt (prompt_review_signal_total com 7 dimensões) |
US7, FR-068 |
| 05_Metricas | 13–15: Métricas de Persistência, Observabilidade e Capacidade (Durações, Disco, Locks) | US7, FR-067, SC-006 |
| 05_Metricas | 16: Baseline de Staging e Relatório Completo Obrigatório | US8, FR-076, FR-077, SC-007 |
| 05_Metricas | 17–20: Logs Estruturados, 3 Dashboards Mínimos e Dimensões por Métrica no Catálogo | US7, FR-063, FR-065, FR-066, FR-067, FR-069 |
| 06_Runbook | 1–6: Princípios Operacionais, Matriz de Responsabilidades (4 papéis) e Pré-requisitos | US9, FR-013, FR-078, FR-084, Key Entities |
| 06_Runbook | 7–11: Checklist de Release, Sequência de Implantação 11 Passos, Preflight e Smoke Test | US9, FR-078, FR-079, FR-080 |
| 06_Runbook | 12–14: Monitoramento, Logs/Correlação e Reprocessamento Idempotente | US7, US9, FR-009, FR-065, FR-081, FR-084 |
| 06_Runbook | 15–25: Diagnósticos de Incidentes (Todos subsistemas + Entrada/Produtor) e Proibição Edição Manual | US9, FR-064, FR-084, Edge Cases |
| 06_Runbook | 26: Rollback Manual (Gatilhos, Procedimento e Retomada) | US9, FR-081 |
| 06_Runbook | 27–29: Troca Certificada de Modelo, Rotação de Credenciais e Backup/Retenção | US9, FR-081, FR-082, FR-083 |
| 06_Runbook | 30–32: Reconciliação, Shutdown Controlado e Prontidão Operacional | US9, FR-081, SC-010 |
| 07_Prompt_Harness | 1–5: 2 Prompts Atômicos, Regras Comuns, Ordem do Contexto (6 blocos) e Exclusões Estritas | US2, US4, FR-057, FR-058, FR-059 |
| 07_Prompt_Harness | 6–8: Prompt article_content_hygiene, Regras Normativas de Reparos e 10 Passos do Harness |
US2, FR-024, FR-025, FR-027, FR-028, FR-029, FR-030 |
| 07_Prompt_Harness | 9: Prompt article_sentiment_tags e Harness de Enriquecimento (Sem corpo, validação tags) |
US4, FR-037, FR-038, FR-039, FR-040 |
| 07_Prompt_Harness | 10–13: Política de Provider, Casos Promptfoo de Reparo/10 Assertions, Langfuse e Aceite | US6, US7, US8, FR-041, FR-042, FR-044, FR-057, FR-060, FR-061, FR-072 |
Assumptions
- Upstream crawling, HTML fetching, and multi-extractor execution (
trafilatura,newspaper4k,readability) as well asselected_extractorcalculation and video/gallery filtering are performed by prior pipeline stages and are out of scope. - Self-healing prompt optimization, automated prompt mutation, LLM-as-a-judge for prompt improvement, canary deployments, and auto-rollback belong to a distinct future subproject; the runtime only emits telemetry review signals (
prompt_review_signal_total). - All LLM providers configured in runtime roles (
runtime_primary,runtime_fallback) are low-cost models certified via Promptfoo and supporting structured JSON schema outputs. - The local filesystem and SQLite (WAL mode, short transactions, configurable lock timeout) provide the persistence and concurrency foundation for the target workload of 100 articles/hour.
- Execution occurs in the Python version supported by the repository with standard dependencies specified in the project lockfile.