feat(runtime): implement single-article consolidation runtime and modularize codebase
This commit is contained in:
@@ -0,0 +1,129 @@
|
||||
# Requirements Readiness Checklist: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Purpose**: Formal requirements-quality review and readiness checklist covering functional completeness, architectural constraints, security invariants, operational resilience, and contractual consistency across the runtime specification (FR-001 to FR-084).
|
||||
**Created**: 2026-08-23
|
||||
**Feature**: [`spec.md`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/spec.md)
|
||||
|
||||
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
|
||||
|
||||
---
|
||||
|
||||
## 1. Requirement Completeness
|
||||
|
||||
- [x] CHK001 Are extraction payload ingestion and field retention requirements specified for all three supported extractors (`trafilatura`, `newspaper4k`, `readability`)? [Completeness, Spec §FR-010, §FR-011]
|
||||
- [x] CHK002 Are structural preservation requirements explicitly defined for all candidate element types (paragraphs, headings, lists, quotes, links, images)? [Completeness, Spec §FR-014, §FR-030]
|
||||
- [x] CHK003 Are the 10 sequential validation steps of the hygiene harness fully enumerated and ordered in the specification? [Completeness, Spec §FR-024, §FR-025]
|
||||
- [x] CHK004 Are the 5 allowable text micro-repair categories exhaustively defined with explicit acceptance/rejection criteria? [Completeness, Spec §FR-027, §FR-028]
|
||||
- [x] CHK005 Are requirements for ECP relevance classification handling defined for all 4 decision categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`)? [Completeness, Spec §FR-031, §FR-033, §FR-034]
|
||||
- [x] CHK006 Are sentiment classification and native language tag enrichment requirements documented with strict input/output bounds? [Completeness, Spec §FR-037, §FR-038, §FR-039, §FR-040]
|
||||
- [x] CHK007 Are state machine lifecycle transitions and persistence requirements defined for all valid paths from `received` to terminal states? [Completeness, Spec §FR-046, §FR-050]
|
||||
- [x] CHK008 Are all 9 versioned contract schemas identified and cross-referenced with explicit versioning rules? [Completeness, Spec §FR-004]
|
||||
|
||||
---
|
||||
|
||||
## 2. Requirement Clarity & Precision
|
||||
|
||||
- [x] CHK009 Is the input size threshold quantified with an exact byte limit and unambiguous pre-provider failure behavior? [Clarity, Spec §FR-056]
|
||||
- [x] CHK010 Is the definition of "sensitive entities" in text micro-repairs unambiguously clarified to prevent ungrounded modifications to names, dates, numbers, and facts? [Clarity, Spec §FR-027, §FR-028, §FR-029]
|
||||
- [x] CHK011 Are the minimal ECP identity fields supplied to the enrichment prompt strictly limited to `qid` and `canonical_name` without vague contextual keyword lists? [Clarity, Spec §FR-037, §FR-059]
|
||||
- [x] CHK012 Is the candidate equivalence mapping defined explicitly as non-destructive evidence rather than automatic deduplication? [Clarity, Spec §FR-014, §FR-020]
|
||||
- [x] CHK013 Is the zero-regex policy quantified with unambiguous static AST, JSON schema, and Promptfoo evaluation constraints? [Clarity, Spec §FR-015, §FR-070]
|
||||
- [x] CHK014 Are the exit codes of the CLI interface explicitly mapped to specific execution outcomes without ambiguity between article validation and configuration errors? [Clarity, Spec §FR-003, §FR-047]
|
||||
|
||||
---
|
||||
|
||||
## 3. Requirement Consistency & Alignment
|
||||
|
||||
- [x] CHK015 Do state transition definitions align consistently between textual requirements and formal data model entity specifications? [Consistency, Spec §FR-046]
|
||||
- [x] CHK016 Are the error codes in the output manifest strictly consistent with the normative 16-code error catalog? [Consistency, Spec §FR-047]
|
||||
- [x] CHK017 Is the terminal state `ecp_rejected` consistently specified as producing an output manifest with `generate_markdown: false` and exactly zero Markdown files? [Consistency, Spec §FR-034, §FR-046, §FR-050]
|
||||
- [x] CHK018 Do the prompt context specifications in §FR-024 and §FR-037 align with the 6-block prompt architecture defined in the harness specification? [Consistency, Spec §FR-058, §FR-059]
|
||||
- [x] CHK019 Are the decoupling requirements between semantic schema invalidity (fallback trigger) and grounding violations (immediate invalidation) consistently preserved across all hygiene requirements? [Consistency, Spec §FR-026, §US2]
|
||||
|
||||
---
|
||||
|
||||
## 4. Acceptance Criteria & Measurability
|
||||
|
||||
- [x] CHK020 Are all 11 release invariants defined with measurable zero-tolerance thresholds (count = 0)? [Measurability, Spec §FR-075]
|
||||
- [x] CHK021 Can the prompt parity invariant between runtime production prompts and Promptfoo test suites be objectively verified by SHA-256 byte comparison? [Measurability, Spec §FR-057, §FR-077]
|
||||
- [x] CHK022 Are staging performance and latency SLOs formulated as measurable empirical calibration gates prior to production release? [Measurability, Spec §FR-076]
|
||||
- [x] CHK023 Can the batch wrapper rejection rule (`"articles": false`) be objectively evaluated against any composite JSON payload? [Measurability, Spec §FR-003, §FR-007, Contract 1]
|
||||
- [x] CHK024 Is the holdout dataset evaluation criterion objectively separated from prompt few-shot development data? [Measurability, Spec §FR-073]
|
||||
|
||||
---
|
||||
|
||||
## 5. Scenario & Flow Coverage
|
||||
|
||||
- [x] CHK025 Are requirements defined for the primary happy path of direct inherence resulting in published Markdown and manifest? [Coverage, Spec §US1, §FR-031, §FR-033, §FR-037, §FR-038, §FR-039, §FR-040, §FR-047, §FR-048, §FR-049, §FR-050]
|
||||
- [x] CHK026 Are requirements defined for alternate flows involving primary model failure and automated fallback to the secondary provider? [Coverage, Spec §US6, §FR-043, §FR-044, §FR-045]
|
||||
- [x] CHK027 Are requirements defined for exception flows involving unparseable JSON inputs, missing extractors, and schema violations? [Coverage, Spec §US2, §FR-005, §FR-006, §FR-007, §FR-047]
|
||||
- [x] CHK028 Are requirements defined for recovery flows involving interrupted writes and process crash reconciliation? [Coverage, Spec §US8, §FR-009, §FR-050, §FR-081]
|
||||
- [x] CHK029 Are requirements defined for idempotency replay when identical fingerprints are submitted concurrently or sequentially? [Coverage, Spec §US3, §FR-009]
|
||||
|
||||
---
|
||||
|
||||
## 6. Edge Case & Boundary Coverage
|
||||
|
||||
- [x] CHK030 Are requirements specified for handling articles with empty body text, missing titles, or missing source URLs? [Edge Case, Spec §FR-007]
|
||||
- [x] CHK031 Are boundary conditions specified for documents exceeding maximum allowed input byte limits? [Edge Case, Spec §FR-056]
|
||||
- [x] CHK032 Are requirements specified for limited technical retries on timeout, connection interruption/reset, HTTP 429 with backoff up to the configured limit, HTTP 5xx, and empty technical responses? [Edge Case, Spec §FR-043]
|
||||
- [x] CHK033 Are boundary constraints defined for the minimum (3) and maximum (8) allowable tags in enrichment responses? [Edge Case, Spec §FR-038]
|
||||
- [x] CHK034 Is the behavior specified for corrupted Unicode or mojibake in proper names versus factual semantic edits? [Edge Case, Spec §FR-029]
|
||||
|
||||
---
|
||||
|
||||
## 7. Non-Functional & Security Requirements (SEC-001 to SEC-008)
|
||||
|
||||
- [x] CHK035 Are prompt injection resistance requirements specified to prevent instructions within article bodies from overriding system directives (SEC-001)? [Security, Spec §FR-051, §FR-058, §FR-059, §FR-071]
|
||||
- [x] CHK036 Are credential and secret redaction requirements defined for technical stderr logs, trace attributes, and manifests (SEC-002)? [Security, Spec §FR-055, §FR-063, §FR-071]
|
||||
- [x] CHK037 Are filesystem path traversal prevention requirements documented for article paths and output filenames (SEC-003)? [Security, Spec §FR-053, §FR-054, §FR-071]
|
||||
- [x] CHK038 Are requirements defined to prevent local filesystem exhaustion and unbounded temporary file accumulation (SEC-004)? [Security, Spec §FR-050, §FR-056, §FR-078, §FR-081, §FR-082]
|
||||
- [x] CHK039 Are untrusted input size limits specified to prevent denial-of-service via large payloads (SEC-005)? [Security, Spec §FR-056, §FR-071]
|
||||
- [x] CHK040 Are requirements defined to prevent schema poisoning and local duplicate validation definitions (SEC-006)? [Security, Spec §FR-002, §FR-005, §FR-071]
|
||||
- [x] CHK041 Are requirements specified for handling SQLite lock contention and database lock timeouts (SEC-007)? [Security, Spec §FR-009, §FR-046, §FR-071]
|
||||
- [x] CHK042 Are requirements defined for secure telemetry degradation when observability endpoints are unreachable (SEC-008)? [Security, Spec §FR-064, §FR-071]
|
||||
|
||||
---
|
||||
|
||||
## 8. Operational Resilience & Lifecycle Governance (FR-081, FR-084)
|
||||
|
||||
- [x] CHK043 Are consistent database backup and restore requirements documented using native SQLite APIs without distributed database dependencies? [Resilience, Spec §FR-081]
|
||||
- [x] CHK044 Are graceful shutdown requirements defined for `SIGTERM` and `SIGINT` signals to flush in-flight telemetry and prevent SQLite state corruption? [Resilience, Spec §FR-081]
|
||||
- [x] CHK045 Are credential rotation and certified model rotation procedures testable and verifiable via preflight configuration checks? [Resilience, Spec §FR-081, §FR-083]
|
||||
- [x] CHK046 Are rollback procedures specified for reverting releases while preserving offline telemetry and historical state? [Resilience, Spec §FR-081]
|
||||
- [x] CHK047 Is the multi-stakeholder responsibility matrix (Orchestrator, Operations, Engineering, Curator/Eval) unambiguously mapped without overlapping operational boundaries? [Governance, Spec §US9, §FR-084]
|
||||
|
||||
---
|
||||
|
||||
## 9. Dependencies & Contract Traceability
|
||||
|
||||
- [x] CHK048 Are all runtime dependencies evaluated against the 7 mandatory criteria: requirement served, standard-library alternative, security impact, maintenance impact, license, size impact, and startup impact? [Governance, Spec §FR-002]
|
||||
- [x] CHK049 Is the local canonical ECP schema resolution specified via `referencing.Registry` without network HTTP lookups? [Governance, Spec §FR-002, §FR-005]
|
||||
- [x] CHK050 Is the release metadata artifact (`src/core/release-metadata.json`) specified as the immutable verification source for config hashes, prompt hashes, and model certifications? [Governance, Spec §FR-042, §FR-077, §FR-078]
|
||||
|
||||
---
|
||||
|
||||
## 10. Scope Boundaries, Quality Gates & Operations (FR-001 to FR-084)
|
||||
|
||||
- [x] CHK051 Are the master simplicity constraints explicitly reviewable, including minimum necessary code, no anticipatory self-healing, no API, no internal batch or worker pools, and no LangChain, LangGraph, agents, planners, workflow frameworks, Postgres, external queues, or object storage inside the runtime? [Scope, Spec §FR-001, §FR-002, §FR-003, §FR-046]
|
||||
- [x] CHK052 Are requirements explicit that the ECP snapshot is mandatory, selected_extractor is never recalculated or substituted, unknown input fields are preserved, and all terminating local validations occur before any remote call? [Completeness, Spec §FR-005, §FR-006, §FR-007, §FR-010, §FR-011, §FR-012]
|
||||
- [x] CHK053 Are fingerprint composition, functional configuration inclusion, operational secret exclusion, and reproducible packaging requirements completely and unambiguously defined? [Precision, Spec §FR-008, §FR-013]
|
||||
- [x] CHK054 Are deterministic URL, publication date, title, subtitle, and author resolution rules fully specified, including priority orders, omission behavior, and the prohibition on splitting author strings by delimiters? [Completeness, Spec §FR-017, §FR-018, §FR-019]
|
||||
- [x] CHK055 Is mandatory LLM hygiene required even under complete extractor consensus, with output limited to candidate IDs and repair diffs and with all editorial preservation rules explicitly defined? [Completeness, Spec §FR-022, §FR-023, §FR-024, §FR-030]
|
||||
- [x] CHK056 Are the exact manifest, conditional Markdown, YAML front matter, canonical body rendering, hashing, atomic persistence, SQLite consistency, and reconciliation requirements completely specified? [Completeness, Spec §FR-047, §FR-048, §FR-049, §FR-050]
|
||||
- [x] CHK057 Are Langfuse spans, per-attempt generations, trace-content policy, redaction, pending telemetry, structured logs, metric cardinality, normative metrics, dashboards, and non-executing prompt review signals completely specified? [Observability, Spec §FR-060 to §FR-069]
|
||||
- [x] CHK058 Are Promptfoo change triggers, 20-case regression, golden-set ground truth, holdout isolation, 10 fault-injection scenarios, per-slice quality gates, 11 zero-tolerance invariants, load test, and release evidence requirements all objectively verifiable? [Quality Gates, Spec §FR-070 to §FR-077]
|
||||
- [x] CHK059 Are preflight, smoke test, 11-step deployment, retention, reconciliation, orphan cleanup, rotations, incident response, reprocessing rules, and prohibitions on manual production artifact editing completely specified? [Operations, Spec §FR-078 to §FR-084]
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied
|
||||
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation
|
||||
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers
|
||||
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`
|
||||
- Add comments or findings inline
|
||||
- Link to relevant resources or documentation
|
||||
- Items are numbered sequentially (CHK001 to CHK059) for easy reference
|
||||
@@ -0,0 +1,54 @@
|
||||
# Specification Quality Checklist: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-08-23
|
||||
**Feature**: [spec.md](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Verification Notes for the 4 Precision Adjustments
|
||||
|
||||
1. **FR-057 Assertions do Promptfoo Completas (Doc 04 §22.2 e Doc 07 §11.3)**:
|
||||
- JSON Schema validation
|
||||
- Validadores Python customizados sem regex
|
||||
- IDs pertencentes ao contexto
|
||||
- Conjuntos e ordem esperada
|
||||
- URLs pertencentes à entrada
|
||||
- Precisão e recall
|
||||
- Enum e cardinalidade
|
||||
- Diffs com bibliotecas de sequência/Unicode sem regex
|
||||
- Comparação com a verdade de referência
|
||||
- Métricas de custo e latência
|
||||
2. **FR-073 Contrato das Tags na Golden Set**:
|
||||
- Atualizado para `accepted tags or a closed evaluation rubric`.
|
||||
3. **US9 (Cenário 4) e FR-084 / Runbook Matriz de Responsabilidades Exata**:
|
||||
- *Orchestrator*: provide article/ECP, control concurrency, and consume the manifest.
|
||||
- *Operations*: deploy, monitor, recover, and execute rollback.
|
||||
- *Engineering*: correct code, prompts, schemas, or integrations through the normal release process.
|
||||
- *Curator/Eval*: maintain the golden set and approve quality.
|
||||
4. **US2 Cenário 4 e FR-026 Desacoplamento entre Schema Inválido e Grounding Violation**:
|
||||
- `GROUNDING_VIOLATION` restrito a IDs, URLs, imagens ou conteúdo sem origem nos candidatos de entrada (invalida toda a resposta e aciona fallback).
|
||||
- Schema inválido tratado como falha semântica acionando fallback; esgotadas as opções de fallback, aplica-se `HYGIENE_FAILED`.
|
||||
@@ -0,0 +1,97 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/article-input.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "ArticleInputUnit",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"selected_extractor"
|
||||
],
|
||||
"properties": {
|
||||
"articles": false,
|
||||
"selected_extractor": {
|
||||
"type": "string",
|
||||
"enum": ["trafilatura", "newspaper4k", "readability"]
|
||||
},
|
||||
"crawled_url": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"error_message": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"extraction_status": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"http_status": {
|
||||
"type": ["integer", "null"]
|
||||
},
|
||||
"input_meta": {
|
||||
"type": ["object", "null"],
|
||||
"properties": {
|
||||
"url": { "type": ["string", "null"] },
|
||||
"titulo": { "type": ["string", "null"] },
|
||||
"subtitulo": { "type": ["string", "null"] },
|
||||
"quando_publicado": { "type": ["string", "null"] }
|
||||
},
|
||||
"additionalProperties": true
|
||||
},
|
||||
"page_title": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"trafilatura": {
|
||||
"type": ["object", "null"],
|
||||
"properties": {
|
||||
"title": { "type": ["string", "null"] },
|
||||
"author": { "type": ["string", "null"] },
|
||||
"date": { "type": ["string", "null"] },
|
||||
"description": { "type": ["string", "null"] },
|
||||
"text": { "type": ["string", "null"] },
|
||||
"markdown": { "type": ["string", "null"] },
|
||||
"canonical_url": { "type": ["string", "null"] },
|
||||
"image": { "type": ["string", "null"] },
|
||||
"language": { "type": ["string", "null"] },
|
||||
"sitename": { "type": ["string", "null"] },
|
||||
"categories": { "type": "array", "items": { "type": "string" } },
|
||||
"tags": { "type": "array", "items": { "type": "string" } },
|
||||
"raw_json": { "type": ["object", "string", "null"] },
|
||||
"pagetype": { "type": ["string", "null"] },
|
||||
"error": { "type": ["string", "null"] }
|
||||
},
|
||||
"additionalProperties": true
|
||||
},
|
||||
"newspaper4k": {
|
||||
"type": ["object", "null"],
|
||||
"properties": {
|
||||
"title": { "type": ["string", "null"] },
|
||||
"authors": { "type": "array", "items": { "type": "string" } },
|
||||
"publish_date": { "type": ["string", "null"] },
|
||||
"meta_description": { "type": ["string", "null"] },
|
||||
"text": { "type": ["string", "null"] },
|
||||
"article_html": { "type": ["string", "null"] },
|
||||
"canonical_link": { "type": ["string", "null"] },
|
||||
"top_image": { "type": ["string", "null"] },
|
||||
"images": { "type": "array", "items": { "type": "string" } },
|
||||
"meta_data": { "type": "object" },
|
||||
"meta_lang": { "type": ["string", "null"] },
|
||||
"meta_site_name": { "type": ["string", "null"] },
|
||||
"tags": { "type": "array", "items": { "type": "string" } },
|
||||
"keywords": { "type": "array", "items": { "type": "string" } },
|
||||
"error": { "type": ["string", "null"] }
|
||||
},
|
||||
"additionalProperties": true
|
||||
},
|
||||
"readability": {
|
||||
"type": ["object", "null"],
|
||||
"properties": {
|
||||
"title": { "type": ["string", "null"] },
|
||||
"short_title": { "type": ["string", "null"] },
|
||||
"author": { "type": ["string", "null"] },
|
||||
"cleaned_text": { "type": ["string", "null"] },
|
||||
"cleaned_html": { "type": ["string", "null"] },
|
||||
"error": { "type": ["string", "null"] }
|
||||
},
|
||||
"additionalProperties": true
|
||||
}
|
||||
},
|
||||
"additionalProperties": true
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/candidates-payload.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "CandidatesPayload",
|
||||
"description": "Normalized candidate payload supplied to LLM hygiene prompt (projection of internal CandidateObjects)",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"language",
|
||||
"selected_extractor",
|
||||
"metadata_candidates",
|
||||
"block_candidates",
|
||||
"link_candidates",
|
||||
"image_candidates"
|
||||
],
|
||||
"properties": {
|
||||
"language": {
|
||||
"type": "string"
|
||||
},
|
||||
"selected_extractor": {
|
||||
"type": "string",
|
||||
"enum": ["trafilatura", "newspaper4k", "readability"]
|
||||
},
|
||||
"metadata_candidates": {
|
||||
"type": "object",
|
||||
"required": ["title_candidates", "subtitle_candidates", "author_candidates"],
|
||||
"properties": {
|
||||
"title_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "source", "text"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"source": { "type": "string" },
|
||||
"text": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"subtitle_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "source", "text"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"source": { "type": "string" },
|
||||
"text": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"author_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "source", "text"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"source": { "type": "string" },
|
||||
"text": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"block_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "type", "order_index", "text", "source_extractor"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"type": {
|
||||
"type": "string",
|
||||
"enum": ["paragraph", "heading", "list_item", "quote"]
|
||||
},
|
||||
"order_index": { "type": "integer" },
|
||||
"text": { "type": "string" },
|
||||
"source_extractor": {
|
||||
"type": "string",
|
||||
"enum": ["trafilatura", "newspaper4k", "readability"]
|
||||
},
|
||||
"equivalences": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" }
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"link_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "url", "anchor_text", "parent_block_id"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"url": { "type": "string" },
|
||||
"anchor_text": { "type": "string" },
|
||||
"parent_block_id": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"image_candidates": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["candidate_id", "url"],
|
||||
"properties": {
|
||||
"candidate_id": { "type": "string" },
|
||||
"url": { "type": "string" },
|
||||
"alt": { "type": ["string", "null"] },
|
||||
"caption": { "type": ["string", "null"] },
|
||||
"parent_block_id": { "type": ["string", "null"] }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
@@ -0,0 +1,115 @@
|
||||
# CLI Interface Contract: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Date**: 2026-08-23
|
||||
**Contract Version**: `1.0.0`
|
||||
**Status**: Complete
|
||||
|
||||
---
|
||||
|
||||
## 1. Invocation Model
|
||||
|
||||
The runtime executes as a single-article ephemeral Python CLI command. It processes exactly one article input payload and one ECP profile snapshot per invocation, writing output artifacts and state atomically based on paths declared in the versioned configuration, and exiting cleanly with standard exit codes.
|
||||
|
||||
---
|
||||
|
||||
## 2. Command Synopsis
|
||||
|
||||
### 2.1 Main Consolidation Command
|
||||
|
||||
```bash
|
||||
python -m src.cli.consolidate \
|
||||
--input-article <path/to/article.json> \
|
||||
--ecp-snapshot <path/to/ecp_snapshot.json> \
|
||||
--config <path/to/runtime_config.json>
|
||||
```
|
||||
|
||||
### 2.2 Operational Commands
|
||||
|
||||
```bash
|
||||
# Preflight Validation
|
||||
python -m src.cli.preflight --config <path/to/runtime_config.json>
|
||||
|
||||
# Smoke Test
|
||||
python -m src.cli.smoke --config <path/to/runtime_config.json> --fixture <path/to/fixture.json>
|
||||
|
||||
# Reconcile State & Telemetry
|
||||
python -m src.cli.reconcile --config <path/to/runtime_config.json>
|
||||
|
||||
# Resend Pending Telemetry
|
||||
python -m src.cli.telemetry_flush --config <path/to/runtime_config.json>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Options & Arguments
|
||||
|
||||
### `src.cli.consolidate`
|
||||
|
||||
| Option | Flag | Type | Required | Description |
|
||||
|:---|:---|:---|:---|:---|
|
||||
| `--input-article` | `-i` | File Path | Yes | Path to single-article JSON input unit |
|
||||
| `--ecp-snapshot` | `-e` | File Path | Yes | Path to canonical ECP profile snapshot JSON |
|
||||
| `--config` | `-c` | File Path | Yes | Path to approved runtime configuration file |
|
||||
|
||||
### `src.cli.smoke`
|
||||
|
||||
| Option | Flag | Type | Required | Description |
|
||||
|:---|:---|:---|:---|:---|
|
||||
| `--config` | `-c` | File Path | Yes | Path to approved runtime configuration file |
|
||||
| `--fixture` | `-f` | File Path | Yes | Path to smoke test input fixture JSON |
|
||||
|
||||
### `src.cli.preflight` / `reconcile` / `telemetry_flush`
|
||||
|
||||
| Option | Flag | Type | Required | Description |
|
||||
|:---|:---|:---|:---|:---|
|
||||
| `--config` | `-c` | File Path | Yes | Path to approved runtime configuration file |
|
||||
|
||||
*Note on Paths & Options*: All persistence paths (output directory, SQLite database) belong strictly to the versioned configuration file to prevent discrepancies with the deterministic execution fingerprint.
|
||||
|
||||
---
|
||||
|
||||
## 4. Exit Codes
|
||||
|
||||
| Code | Meaning | Description |
|
||||
|:---|:---|:---|
|
||||
| `0` | Success / Handled Rejection | Article successfully processed (`completed_text` or `rejected_ecp`). |
|
||||
| `1` | Input Article / ECP Error | Input article or ECP schema failed local validation (`failed_validation`). |
|
||||
| `2` | Configuration / Preflight Error | Preflight verification failed, config hash mismatch, missing credentials, or invalid config file. |
|
||||
| `3` | Processing Failure | LLM hygiene, enrichment, or gateway fallback failed (`failed_processing`). |
|
||||
| `4` | Persistence Failure | File write, rename, or SQLite lock timeout failed. |
|
||||
|
||||
---
|
||||
|
||||
## 5. Output Protocol
|
||||
|
||||
### 5.1 stdout (Machine-Readable JSON)
|
||||
Every execution emits a single structured JSON object on stdout:
|
||||
|
||||
1. **`consolidate` (Fingerprint Established)**: Emits the complete manifest JSON complying with [`manifest-output.schema.json`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/contracts/manifest-output.schema.json).
|
||||
2. **`consolidate` (Article Validation Failure before Fingerprint)**: Emits a structured article failure JSON (Exit Code `1`):
|
||||
```json
|
||||
{
|
||||
"status": "failed_validation",
|
||||
"error_codes": ["INVALID_ARTICLE_SCHEMA"],
|
||||
"message": "Input article payload contains unparseable JSON or schema violation"
|
||||
}
|
||||
```
|
||||
3. **Configuration / Preflight Failure (Exit Code `2`)**: Emits a technical configuration envelope (without inventing article error codes):
|
||||
```json
|
||||
{
|
||||
"status": "configuration_error",
|
||||
"config_error_code": "CONFIG_HASH_MISMATCH",
|
||||
"message": "Runtime configuration SHA-256 does not match certified release-metadata.json"
|
||||
}
|
||||
```
|
||||
4. **Operational Commands**: Emit their specific structured JSON reports (Exit Code `0` on success, `2` on failure):
|
||||
- `preflight`: `{ "status": "ok", "config_version": "1.0.0", "certified_hash_match": true, "checks": [...] }`
|
||||
- `smoke`: `{ "status": "ok", "fingerprint": "...", "duration_ms": 124.5 }`
|
||||
- `reconcile`: `{ "status": "ok", "reconciled_articles": 0, "fixed_states": 0 }`
|
||||
- `telemetry_flush`: `{ "status": "ok", "flushed_events": 5, "remaining_pending": 0 }`
|
||||
|
||||
### 5.2 stderr (Sanitized Technical Logs)
|
||||
Emits structured JSON log lines containing timestamps, log levels, event codes, error details, and trace correlation IDs.
|
||||
- **Structural Sanitization**: Strips HTTP `Authorization`, `Proxy-Authorization`, `X-Api-Key` headers, and token query parameters.
|
||||
- **Exact Token Replacement**: Replaces exact string values of all loaded environment secrets (`GROQ_API_KEY`, `DEEPSEEK_API_KEY`, `LANGFUSE_SECRET_KEY`) with `[REDACTED]`.
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/ecp-snapshot.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "ECPSnapshotContract",
|
||||
"description": "Forwarding reference to the external canonical ECP Snapshot schema without local duplicate validation",
|
||||
"$ref": "https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json"
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/enrichment-response.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "ArticleSentimentTagsResponse",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"sentiment",
|
||||
"tags",
|
||||
"evidence_candidate_ids"
|
||||
],
|
||||
"properties": {
|
||||
"sentiment": {
|
||||
"type": "string",
|
||||
"enum": ["positive", "negative", "neutral"]
|
||||
},
|
||||
"tags": {
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 },
|
||||
"minItems": 3,
|
||||
"maxItems": 8,
|
||||
"uniqueItems": true
|
||||
},
|
||||
"evidence_candidate_ids": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"minItems": 1
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
@@ -0,0 +1,62 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/hygiene-response.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "ArticleContentHygieneResponse",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"title_candidate_id",
|
||||
"subtitle_candidate_id",
|
||||
"author_candidate_id",
|
||||
"kept_block_ids",
|
||||
"kept_link_ids",
|
||||
"kept_image_ids",
|
||||
"repairs"
|
||||
],
|
||||
"properties": {
|
||||
"title_candidate_id": {
|
||||
"type": "string"
|
||||
},
|
||||
"subtitle_candidate_id": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"author_candidate_id": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"kept_block_ids": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"minItems": 1,
|
||||
"uniqueItems": true
|
||||
},
|
||||
"kept_link_ids": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"uniqueItems": true
|
||||
},
|
||||
"kept_image_ids": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"uniqueItems": true
|
||||
},
|
||||
"repairs": {
|
||||
"$ref": "repair-operations.schema.json"
|
||||
},
|
||||
"removal_reasons": {
|
||||
"type": "object",
|
||||
"additionalProperties": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"advertisement",
|
||||
"recommendation",
|
||||
"navigation",
|
||||
"newsletter",
|
||||
"player_interface",
|
||||
"duplicate",
|
||||
"non_editorial"
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
@@ -0,0 +1,377 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/manifest-output.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "ArticleConsolidationManifest",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"schema_version",
|
||||
"fingerprint",
|
||||
"source_url",
|
||||
"selected_extractor",
|
||||
"final_status",
|
||||
"generate_markdown",
|
||||
"markdown_path",
|
||||
"markdown_hash",
|
||||
"ecp_classification",
|
||||
"enrichment",
|
||||
"provider_versions",
|
||||
"model_versions",
|
||||
"prompt_versions",
|
||||
"config_version",
|
||||
"trace_id",
|
||||
"error_codes"
|
||||
],
|
||||
"properties": {
|
||||
"schema_version": {
|
||||
"type": "string",
|
||||
"const": "1.0.0"
|
||||
},
|
||||
"fingerprint": {
|
||||
"type": "string",
|
||||
"minLength": 64,
|
||||
"maxLength": 64
|
||||
},
|
||||
"source_url": {
|
||||
"type": "string"
|
||||
},
|
||||
"selected_extractor": {
|
||||
"type": "string",
|
||||
"enum": ["trafilatura", "newspaper4k", "readability"]
|
||||
},
|
||||
"final_status": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"completed_text",
|
||||
"rejected_ecp",
|
||||
"failed_validation",
|
||||
"failed_processing"
|
||||
]
|
||||
},
|
||||
"generate_markdown": {
|
||||
"type": "boolean"
|
||||
},
|
||||
"markdown_path": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"markdown_hash": {
|
||||
"type": ["string", "null"],
|
||||
"minLength": 64,
|
||||
"maxLength": 64
|
||||
},
|
||||
"ecp_classification": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["category", "confidence", "rationale", "evidences"],
|
||||
"properties": {
|
||||
"category": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"DIRECT_INHERENT",
|
||||
"CONTEXTUAL_INHERENT",
|
||||
"TANGENTIAL",
|
||||
"NOT_RELATED"
|
||||
]
|
||||
},
|
||||
"confidence": {
|
||||
"type": "number",
|
||||
"minimum": 0.0,
|
||||
"maximum": 1.0
|
||||
},
|
||||
"rationale": {
|
||||
"type": "string"
|
||||
},
|
||||
"evidences": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"description": "Grounded textual fragments from intermediate Markdown"
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"enrichment": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["sentiment", "tags"],
|
||||
"properties": {
|
||||
"sentiment": {
|
||||
"type": "string",
|
||||
"enum": ["positive", "negative", "neutral"]
|
||||
},
|
||||
"tags": {
|
||||
"type": "array",
|
||||
"items": { "type": "string", "minLength": 1 },
|
||||
"minItems": 3,
|
||||
"maxItems": 8,
|
||||
"uniqueItems": true
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"provider_versions": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["hygiene", "enrichment"],
|
||||
"properties": {
|
||||
"hygiene": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["provider", "model", "role_config_version"],
|
||||
"properties": {
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"role_config_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"enrichment": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["provider", "model", "role_config_version"],
|
||||
"properties": {
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"role_config_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"model_versions": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["runtime_primary", "runtime_fallback"],
|
||||
"properties": {
|
||||
"runtime_primary": {
|
||||
"type": "object",
|
||||
"required": ["provider", "model", "role_config_version"],
|
||||
"properties": {
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"role_config_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"runtime_fallback": {
|
||||
"type": "object",
|
||||
"required": ["provider", "model", "role_config_version"],
|
||||
"properties": {
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"role_config_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"prompt_versions": {
|
||||
"type": ["object", "null"],
|
||||
"required": ["article_content_hygiene", "article_sentiment_tags"],
|
||||
"properties": {
|
||||
"article_content_hygiene": {
|
||||
"type": "object",
|
||||
"required": ["version", "hash"],
|
||||
"properties": {
|
||||
"version": { "type": "string" },
|
||||
"hash": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"article_sentiment_tags": {
|
||||
"type": "object",
|
||||
"required": ["version", "hash"],
|
||||
"properties": {
|
||||
"version": { "type": "string" },
|
||||
"hash": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"config_version": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"trace_id": {
|
||||
"type": ["string", "null"]
|
||||
},
|
||||
"error_codes": {
|
||||
"type": "array",
|
||||
"uniqueItems": true,
|
||||
"items": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"INVALID_ARTICLE_SCHEMA",
|
||||
"MISSING_SELECTED_EXTRACTOR",
|
||||
"INVALID_SELECTED_EXTRACTOR",
|
||||
"SELECTED_EXTRACTOR_UNAVAILABLE",
|
||||
"MISSING_SOURCE_URL",
|
||||
"MISSING_TITLE_CANDIDATE",
|
||||
"MISSING_CONTENT",
|
||||
"INVALID_ECP_SCHEMA",
|
||||
"HYGIENE_FAILED",
|
||||
"GROUNDING_VIOLATION",
|
||||
"INVALID_TEXT_REPAIR",
|
||||
"ECP_CLASSIFICATION_FAILED",
|
||||
"ECP_REJECTED",
|
||||
"ENRICHMENT_FAILED",
|
||||
"PERSISTENCE_FAILED",
|
||||
"TELEMETRY_PENDING"
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"allOf": [
|
||||
{
|
||||
"if": {
|
||||
"properties": { "final_status": { "const": "completed_text" } }
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"generate_markdown": { "const": true },
|
||||
"markdown_path": { "type": "string", "minLength": 1 },
|
||||
"markdown_hash": { "type": "string", "minLength": 64, "maxLength": 64 },
|
||||
"ecp_classification": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"category": { "enum": ["DIRECT_INHERENT", "CONTEXTUAL_INHERENT"] }
|
||||
}
|
||||
},
|
||||
"enrichment": { "type": "object" },
|
||||
"provider_versions": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"hygiene": { "type": "object" },
|
||||
"enrichment": { "type": "object" }
|
||||
}
|
||||
},
|
||||
"model_versions": { "type": "object" },
|
||||
"prompt_versions": { "type": "object" },
|
||||
"error_codes": {
|
||||
"type": "array",
|
||||
"items": { "const": "TELEMETRY_PENDING" },
|
||||
"maxItems": 1,
|
||||
"uniqueItems": true
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": { "final_status": { "const": "rejected_ecp" } }
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"generate_markdown": { "const": false },
|
||||
"markdown_path": { "type": "null" },
|
||||
"markdown_hash": { "type": "null" },
|
||||
"ecp_classification": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"category": { "enum": ["TANGENTIAL", "NOT_RELATED"] }
|
||||
}
|
||||
},
|
||||
"enrichment": { "type": "null" },
|
||||
"provider_versions": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"hygiene": { "type": "object" },
|
||||
"enrichment": { "type": "null" }
|
||||
}
|
||||
},
|
||||
"model_versions": { "type": "object" },
|
||||
"prompt_versions": { "type": "object" },
|
||||
"error_codes": {
|
||||
"type": "array",
|
||||
"items": { "enum": ["ECP_REJECTED", "TELEMETRY_PENDING"] },
|
||||
"contains": { "const": "ECP_REJECTED" },
|
||||
"minContains": 1,
|
||||
"maxContains": 1,
|
||||
"minItems": 1,
|
||||
"maxItems": 2,
|
||||
"uniqueItems": true
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": { "final_status": { "const": "failed_validation" } }
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"generate_markdown": { "const": false },
|
||||
"markdown_path": { "type": "null" },
|
||||
"markdown_hash": { "type": "null" },
|
||||
"provider_versions": { "type": "null" },
|
||||
"error_codes": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"enum": [
|
||||
"INVALID_ARTICLE_SCHEMA",
|
||||
"MISSING_SELECTED_EXTRACTOR",
|
||||
"INVALID_SELECTED_EXTRACTOR",
|
||||
"SELECTED_EXTRACTOR_UNAVAILABLE",
|
||||
"MISSING_SOURCE_URL",
|
||||
"MISSING_TITLE_CANDIDATE",
|
||||
"MISSING_CONTENT",
|
||||
"INVALID_ECP_SCHEMA",
|
||||
"TELEMETRY_PENDING"
|
||||
]
|
||||
},
|
||||
"contains": {
|
||||
"enum": [
|
||||
"INVALID_ARTICLE_SCHEMA",
|
||||
"MISSING_SELECTED_EXTRACTOR",
|
||||
"INVALID_SELECTED_EXTRACTOR",
|
||||
"SELECTED_EXTRACTOR_UNAVAILABLE",
|
||||
"MISSING_SOURCE_URL",
|
||||
"MISSING_TITLE_CANDIDATE",
|
||||
"MISSING_CONTENT",
|
||||
"INVALID_ECP_SCHEMA"
|
||||
]
|
||||
},
|
||||
"minItems": 1,
|
||||
"uniqueItems": true
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"if": {
|
||||
"properties": { "final_status": { "const": "failed_processing" } }
|
||||
},
|
||||
"then": {
|
||||
"properties": {
|
||||
"generate_markdown": { "const": false },
|
||||
"markdown_path": { "type": "null" },
|
||||
"markdown_hash": { "type": "null" },
|
||||
"error_codes": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"enum": [
|
||||
"HYGIENE_FAILED",
|
||||
"GROUNDING_VIOLATION",
|
||||
"INVALID_TEXT_REPAIR",
|
||||
"ECP_CLASSIFICATION_FAILED",
|
||||
"ENRICHMENT_FAILED",
|
||||
"PERSISTENCE_FAILED",
|
||||
"TELEMETRY_PENDING"
|
||||
]
|
||||
},
|
||||
"contains": {
|
||||
"enum": [
|
||||
"HYGIENE_FAILED",
|
||||
"GROUNDING_VIOLATION",
|
||||
"INVALID_TEXT_REPAIR",
|
||||
"ECP_CLASSIFICATION_FAILED",
|
||||
"ENRICHMENT_FAILED",
|
||||
"PERSISTENCE_FAILED"
|
||||
]
|
||||
},
|
||||
"minItems": 1,
|
||||
"uniqueItems": true
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"additionalProperties": false
|
||||
}
|
||||
@@ -0,0 +1,63 @@
|
||||
# Prompts Versioned Contract: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Date**: 2026-08-23
|
||||
**Contract Version**: `1.0.0`
|
||||
**Status**: Complete
|
||||
|
||||
---
|
||||
|
||||
## 1. Versioned Prompt Registry
|
||||
|
||||
The runtime manages exactly two atomic prompt contracts. Each prompt is versioned independently with an immutable semantic version and content hash.
|
||||
|
||||
| Prompt Identifier | File Path | Semantic Version | Input Context Structure | Expected Output Schema | Promptfoo Suite |
|
||||
|:---|:---|:---|:---|:---|:---|
|
||||
| `article_content_hygiene` | `prompts/article_content_hygiene.v1.txt` | `1.0.0` | 6-Block Context (`candidates-payload.schema.json`) | `hygiene-response.schema.json` | `evals/promptfoo.config.yaml` |
|
||||
| `article_sentiment_tags` | `prompts/article_sentiment_tags.v1.txt` | `1.0.0` | 6-Block Context (Title, subtitle, sanitized Markdown body, language, minimal ECP identity: `qid`, `canonical_name`) | `enrichment-response.schema.json` | `evals/promptfoo.config.yaml` |
|
||||
|
||||
---
|
||||
|
||||
## 2. Normative 6-Block Prompt Ordering (Doc 07 §5.3)
|
||||
|
||||
Every prompt sent to the Model Gateway MUST strictly follow the exact 6-block sequence defined in the normative specification:
|
||||
|
||||
1. **Bloco 1: Regras do Sistema (System Role & Policy Constraints)**
|
||||
Defines the agent/system persona, zero-hallucination mandate, and absolute prohibition of free-form text or Markdown invention.
|
||||
2. **Bloco 2: Responsabilidade da Chamada (Task Responsibility)**
|
||||
Declares the single, narrow responsibility of this specific logical invocation (candidate selection & micro-repair vs sentiment & tag extraction).
|
||||
3. **Bloco 3: Schema e Enums (Output Schema & Enums)**
|
||||
Provides the exact JSON schema and closed allowable categories/enums that the model MUST output.
|
||||
4. **Bloco 4: Contexto Estrutural (Structural Document Context)**
|
||||
Provides structural constraints, backbone extractor definition, metadata slots, and document language.
|
||||
5. **Bloco 5: Candidatos e Evidências (Candidates & Evidence Payloads)**
|
||||
Supplies the candidate blocks, links, images, and untrusted article content clearly delimited and separated from instructions (`<article_candidates>...</article_candidates>`).
|
||||
6. **Bloco 6: Pedido Final de Resposta Estruturada (Final Structured Response Directive)**
|
||||
Final closing directive commanding immediate output of the structured JSON response adhering strictly to the schema.
|
||||
|
||||
---
|
||||
|
||||
## 3. Strict Input Context Specifications
|
||||
|
||||
### 3.1 `article_content_hygiene` (FR-024, FR-058)
|
||||
Receives strictly the normalized candidate payload (`candidates-payload.schema.json`) structured across the 6 blocks above. Article text is treated as untrusted data, enclosed in explicit boundary delimiters.
|
||||
|
||||
### 3.2 `article_sentiment_tags` (FR-037, FR-059)
|
||||
Receives strictly the sanitized editorial context and minimal public ECP identity:
|
||||
- Editorial Document: Final title, subtitle (if present), and intermediate sanitized Markdown body;
|
||||
- Document Language;
|
||||
- Target Entity Identity: strictly `qid` and `canonical_name` (zero ad-hoc keywords, zero description or aliases lists, zero raw ECP snapshot);
|
||||
- Output JSON Schema (`enrichment-response.schema.json`).
|
||||
|
||||
---
|
||||
|
||||
## 4. Contractual Invariants
|
||||
|
||||
1. **Parity Guarantee**: The prompt files deployed in production runtime MUST be byte-for-byte identical (identical SHA-256 hash) to the prompt files evaluated by Promptfoo in `evals/`.
|
||||
2. **Preflight Verification**: During startup and preflight, the runtime recalculates the SHA-256 hash of each prompt file and asserts equality against the approved release metadata (`src/core/release-metadata.json`). Any mismatch aborts execution with exit code `2`.
|
||||
3. **Semantic Version Validation**: Semantic versions in prompt headers are parsed using standard semantic version parsers (`packaging.version` / pure tuple comparison), never regular expressions.
|
||||
4. **Structured Invariants**:
|
||||
- Zero free-form body generation instructions.
|
||||
- Zero historical references or self-healing instructions.
|
||||
- All article text inputs explicitly delimited as untrusted data.
|
||||
5. **Contract Testing**: `tests/contract/test_prompts_contract.py` validates prompt file existence, semver compliance via parser, SHA-256 calculation, schema associations, and byte parity with Promptfoo suites.
|
||||
@@ -0,0 +1,44 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/repair-operations.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "TextRepairOperations",
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": [
|
||||
"target_candidate_id",
|
||||
"original_fragment",
|
||||
"replacement_fragment",
|
||||
"category",
|
||||
"rationale"
|
||||
],
|
||||
"properties": {
|
||||
"target_candidate_id": {
|
||||
"type": "string"
|
||||
},
|
||||
"original_fragment": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"replacement_fragment": {
|
||||
"type": "string"
|
||||
},
|
||||
"category": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"encoding",
|
||||
"unicode",
|
||||
"spacing",
|
||||
"punctuation_corruption",
|
||||
"obvious_typo"
|
||||
]
|
||||
},
|
||||
"rationale": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,193 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://schemas.aftech.internal/article-consolidation/v1/runtime-config.schema.json",
|
||||
"x-contract-version": "1.0.0",
|
||||
"title": "RuntimeConfigContract",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"config_version",
|
||||
"paths",
|
||||
"roles",
|
||||
"prompts",
|
||||
"ecp",
|
||||
"limits",
|
||||
"pricing",
|
||||
"langfuse",
|
||||
"sqlite"
|
||||
],
|
||||
"properties": {
|
||||
"config_version": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"description": "Semantic version of the functional configuration; changes with any model/prompt/parameter adjustment and enters fingerprint"
|
||||
},
|
||||
"paths": {
|
||||
"type": "object",
|
||||
"required": ["output_dir", "sqlite_db"],
|
||||
"properties": {
|
||||
"output_dir": { "type": "string" },
|
||||
"sqlite_db": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"roles": {
|
||||
"type": "object",
|
||||
"required": ["runtime_primary", "runtime_fallback"],
|
||||
"properties": {
|
||||
"runtime_primary": {
|
||||
"type": "object",
|
||||
"required": [
|
||||
"role_config_version",
|
||||
"provider",
|
||||
"model",
|
||||
"endpoint_url",
|
||||
"timeout_seconds",
|
||||
"max_retries",
|
||||
"parameters",
|
||||
"hygiene_prompt_version",
|
||||
"hygiene_schema_version",
|
||||
"enrichment_prompt_version",
|
||||
"enrichment_schema_version"
|
||||
],
|
||||
"properties": {
|
||||
"role_config_version": { "type": "string" },
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"endpoint_url": { "type": "string", "format": "uri" },
|
||||
"timeout_seconds": { "type": "number", "minimum": 1 },
|
||||
"max_retries": { "type": "integer", "minimum": 0 },
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"additionalProperties": true
|
||||
},
|
||||
"hygiene_prompt_version": { "type": "string" },
|
||||
"hygiene_schema_version": { "type": "string" },
|
||||
"enrichment_prompt_version": { "type": "string" },
|
||||
"enrichment_schema_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"runtime_fallback": {
|
||||
"type": "object",
|
||||
"required": [
|
||||
"role_config_version",
|
||||
"provider",
|
||||
"model",
|
||||
"endpoint_url",
|
||||
"timeout_seconds",
|
||||
"max_retries",
|
||||
"parameters",
|
||||
"hygiene_prompt_version",
|
||||
"hygiene_schema_version",
|
||||
"enrichment_prompt_version",
|
||||
"enrichment_schema_version"
|
||||
],
|
||||
"properties": {
|
||||
"role_config_version": { "type": "string" },
|
||||
"provider": { "type": "string" },
|
||||
"model": { "type": "string" },
|
||||
"endpoint_url": { "type": "string", "format": "uri" },
|
||||
"timeout_seconds": { "type": "number", "minimum": 1 },
|
||||
"max_retries": { "type": "integer", "minimum": 0 },
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"additionalProperties": true
|
||||
},
|
||||
"hygiene_prompt_version": { "type": "string" },
|
||||
"hygiene_schema_version": { "type": "string" },
|
||||
"enrichment_prompt_version": { "type": "string" },
|
||||
"enrichment_schema_version": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"prompts": {
|
||||
"type": "object",
|
||||
"required": ["article_content_hygiene", "article_sentiment_tags"],
|
||||
"properties": {
|
||||
"article_content_hygiene": {
|
||||
"type": "object",
|
||||
"required": ["path", "version", "hash"],
|
||||
"properties": {
|
||||
"path": { "type": "string" },
|
||||
"version": { "type": "string" },
|
||||
"hash": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"article_sentiment_tags": {
|
||||
"type": "object",
|
||||
"required": ["path", "version", "hash"],
|
||||
"properties": {
|
||||
"path": { "type": "string" },
|
||||
"version": { "type": "string" },
|
||||
"hash": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"ecp": {
|
||||
"type": "object",
|
||||
"required": ["canonical_schema_reference", "classifier_module"],
|
||||
"properties": {
|
||||
"canonical_schema_reference": { "type": "string" },
|
||||
"classifier_module": { "type": "string" }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"limits": {
|
||||
"type": "object",
|
||||
"required": ["max_input_bytes", "context_strategy"],
|
||||
"properties": {
|
||||
"max_input_bytes": {
|
||||
"type": "integer",
|
||||
"minimum": 1024,
|
||||
"description": "Maximum byte size of raw input article JSON payload before candidate extraction"
|
||||
},
|
||||
"context_strategy": {
|
||||
"type": "string",
|
||||
"const": "fail_before_provider",
|
||||
"description": "Strategy applied when input exceeds limit; strictly fails before calling provider"
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"pricing": {
|
||||
"type": "object",
|
||||
"description": "Operational token pricing for cost reporting; does NOT alter functional execution fingerprint",
|
||||
"required": ["primary_input_1k", "primary_output_1k", "fallback_input_1k", "fallback_output_1k"],
|
||||
"properties": {
|
||||
"primary_input_1k": { "type": "number", "minimum": 0.0 },
|
||||
"primary_output_1k": { "type": "number", "minimum": 0.0 },
|
||||
"fallback_input_1k": { "type": "number", "minimum": 0.0 },
|
||||
"fallback_output_1k": { "type": "number", "minimum": 0.0 }
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"langfuse": {
|
||||
"type": "object",
|
||||
"required": ["environment", "trace_content_policy"],
|
||||
"properties": {
|
||||
"environment": { "type": "string" },
|
||||
"trace_content_policy": {
|
||||
"type": "string",
|
||||
"enum": ["metadata_only", "full_redacted"]
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
},
|
||||
"sqlite": {
|
||||
"type": "object",
|
||||
"required": ["busy_timeout_ms"],
|
||||
"properties": {
|
||||
"busy_timeout_ms": { "type": "integer", "minimum": 100 }
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
},
|
||||
"additionalProperties": false
|
||||
}
|
||||
@@ -0,0 +1,263 @@
|
||||
# Data Model: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Date**: 2026-08-23
|
||||
**Status**: Complete
|
||||
|
||||
---
|
||||
|
||||
## 1. Domain Entities & Relationships
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
ArticleInputUnit ||--o{ CandidateObject : extracts
|
||||
ArticleInputUnit ||--|| ECPSnapshot : references
|
||||
ArticleInputUnit ||--|| StateMachineRecord : tracks
|
||||
CandidateObject ||--o{ TextRepairOperation : receives
|
||||
StateMachineRecord ||--o| OutputManifest : persists
|
||||
StateMachineRecord ||--o| PublishedMarkdown : renders
|
||||
StateMachineRecord ||--o{ StateTransitionLog : logs
|
||||
StateMachineRecord ||--o{ TelemetryEvent : queues
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Entity Definitions
|
||||
|
||||
### 2.1 Article Input Unit (`article_input`)
|
||||
Represents the incoming single article JSON payload. Rejects explicit batch wrappers via `"articles": false`. Unknown fields in the input are preserved in the original object without alteration.
|
||||
|
||||
| Field | Type | Description | Required |
|
||||
|:---|:---|:---|:---|
|
||||
| `selected_extractor` | enum | `trafilatura` \| `newspaper4k` \| `readability` | Yes |
|
||||
| `crawled_url` | string (URL) \| null | Crawled URL | No |
|
||||
| `error_message` | string \| null | Upstream error message | No |
|
||||
| `extraction_status` | string \| null | Upstream status | No |
|
||||
| `http_status` | integer \| null | HTTP response status code | No |
|
||||
| `input_meta` | object \| null | Metadata map (`url`, `titulo`, `subtitulo`, `quando_publicado`) | No |
|
||||
| `page_title` | string \| null | Raw HTML page title | No |
|
||||
| `trafilatura` | object \| null | Trafilatura extraction output (accepts `raw_json` as object, string, or null) | No |
|
||||
| `newspaper4k` | object \| null | Newspaper4k extraction output | No |
|
||||
| `readability` | object \| null | Readability extraction output | No |
|
||||
| `articles` | false | Explicitly forbidden (batch wrapper rejection) | No |
|
||||
|
||||
*Note on Validation*: The runtime validates that the collective extraction sources provide at least one resolvable source URL, at least one non-empty candidate title, processable text, and usable content in `selected_extractor`. Total byte size is checked against `limits.max_input_bytes` before invoking remote providers (failing with `INVALID_ARTICLE_SCHEMA` if exceeded).
|
||||
|
||||
---
|
||||
|
||||
### 2.2 Entity Context Profile Snapshot (`ecp_snapshot`)
|
||||
Represents the complete, integral ECP Snapshot received by the runtime. The runtime validates the payload locally against the monorepo's canonical ECP schema (`src/adapters/ecp/schemas/ecp-profile.schema.json`) registered in `referencing.Registry` matching `$ref: "https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json"` without HTTP lookups. It extracts identity and version metadata (`qid`, `canonical_name`, `version`) for manifest and trace recording.
|
||||
|
||||
| Field | Type | Description | Required |
|
||||
|:---|:---|:---|:---|
|
||||
| *(opaque payload)* | object | Complete canonical ECP snapshot validated via `referencing.Registry` | Yes |
|
||||
| `qid` | string | Extracted canonical Wikidata / Entity QID (e.g. `Q148`) | Extracted |
|
||||
| `canonical_name` | string | Extracted entity canonical name | Extracted |
|
||||
| `version` | string | Extracted semantic version of the referenced ECP profile | Extracted |
|
||||
|
||||
---
|
||||
|
||||
### 2.3 Candidate Object (`candidate_object`)
|
||||
Single structural element extracted from the article payloads. Preserves all candidates without destructive deduplication.
|
||||
|
||||
| Field | Type | Description | Required |
|
||||
|:---|:---|:---|:---|
|
||||
| `candidate_id` | string | Opaque unique ID (e.g. `cand_blk_001`, `cand_title_001`) | Yes |
|
||||
| `type` | enum | `title` \| `subtitle` \| `author` \| `date` \| `paragraph` \| `heading` \| `list_item` \| `quote` \| `link` \| `image` | Yes |
|
||||
| `extractor_source` | enum | `trafilatura` \| `newspaper4k` \| `readability` \| `input_meta` \| `page_title` | Yes |
|
||||
| `source_field` | string | Origin field (e.g. `text`, `article_html`, `title`) | Yes |
|
||||
| `original_text_or_url` | string | Exact original text or URL content | Yes |
|
||||
| `structural_representation` | string | Markdown/HTML/AST structural snippet | Yes |
|
||||
| `order_index` | integer | Position index in extractor backbone | Yes |
|
||||
| `parent_candidate_id` | string \| null | ID of parent element (for nested list items, blockquotes, etc.) | No |
|
||||
| `equivalences` | array[string] | List of candidate IDs representing equivalent content from other extractors | Yes |
|
||||
| `content_hash` | string (64-char hex) | Deterministic content hash | Yes |
|
||||
| `structural_flags` | object | Purely structural flags (e.g. `{"heading_level": 2}`) | Yes |
|
||||
|
||||
*Note on Projection*: The LLM prompt receives `CandidatesPayload`, which is a clean, normalized projection of these internal `CandidateObject` instances.
|
||||
|
||||
---
|
||||
|
||||
### 2.4 Text Repair Operation (`text_repair`)
|
||||
Micro-repair proposed by the LLM and validated by the harness.
|
||||
|
||||
| Field | Type | Description | Required |
|
||||
|:---|:---|:---|:---|
|
||||
| `target_candidate_id` | string | ID of the target block or metadata candidate | Yes |
|
||||
| `original_fragment` | string | Exact substring in candidate to replace | Yes |
|
||||
| `replacement_fragment` | string | Validated replacement text | Yes |
|
||||
| `category` | enum | `encoding` \| `unicode` \| `spacing` \| `punctuation_corruption` \| `obvious_typo` | Yes |
|
||||
| `rationale` | string | Short explanation | Yes |
|
||||
| `is_accepted` | boolean | Validation outcome from harness | Yes |
|
||||
| `rejection_reason` | string \| null | Code if rejected (`SENSITIVE_ENTITY`, `AMBIGUOUS_TARGET`, `OUT_OF_CATEGORY`, etc.) | No |
|
||||
|
||||
---
|
||||
|
||||
### 2.5 State Machine Record (`state_record`)
|
||||
SQLite table `article_states` storing runtime execution status.
|
||||
|
||||
| Column | SQLite Type | Description |
|
||||
|:---|:---|:---|
|
||||
| `fingerprint` | TEXT (PK, 64-char hex) | Deterministic content hash of the execution |
|
||||
| `source_url` | TEXT | Resolved source URL |
|
||||
| `selected_extractor` | TEXT | Extractor used as backbone |
|
||||
| `current_state` | TEXT | `received` \| `validated` \| `content_cleaned` \| `ecp_approved` \| `ecp_rejected` \| `enriched` \| `completed_text` \| `failed` |
|
||||
| `final_status` | TEXT | `completed_text` \| `rejected_ecp` \| `failed_validation` \| `failed_processing` \| NULL |
|
||||
| `generate_markdown` | INTEGER | 1 if Markdown generated, 0 otherwise |
|
||||
| `markdown_path` | TEXT | Path to generated `.md` file (or NULL) |
|
||||
| `manifest_path` | TEXT | Path to generated `.result.json` file |
|
||||
| `markdown_hash` | TEXT (64-char hex) | Content hash of generated `.md` file (or NULL) |
|
||||
| `manifest_hash` | TEXT (64-char hex) | Content hash of generated `.result.json` file |
|
||||
| `ecp_category` | TEXT | `DIRECT_INHERENT` \| `CONTEXTUAL_INHERENT` \| `TANGENTIAL` \| `NOT_RELATED` \| NULL |
|
||||
| `ecp_confidence` | REAL | Confidence score (0.0 to 1.0) |
|
||||
| `functional_versions_json` | TEXT (JSON) | Consolidated versions of contracts, ECP reference, config, prompts, and models |
|
||||
| `trace_id` | TEXT | Langfuse trace identifier |
|
||||
| `terminal_error_code` | TEXT | Normative error code if failed |
|
||||
| `error_metadata_json` | TEXT (JSON) | Sanitized minimal error metadata (stack trace in technical log only) |
|
||||
| `created_at` | TEXT (ISO 8601) | Timestamp of ingestion |
|
||||
| `updated_at` | TEXT (ISO 8601) | Timestamp of last transition |
|
||||
|
||||
---
|
||||
|
||||
### 2.6 State Transition Log (`state_transitions`)
|
||||
SQLite table `state_transitions` tracking execution lifecycle.
|
||||
|
||||
| Column | SQLite Type | Description |
|
||||
|:---|:---|:---|
|
||||
| `id` | INTEGER (PK AUTO) | Unique transition ID |
|
||||
| `fingerprint` | TEXT (FK) | Reference to `article_states.fingerprint` |
|
||||
| `from_state` | TEXT | Starting state |
|
||||
| `to_state` | TEXT | Destination state |
|
||||
| `start_time` | TEXT (ISO 8601) | Transition start timestamp |
|
||||
| `end_time` | TEXT (ISO 8601) | Transition end timestamp |
|
||||
| `duration_ms` | REAL | Elapsed milliseconds |
|
||||
| `result` | TEXT | `success` \| `failure` \| `skipped` |
|
||||
| `metadata_json` | TEXT (JSON) | Transition context metadata |
|
||||
|
||||
---
|
||||
|
||||
### 2.7 Telemetry Event (`pending_telemetry`)
|
||||
SQLite table `pending_telemetry` for resilient deferred delivery to Langfuse when the network/service is unreachable.
|
||||
|
||||
| Column | SQLite Type | Description |
|
||||
|:---|:---|:---|
|
||||
| `event_id` | TEXT (PK) | UUID / Unique event ID |
|
||||
| `fingerprint` | TEXT | Associated article fingerprint |
|
||||
| `trace_id` | TEXT | Associated trace ID |
|
||||
| `event_type` | TEXT | `span` \| `generation` \| `score` |
|
||||
| `payload_json` | TEXT (JSON) | Sanitized telemetry event payload |
|
||||
| `created_at` | TEXT (ISO 8601) | Creation timestamp |
|
||||
| `retry_count` | INTEGER | Number of transmission attempts |
|
||||
| `last_error` | TEXT | Last error message |
|
||||
|
||||
---
|
||||
|
||||
### 2.8 Release Metadata Contract (`src/core/release-metadata.json`)
|
||||
Immutable packaged artifact recording certified configurations for preflight verification. `runtime_config_sha256` is strictly calculated as the exact file byte SHA-256 hash (`hashlib.sha256(Path(config_path).read_bytes()).hexdigest()`).
|
||||
|
||||
```json
|
||||
{
|
||||
"release_version": "1.0.0",
|
||||
"runtime_config_sha256": "06a2769f7aa3a15a1e61880f171ecebc0d29094ab2499616243e59c0aecf340f",
|
||||
"prompts_hashes": {
|
||||
"article_content_hygiene": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0",
|
||||
"article_sentiment_tags": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0"
|
||||
},
|
||||
"schemas_versions": {
|
||||
"article_input": "1.0.0",
|
||||
"ecp_snapshot": "1.0.0",
|
||||
"runtime_config": "1.0.0",
|
||||
"candidates_payload": "1.0.0",
|
||||
"hygiene_response": "1.0.0",
|
||||
"repair_operations": "1.0.0",
|
||||
"enrichment_response": "1.0.0",
|
||||
"manifest_output": "1.0.0"
|
||||
},
|
||||
"certified_models": {
|
||||
"runtime_primary": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
|
||||
"runtime_fallback": { "provider": "deepseek", "model": "deepseek-v4-flash", "role_config_version": "1.0.0" }
|
||||
},
|
||||
"ecp_classifier_config_hash": "a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 2.9 Output Manifest (`<fingerprint>.result.json`)
|
||||
Structure of the machine-readable output manifest complying with `manifest-output.schema.json`.
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "1.0.0",
|
||||
"fingerprint": "a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890",
|
||||
"source_url": "https://example.com/noticia-123",
|
||||
"selected_extractor": "trafilatura",
|
||||
"final_status": "completed_text",
|
||||
"generate_markdown": true,
|
||||
"markdown_path": "out/a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890.md",
|
||||
"markdown_hash": "06a2769f7aa3a15a1e61880f171ecebc0d29094ab2499616243e59c0aecf340f",
|
||||
"ecp_classification": {
|
||||
"category": "DIRECT_INHERENT",
|
||||
"confidence": 0.95,
|
||||
"rationale": "Article directly analyzes the economic policy of the entity.",
|
||||
"evidences": ["trecho textual fundamentado 1", "trecho textual fundamentado 2"]
|
||||
},
|
||||
"enrichment": {
|
||||
"sentiment": "positive",
|
||||
"tags": ["Economia", "Política Monetária", "Inflação"]
|
||||
},
|
||||
"provider_versions": {
|
||||
"hygiene": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
|
||||
"enrichment": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" }
|
||||
},
|
||||
"model_versions": {
|
||||
"runtime_primary": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
|
||||
"runtime_fallback": { "provider": "deepseek", "model": "deepseek-v4-flash", "role_config_version": "1.0.0" }
|
||||
},
|
||||
"prompt_versions": {
|
||||
"article_content_hygiene": { "version": "1.0.0", "hash": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0" },
|
||||
"article_sentiment_tags": { "version": "1.0.0", "hash": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0" }
|
||||
},
|
||||
"config_version": "1.0.0",
|
||||
"trace_id": "trace_run_20260823_001",
|
||||
"error_codes": []
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 2.10 Published Markdown (`<fingerprint>.md`)
|
||||
Format of the rendered Markdown document with YAML front matter.
|
||||
|
||||
```markdown
|
||||
---
|
||||
title: "Título Principal do Artigo Publicado"
|
||||
subtitle: "Subtítulo editorial detalhado"
|
||||
author: "Nome do Autor"
|
||||
published_at: "2026-08-23T14:00:00Z"
|
||||
source_url: "https://example.com/noticia-123"
|
||||
sentiment: positive
|
||||
tags:
|
||||
- Economia
|
||||
- Política Monetária
|
||||
- Inflação
|
||||
ecp_qid: "Q148"
|
||||
ecp_canonical_name: "Entidade Alvo"
|
||||
ecp_category: DIRECT_INHERENT
|
||||
ecp_confidence: 0.95
|
||||
---
|
||||
|
||||
# Título Principal do Artigo Publicado
|
||||
|
||||
*Subtítulo editorial detalhado*
|
||||
|
||||
Primeiro parágrafo do artigo com [link grounded](https://example.com/referencia) e texto limpo.
|
||||
|
||||
## Intertítulo Editorial
|
||||
|
||||
Segundo parágrafo contendo citação textual sem alterações indevidas.
|
||||
|
||||

|
||||
|
||||
Parágrafo de encerramento sem notas de rodapé publicitárias ou chamadas de redes sociais.
|
||||
```
|
||||
@@ -0,0 +1,199 @@
|
||||
# Implementation Plan: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Branch**: `006-article-consolidation-runtime` | **Date**: 2026-08-23 | **Spec**: [`specs/006-article-consolidation-runtime/spec.md`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/spec.md)
|
||||
|
||||
**Input**: Feature specification from `specs/006-article-consolidation-runtime/spec.md`
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
The Article Consolidation and Hygiene Runtime is an ephemeral, deterministic Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot. It coordinates extraction payloads from three extractors (`trafilatura`, `newspaper4k`, `readability`), performs 100% LLM extractive hygiene via certified low-cost models, validates grounded candidate selections and micro-repairs without regular expressions or manual keyword dictionaries, applies an ECP relevance gate via `src.classifier.InherenceClassifier`, adds entity-relative sentiment and native-language tags, renders canonical Markdown with YAML front matter, persists machine-readable manifests and state atomically in SQLite (WAL), and transmits sanitized observability telemetry directly to Langfuse.
|
||||
|
||||
---
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python >= 3.10 (strictly matching `requires-python` in `pyproject.toml`)
|
||||
**Primary Dependencies**:
|
||||
- Standard library (`json`, `urllib.parse`, `unicodedata`, `difflib`, `sqlite3`, `pathlib`, `hashlib`, `typing`, `signal`)
|
||||
- `beautifulsoup4` (DOM parsing)
|
||||
- `marko` (CommonMark AST parsing)
|
||||
- `jsonschema` + `referencing` (JSON Schema Draft 2020-12 validation with immutable local schema registry)
|
||||
- `python-dateutil` (ISO 8601 date parsing)
|
||||
- `pyyaml` (Safe YAML front matter serialization)
|
||||
- `httpx` (HTTP client for Model Gateway)
|
||||
- `langfuse` (Observability SDK >= 4.7)
|
||||
- Existing monorepo modules (`src.language`, `src.classifier.InherenceClassifier`)
|
||||
|
||||
**Storage**: SQLite 3 (WAL mode, configurable busy timeout, short transactions, native backup API) + Local filesystem (atomic temporary files and renames)
|
||||
**Testing**: `pytest` (unit, contract, mock integration, security, load, fault injection, operations resilience), static zero-regex multi-parser analyzer (Python AST + JSON pattern check + Promptfoo YAML check), `promptfoo` (offline prompt evaluations in dev/CI)
|
||||
**Target Platform**: Ambiente suportado pelo repositório e pelo deployment pipeline
|
||||
**Project Type**: Python CLI Tool (Ephemeral process, no API, no internal worker pool)
|
||||
**Performance Goals**: Sustained throughput >= 100 articles/hour under external orchestrator concurrency; p50, p95, and p99 latency and cost empirically measured and approved in staging before go-live
|
||||
**Constraints**: Zero regular expressions (`re`) in text processing, schemas, and assertions; zero manual keyword dictionaries; zero expensive/powerful models in runtime roles (or internal ECP classifier); 11 critical release invariants (all 0); secret exposure = 0; strict 10-step hygiene harness; input size limit enforcement (`INVALID_ARTICLE_SCHEMA` on overflow); no generic unused abstractions
|
||||
**Scale/Scope**: 1 article per CLI invocation, 9 versioned contracts, 10 fault injection scenarios, 8 security scenarios (SEC-001 to SEC-008), operational resilience testing (FR-081), complete golden set with holdout, contract tests over all 20 real reference units
|
||||
|
||||
---
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Principle / Rule | Compliance Status | Description |
|
||||
|:---|:---|:---|
|
||||
| **I. Simplicity Mandate (FR-001)** | **PASS** | Minimum sufficient code, no generic unused abstractions, small responsibilities share modules, zero redundant local metrics system. |
|
||||
| **II. Strict Dependency Policy (FR-002)** | **PASS** | Standard library prioritized; 7-factor qualitative evaluation completed; no agent frameworks, no trivial libraries, no duplicate ECP schema, no custom regex parsers. |
|
||||
| **III. CLI Ephemeral Interface (FR-003)** | **PASS** | Python CLI processing 1 article per invocation, no internal batch loops, no internal worker pools, external concurrency. |
|
||||
| **IV. Regex Prohibition (FR-015, FR-070)** | **PASS** | Multi-parser static check in CI enforces zero `re` calls in Python AST, zero `pattern` keys in JSON schemas, and zero regex assertions in Promptfoo YAML. |
|
||||
| **V. Agnostic Gateway & Cheap Models (FR-041, FR-042)** | **PASS** | Logical roles `runtime_primary` and `runtime_fallback` limited to certified cheap models; zero powerful models in runtime or internal classifier. |
|
||||
| **VI. Zero Hallucination Grounding (FR-023, FR-026)** | **PASS** | LLM returns only candidate IDs and bounded diffs; harness strictly enforces candidate grounding and reverses ungrounded micro-repairs. |
|
||||
| **VII. Atomic Persistence & Idempotency (FR-009, FR-050)** | **PASS** | SQLite WAL mode + file atomic renames; deterministic fingerprinting; hash-based crash reconciliation. |
|
||||
|
||||
---
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation & 9 Versioned Contracts (this feature)
|
||||
|
||||
```text
|
||||
specs/006-article-consolidation-runtime/
|
||||
├── spec.md # Feature specification (v1.0.0)
|
||||
├── plan.md # Implementation plan (this file)
|
||||
├── research.md # Technical research & decisions (Phase 0)
|
||||
├── data-model.md # Entity definitions, SQLite tables & formats (Phase 1)
|
||||
├── quickstart.md # Runnable verification guide (Phase 1)
|
||||
├── contracts/ # All 9 independent versioned contracts (Phase 1)
|
||||
│ ├── article-input.schema.json # Contract 1: Article Input Unit (v1.0.0)
|
||||
│ ├── ecp-snapshot.schema.json # Contract 2: ECP Snapshot Canonical Reference (v1.0.0)
|
||||
│ ├── runtime-config.schema.json # Contract 3: Runtime Configuration (v1.0.0)
|
||||
│ ├── candidates-payload.schema.json # Contract 4: Candidate Payload to LLM (v1.0.0)
|
||||
│ ├── hygiene-response.schema.json # Contract 5: Hygiene LLM Response (v1.0.0)
|
||||
│ ├── repair-operations.schema.json # Contract 6: Repair Operations Schema (v1.0.0)
|
||||
│ ├── enrichment-response.schema.json # Contract 7: Enrichment LLM Response (v1.0.0)
|
||||
│ ├── manifest-output.schema.json # Contract 8: Output Manifest (v1.0.0)
|
||||
│ ├── prompts-contract.md # Contract 9: Versioned Prompts Contract (v1.0.0)
|
||||
│ └── cli-interface.md # Interface: CLI Command Interface (v1.0.0)
|
||||
└── checklists/
|
||||
└── requirements.md # Quality checklist
|
||||
```
|
||||
|
||||
### Source Code & Fixtures (repository layout)
|
||||
|
||||
```text
|
||||
src/
|
||||
├── __init__.py
|
||||
├── cli/
|
||||
│ ├── __init__.py
|
||||
│ ├── consolidate.py # Main single-article CLI entry point (--input-article, --ecp-snapshot, --config) with SIGTERM/SIGINT graceful shutdown
|
||||
│ ├── preflight.py # Preflight validation CLI (--config verified against src/core/release-metadata.json)
|
||||
│ ├── smoke.py # Smoke test CLI (--config, --fixture)
|
||||
│ ├── reconcile.py # State & artifact reconciliation CLI (--config)
|
||||
│ └── telemetry_flush.py # Deferred telemetry flush CLI (--config)
|
||||
├── core/
|
||||
│ ├── __init__.py
|
||||
│ ├── config.py # Runtime configuration loading & SHA-256 release metadata verification (exact byte hash)
|
||||
│ ├── limits.py # Input size limit validation (FR-056)
|
||||
│ ├── fingerprint.py # Deterministic fingerprint calculator
|
||||
│ ├── state_machine.py # Python explicit state machine & SQLite logger
|
||||
│ └── release-metadata.json # Packaged release metadata (SHA-256 hashes of config, prompts, schemas, ECP config)
|
||||
├── candidate/
|
||||
│ ├── __init__.py
|
||||
│ ├── parser.py # DOM, CommonMark AST, JSON-LD structural parsing
|
||||
│ ├── equivalence.py # Backbone ordering & equivalence mapping (no deletion)
|
||||
│ └── models.py # CandidateObject definitions
|
||||
├── hygiene/
|
||||
│ ├── __init__.py
|
||||
│ ├── harness.py # 10-step validation harness
|
||||
│ ├── repairs.py # Controlled micro-repair validator (difflib/unicodedata)
|
||||
│ └── assembler.py # Grounded intermediate Markdown assembler
|
||||
├── ecp/
|
||||
│ ├── __init__.py
|
||||
│ └── adapter.py # ECP adapter invoking src.classifier.InherenceClassifier & referencing.Registry local schema
|
||||
├── enrichment/
|
||||
│ ├── __init__.py
|
||||
│ └── harness.py # Entity sentiment & native tags validator
|
||||
├── gateway/
|
||||
│ ├── __init__.py
|
||||
│ ├── client.py # Agnostic Model Gateway (primary/fallback)
|
||||
│ └── adapters.py # Provider-specific minimal HTTP adapters (Groq, DeepSeek)
|
||||
├── storage/
|
||||
│ ├── __init__.py
|
||||
│ ├── sqlite_store.py # SQLite WAL state store, transition logs, and native backup/restore API
|
||||
│ ├── file_store.py # Atomic filesystem writer (temp + rename)
|
||||
│ └── markdown_renderer.py # Canonical YAML front matter & body renderer
|
||||
└── observability/
|
||||
├── __init__.py
|
||||
├── langfuse_tracer.py # Spans, generations, metrics, scores & secret redaction
|
||||
└── structured_logger.py # Sanitized JSON logger
|
||||
|
||||
prompts/
|
||||
├── article_content_hygiene.v1.txt
|
||||
└── article_sentiment_tags.v1.txt
|
||||
|
||||
evals/
|
||||
├── promptfoo.config.yaml # Promptfoo test suite configuration
|
||||
├── golden_set/ # Full golden dataset with reference truths
|
||||
├── holdout/ # Holdout dataset (never used in few-shot)
|
||||
└── reference_20/ # 20 reference regression cases
|
||||
|
||||
examples/
|
||||
├── sample_article_valid.json # Executable fixture: valid inherent article unit
|
||||
├── sample_article_tangential.json # Executable fixture: non-inherent article unit
|
||||
└── sample_ecp_snapshot.json # Executable fixture: canonical ECP snapshot
|
||||
|
||||
runtime_config.local.json # Executable local validation configuration fixture
|
||||
|
||||
tests/
|
||||
├── contract/ # Contract tests for all 9 versioned contracts (evaluated over 20 real reference units)
|
||||
│ ├── test_article_input_contract.py
|
||||
│ ├── test_ecp_snapshot_contract.py
|
||||
│ ├── test_runtime_config_contract.py
|
||||
│ ├── test_candidates_payload_contract.py
|
||||
│ ├── test_hygiene_response_contract.py
|
||||
│ ├── test_repair_operations_contract.py
|
||||
│ ├── test_enrichment_response_contract.py
|
||||
│ ├── test_manifest_output_contract.py
|
||||
│ └── test_prompts_contract.py
|
||||
├── unit/
|
||||
│ ├── test_fingerprint.py
|
||||
│ ├── test_input_limits.py
|
||||
│ ├── test_candidate_parser.py
|
||||
│ ├── test_equivalence_mapping.py
|
||||
│ ├── test_hygiene_harness.py
|
||||
│ ├── test_repairs_validator.py
|
||||
│ ├── test_ecp_adapter.py
|
||||
│ ├── test_enrichment_harness.py
|
||||
│ ├── test_model_gateway.py
|
||||
│ ├── test_sqlite_store.py
|
||||
│ └── test_file_store.py
|
||||
├── integration/
|
||||
│ ├── test_e2e_pipeline_mock.py
|
||||
│ ├── test_idempotency_concurrency.py
|
||||
│ └── test_operations_resilience.py # Backup/restore, graceful shutdown, signals, pending telemetry, rollback, credential rotation, certified model rotation, and uncertified rotation rejection (FR-081)
|
||||
├── security/
|
||||
│ └── test_security_scenarios.py # Parametrized tests for SEC-001 to SEC-008
|
||||
├── load/
|
||||
│ └── test_load_100_art_per_hour.py # 100 articles/hour staging load benchmark
|
||||
├── fault_injection/
|
||||
│ └── test_fault_injection_scenarios.py # 10 normative fault scenarios (FLT-001 to FLT-010)
|
||||
└── scripts/
|
||||
└── check_zero_regex.py # Multi-parser static check (Python AST + JSON pattern + Promptfoo YAML)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
> **No architectural violations detected. All modules adhere strictly to simplicity and dependency rules.**
|
||||
|
||||
| Component | Design Choice | Simplicity Justification |
|
||||
|:---|:---|:---|
|
||||
| Orchestration | Explicit Python state machine | Eliminates LangGraph / LangChain overhead while providing full deterministic lifecycle tracking. |
|
||||
| Model Gateway | Direct 2-adapter primary/fallback | Eliminates smart router complexity while guaranteeing cheap model enforcement and fallback. |
|
||||
| Candidate Model | Unified `CandidateObject` dictionary/dataclass | Avoids class hierarchies for candidate types while capturing all required structural flags. |
|
||||
| Persistence | SQLite (WAL) + Local Filesystem | Eliminates distributed database/queue dependencies; provides ACID state within local process boundaries. |
|
||||
| ECP Integration | Direct invocation of `src.classifier.InherenceClassifier` | Leverages existing monorepo module without creating redundant remote microservices. |
|
||||
| Observability | Native Langfuse dashboards + SQLite outage queue | Eliminates redundant local metric systems while fulfilling all reporting requirements. |
|
||||
| Operations | Native SQLite backup API + Signal handling in CLI | Fulfills FR-081 requirements using Python standard library without external daemons. |
|
||||
| Regex Prohibition | `unicodedata`, `difflib`, NLP tokenizers | Multi-parser validation ensures zero regex across code, schemas, and evals without runtime overhead. |
|
||||
@@ -0,0 +1,202 @@
|
||||
# Quickstart Guide: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Date**: 2026-08-23
|
||||
**Status**: Complete
|
||||
|
||||
---
|
||||
|
||||
## 1. Prerequisites & Environment Setup
|
||||
|
||||
The runtime is built using the repository's standard Python packaging defined in `pyproject.toml` (`requires-python = ">=3.10"`).
|
||||
|
||||
### Installation
|
||||
|
||||
```bash
|
||||
# Create and activate virtual environment
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate # ou .venv\Scripts\activate
|
||||
|
||||
# Install package and dependencies in editable mode
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
### Environment Secrets
|
||||
|
||||
Set provider credentials and Langfuse secrets via environment variables:
|
||||
|
||||
```bash
|
||||
# Model Gateway Provider Keys
|
||||
export GROQ_API_KEY="gsk_samplekey1234567890abcdef"
|
||||
export DEEPSEEK_API_KEY="sk_samplekey1234567890abcdef"
|
||||
|
||||
# Observability
|
||||
export LANGFUSE_PUBLIC_KEY="pk-lf-samplepublickey123456"
|
||||
export LANGFUSE_SECRET_KEY="sk-lf-samplesecretkey123456"
|
||||
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Local Validation Configuration (Fixture Only)
|
||||
|
||||
> [!NOTE]
|
||||
> The configuration below is a local development fixture (`runtime_config.local.json`), certified against packaged `src/core/release-metadata.json`. Timeouts, retries, pricing, and limits shown here are exclusive to local testing fixtures and do not represent production calibrated defaults.
|
||||
|
||||
```json
|
||||
{
|
||||
"config_version": "1.0.0",
|
||||
"paths": {
|
||||
"output_dir": "./out",
|
||||
"sqlite_db": "./out/runtime_state.db"
|
||||
},
|
||||
"roles": {
|
||||
"runtime_primary": {
|
||||
"role_config_version": "1.0.0",
|
||||
"provider": "groq",
|
||||
"model": "openai/gpt-oss-20b",
|
||||
"endpoint_url": "https://api.groq.com/openai/v1/chat/completions",
|
||||
"timeout_seconds": 15,
|
||||
"max_retries": 2,
|
||||
"parameters": { "temperature": 0.0 },
|
||||
"hygiene_prompt_version": "1.0.0",
|
||||
"hygiene_schema_version": "1.0.0",
|
||||
"enrichment_prompt_version": "1.0.0",
|
||||
"enrichment_schema_version": "1.0.0"
|
||||
},
|
||||
"runtime_fallback": {
|
||||
"role_config_version": "1.0.0",
|
||||
"provider": "deepseek",
|
||||
"model": "deepseek-v4-flash",
|
||||
"endpoint_url": "https://api.deepseek.com/v1/chat/completions",
|
||||
"timeout_seconds": 15,
|
||||
"max_retries": 2,
|
||||
"parameters": { "temperature": 0.0 },
|
||||
"hygiene_prompt_version": "1.0.0",
|
||||
"hygiene_schema_version": "1.0.0",
|
||||
"enrichment_prompt_version": "1.0.0",
|
||||
"enrichment_schema_version": "1.0.0"
|
||||
}
|
||||
},
|
||||
"prompts": {
|
||||
"article_content_hygiene": {
|
||||
"path": "prompts/article_content_hygiene.v1.txt",
|
||||
"version": "1.0.0",
|
||||
"hash": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0"
|
||||
},
|
||||
"article_sentiment_tags": {
|
||||
"path": "prompts/article_sentiment_tags.v1.txt",
|
||||
"version": "1.0.0",
|
||||
"hash": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0"
|
||||
}
|
||||
},
|
||||
"ecp": {
|
||||
"canonical_schema_reference": "src/adapters/ecp/schemas/ecp-profile.schema.json",
|
||||
"classifier_module": "src.classifier.InherenceClassifier"
|
||||
},
|
||||
"limits": {
|
||||
"max_input_bytes": 5242880,
|
||||
"context_strategy": "fail_before_provider"
|
||||
},
|
||||
"pricing": {
|
||||
"primary_input_1k": 0.0001,
|
||||
"primary_output_1k": 0.0002,
|
||||
"fallback_input_1k": 0.0001,
|
||||
"fallback_output_1k": 0.0002
|
||||
},
|
||||
"langfuse": {
|
||||
"environment": "local_dev",
|
||||
"trace_content_policy": "metadata_only"
|
||||
},
|
||||
"sqlite": {
|
||||
"busy_timeout_ms": 5000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Operational Execution Scenarios
|
||||
|
||||
### Scenario A: Preflight Validation
|
||||
Validates local configuration file exact byte hash against packaged `src/core/release-metadata.json`, prompt file hashes, schema compatibility, filesystem permissions, clock sync, and validated credentials before accepting live traffic.
|
||||
|
||||
```bash
|
||||
python -m src.cli.preflight --config runtime_config.local.json
|
||||
```
|
||||
*Expected Outcome*: Exit code `0`, structured JSON confirmation on stdout that all preflight checks passed.
|
||||
|
||||
---
|
||||
|
||||
### Scenario B: End-to-End Processing (Approved Article)
|
||||
Executes consolidation on an article with direct ECP inherence.
|
||||
|
||||
```bash
|
||||
python -m src.cli.consolidate \
|
||||
--input-article examples/sample_article_valid.json \
|
||||
--ecp-snapshot examples/sample_ecp_snapshot.json \
|
||||
--config runtime_config.local.json
|
||||
```
|
||||
*Expected Outcome*:
|
||||
- Exit code `0`.
|
||||
- Output manifest `./out/<fingerprint>.result.json` created with `final_status: "completed_text"` and `generate_markdown: true`.
|
||||
- Published Markdown `./out/<fingerprint>.md` created with YAML front matter.
|
||||
- SQLite state table updated to `completed_text`.
|
||||
- stdout emits the complete JSON manifest matching `manifest-output.schema.json`.
|
||||
|
||||
---
|
||||
|
||||
### Scenario C: ECP Rejection (Zero Markdown)
|
||||
Executes consolidation on a non-inherent article (`TANGENTIAL` or `NOT_RELATED`).
|
||||
|
||||
```bash
|
||||
python -m src.cli.consolidate \
|
||||
--input-article examples/sample_article_tangential.json \
|
||||
--ecp-snapshot examples/sample_ecp_snapshot.json \
|
||||
--config runtime_config.local.json
|
||||
```
|
||||
*Expected Outcome*:
|
||||
- Exit code `0`.
|
||||
- Manifest `./out/<fingerprint>.result.json` created with `final_status: "rejected_ecp"` and `generate_markdown: false`.
|
||||
- **Zero `<fingerprint>.md` file created for this execution fingerprint**.
|
||||
- SQLite state table updated to `ecp_rejected`.
|
||||
- stdout emits the complete JSON manifest matching `manifest-output.schema.json`.
|
||||
|
||||
---
|
||||
|
||||
### Scenario D: Idempotent Execution & Replay
|
||||
Re-executes the command on a previously completed article fingerprint.
|
||||
|
||||
```bash
|
||||
python -m src.cli.consolidate \
|
||||
--input-article examples/sample_article_valid.json \
|
||||
--ecp-snapshot examples/sample_ecp_snapshot.json \
|
||||
--config runtime_config.local.json
|
||||
```
|
||||
*Expected Outcome*:
|
||||
- Immediate return without re-invoking LLM providers.
|
||||
- Output JSON manifest on stdout matches previously persisted result.
|
||||
|
||||
---
|
||||
|
||||
## 4. Automated Quality Evaluation & Test Commands
|
||||
|
||||
```bash
|
||||
# 1. Run static multi-parser zero-regex check (Python AST + JSON pattern check + Promptfoo YAML check)
|
||||
python -m tests.scripts.check_zero_regex
|
||||
|
||||
# 2. Run unit and all 9 contract tests (including the 20 real reference units)
|
||||
pytest tests/unit tests/contract -v
|
||||
|
||||
# 3. Run security tests (SEC-001 to SEC-008)
|
||||
pytest tests/security -v
|
||||
|
||||
# 4. Run load test (100 articles/hour benchmark)
|
||||
pytest tests/load -v
|
||||
|
||||
# 5. Run mock integration tests, fault injection (10 scenarios), and operations resilience (FR-081)
|
||||
pytest tests/integration tests/fault_injection -v
|
||||
|
||||
# 6. Run offline Promptfoo evaluations
|
||||
npx promptfoo eval -c evals/promptfoo.config.yaml
|
||||
```
|
||||
@@ -0,0 +1,129 @@
|
||||
# Research & Technical Decisions: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Date**: 2026-08-23
|
||||
**Status**: Complete
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary & Architectural Scope
|
||||
|
||||
The Article Consolidation and Hygiene Runtime is an ephemeral Python CLI tool that processes exactly one article unit per run alongside a canonical Entity Context Profile (ECP) snapshot.
|
||||
|
||||
The runtime adheres strictly to the simplicity mandate (FR-001) and dependency policy (FR-002):
|
||||
- Direct Python state machine orchestration (no LangChain, LangGraph, agent frameworks, or workflow engines).
|
||||
- SQLite in WAL mode with short transactions and configurable lock timeout.
|
||||
- Local filesystem for atomic staged file writes and renames.
|
||||
- Agnostic Model Gateway managing two logical roles (`runtime_primary` and `runtime_fallback`) configured with certified low-cost models (defaults: Groq with `openai/gpt-oss-20b` and DeepSeek with `deepseek-v4-flash`).
|
||||
- 100% LLM extractive content hygiene with strict 10-step validation harness.
|
||||
- Grounded candidate selection and controlled micro-repairs without regular expressions (`re`) or manual keyword dictionaries.
|
||||
- Input size limit validation (FR-056) failing in a controlled manner with `INVALID_ARTICLE_SCHEMA` before provider calls if exceeded.
|
||||
- Mandatory ECP relevance gate evaluated directly via the monorepo's canonical ECP classifier module (`src.classifier.classify_text`).
|
||||
- Front matter enrichment (entity sentiment and 3–8 native language tags) post-ECP approval.
|
||||
- Direct observability via Langfuse (spans, generations, scores, traces), with durable local queuing in SQLite (`pending_telemetry`) during network outages.
|
||||
|
||||
---
|
||||
|
||||
## 2. Definitive Dependency Selections & Qualitative Impact Analysis
|
||||
|
||||
In accordance with FR-002, all runtime dependencies are closed, definitive, and qualitatively evaluated against the 7 mandatory criteria:
|
||||
|
||||
| Component / Task | Chosen Solution | Standard Library Alternative | Security Impact | Maintenance Impact | License | Size Impact | Startup Impact |
|
||||
|:---|:---|:---|:---|:---|:---|:---|:---|
|
||||
| **HTML DOM Parsing** | `beautifulsoup4` (with `html.parser`) | `html.parser` directly | Low. Pure Python, robust against malformed HTML. | Low. Stable and mature. | MIT | Minimal | Negligible |
|
||||
| **Markdown AST Parsing** | `marko` | None in stdlib for CommonMark AST | Low. Standard CommonMark AST compliance. | Low. Pure Python parser. | MIT | Minimal | Negligible |
|
||||
| **JSON Schema Validation** | `jsonschema` + `referencing` | Manual validation code | Low. Standard Draft 2020-12 validator with immutable local schema registry. | Low. Reference standard. | MIT | Minimal | Negligible |
|
||||
| **HTTP Client / Gateway** | `httpx` | `urllib.request` | Low. Modern HTTP client with connection pooling. | Low. High adoption. | BSD-3-Clause | Minimal | Negligible |
|
||||
| **Date Parsing & ISO 8601** | `python-dateutil` | `datetime.fromisoformat` | Low. Handles diverse timezone & date formats. | Low. Industry standard. | Apache 2.0 / BSD | Minimal | Negligible |
|
||||
| **YAML Serialization** | `pyyaml` | None in stdlib | Low. Safe dump (`yaml.safe_dump`) prevents code execution. | Low. Standard YAML library. | MIT | Minimal | Negligible |
|
||||
| **Language Detection** | `src.language` (existing monorepo) | N/A | None. Reuses existing repository module. | Zero new dependency. | Monorepo | Zero | Zero |
|
||||
| **Diffing without Regex** | `difflib.SequenceMatcher` | Stdlib `difflib` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||||
| **Unicode Normalization** | `unicodedata` | Stdlib `unicodedata` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||||
| **URL Parsing** | `urllib.parse` | Stdlib `urllib.parse` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||||
| **State Store & Queue** | `sqlite3` | Stdlib `sqlite3` | None. Standard library. | None. Standard library. | Python | Zero | Zero |
|
||||
| **Observability SDK** | `langfuse` | Direct HTTP calls | Low. Official SDK (>=4.7). | Low. Active upstream support. | MIT | Minimal | Negligible |
|
||||
|
||||
*Note on Durable Queuing*: Durable local queuing during observability outages is provided directly by the SQLite table `pending_telemetry`, not delegated to SDK memory buffers.
|
||||
|
||||
---
|
||||
|
||||
## 3. Concrete Architectural & Technical Decisions
|
||||
|
||||
### Decision 1: Direct Python State Machine Orchestration (FR-003, FR-046)
|
||||
- **Decision**: Orchestration is implemented as an explicit Python class `StateMachine` managing state transitions in SQLite:
|
||||
- `received → validated → content_cleaned`
|
||||
- `content_cleaned → ecp_approved → enriched → completed_text`
|
||||
- `content_cleaned → ecp_rejected` (terminal state, zero Markdown files generated)
|
||||
- Valid terminal failures → `failed`
|
||||
- **Transition Recording**: Every state transition records `start_time`, `end_time`, `duration_ms`, and `result` into SQLite table `state_transitions`.
|
||||
|
||||
### Decision 2: Certified Model Gateway & Preflight Certification (FR-041, FR-042, FR-044)
|
||||
- **Decision**: Agnostic Model Gateway (`src.gateway.client`) supporting two certified logical roles:
|
||||
- `runtime_primary`: Default configured as Groq with `openai/gpt-oss-20b` (or certified low-cost equivalent).
|
||||
- `runtime_fallback`: Default configured as DeepSeek with `deepseek-v4-flash` (or certified low-cost equivalent).
|
||||
- **Release Metadata Certification**: Build/packaging produces an immutable `src/core/release-metadata.json` packaged with the release containing:
|
||||
- `release_version`
|
||||
- `runtime_config_sha256` (SHA-256 of the approved functional config)
|
||||
- `prompts_hashes` (SHA-256 of each prompt file)
|
||||
- `schemas_versions` (contract version identifiers)
|
||||
- `certified_models` (logical role to approved provider/model mappings)
|
||||
- `ecp_classifier_config_hash` (SHA-256 of the deterministic configuration summary exposed by the ECP classifier adapter/module, asserting cheap/deterministic parameters)
|
||||
Preflight compares the loaded configuration against `src/core/release-metadata.json`. Any mismatch aborts with exit code `2`.
|
||||
- **Retry & Fallback Policy**:
|
||||
- Technical retries for transient errors only (timeout, connection reset/interruption, HTTP 429 with backoff, HTTP 5xx, empty technical response).
|
||||
- Semantic failures (invalid schema, grounding violation) transition immediately to `runtime_fallback` without retrying on the same model.
|
||||
- **Pricing & Operational Parameters**: Token prices are operational configuration parameters used for cost calculation and do NOT alter the functional execution fingerprint.
|
||||
|
||||
### Decision 3: Input Size Limit Enforcement (FR-056)
|
||||
- **Decision**: Prior to candidate extraction or remote provider calls, input size is checked against `limits.max_input_bytes`. If exceeded, execution terminates immediately with controlled error `INVALID_ARTICLE_SCHEMA` (with structured detail `"INPUT_EXCEEDS_SIZE_LIMIT"`). Strategy is strictly `fail_before_provider`. Automatic unapproved truncation is strictly prohibited.
|
||||
|
||||
### Decision 4: Candidate Equivalence Mapping without Deduplication (FR-014, FR-020)
|
||||
- **Decision**:
|
||||
- The runtime preserves all structurally valid candidate elements from all extractors.
|
||||
- `difflib.SequenceMatcher` calculates cross-extractor sequence similarity during candidate preparation to populate the `equivalences` list, providing consensus evidence to the LLM without deleting or merging candidates.
|
||||
- The `selected_extractor` provides the structural backbone order; secondary extractors provide alternative candidates accessible via explicit LLM ID selection.
|
||||
- Candidate IDs are opaque, stable within execution, and carry no quality judgment.
|
||||
|
||||
### Decision 5: 10-Step Hygiene Harness and Controlled Micro-Repairs (FR-024 to FR-030)
|
||||
- **Decision**:
|
||||
- LLM returns ONLY candidate IDs and micro-repair operations in `article_content_hygiene`.
|
||||
- Harness executes the 10-step sequence: 1. parse JSON; 2. validate schema; 3. validate metadata IDs; 4. validate block IDs; 5. validate backbone order; 6. validate links/images; 7. validate repairs individually; 8. assemble intermediate document; 9. validate grounding; 10. validate minimum content.
|
||||
- Micro-repairs are validated across 5 closed categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`) using `unicodedata` and `difflib` without regex.
|
||||
- Semantic changes in sensitive entities (names, numbers, dates, scores, quotes, facts) are rejected; verifiable encoding defects (e.g. mojibake in a proper name) are accepted.
|
||||
- Ungrounded candidate IDs trigger `GROUNDING_VIOLATION`; schema failures trigger semantic fallback; exhausted options trigger `HYGIENE_FAILED`.
|
||||
|
||||
### Decision 6: Local Resolution of Canonical ECP Schema & Inherence Gate (FR-031 to FR-036)
|
||||
- **Decision**:
|
||||
- The runtime receives the integral ECP Snapshot.
|
||||
- The local resolver loads the canonical schema file declared in `runtime-config.ecp.canonical_schema_reference` (`src/adapters/ecp/schemas/ecp-profile.schema.json`), verifies that its `$id` matches the `$ref` (`https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json`) in `ecp-snapshot.schema.json`, and registers it locally via `referencing.Registry`. Network HTTP retrieval is strictly disabled.
|
||||
- Invocations call the existing monorepo module `src.classifier.InherenceClassifier` directly through the runtime adapter. The adapter verifies classifier configuration matches `ecp_classifier_config_hash`.
|
||||
- Output validation verifies `category`, `is_inherent`, `confidence`, `rationale`, and textual `evidences` (grounded substrings in intermediate Markdown).
|
||||
- `DIRECT_INHERENT` / `CONTEXTUAL_INHERENT` → `ecp_approved`.
|
||||
- `TANGENTIAL` / `NOT_RELATED` → `ecp_rejected` (manifest persisted with status `rejected_ecp`, zero Markdown files generated).
|
||||
|
||||
### Decision 7: Atomic Storage, Backup & Graceful Shutdown (FR-009, FR-046, FR-050, FR-081)
|
||||
- **Decision**:
|
||||
- SQLite in WAL mode (`PRAGMA journal_mode=WAL;`, `PRAGMA busy_timeout=<configured_ms>;`).
|
||||
- Native backup and restore implemented in `src/storage/sqlite_store.py` via `sqlite3.Connection.backup`.
|
||||
- Output files (`.md` and `.result.json`) are written to temporary files in the target directory, flushed, verified by 64-character content hash, and atomically renamed (`os.replace`).
|
||||
- Interrupted write reconciliation (ID-009): If process crashes after Markdown write but before SQLite final state update, replay matches content hash, avoids duplicate calls/outputs, and updates SQLite state to `completed_text`.
|
||||
- Signal Handling: `src/cli/consolidate.py` traps `SIGTERM`/`SIGINT` to safely finish in-flight filesystem operations and flush/persist pending telemetry in SQLite before process exit.
|
||||
|
||||
### Decision 8: Observability, Metrics & Dashboards in Langfuse (FR-060 to FR-069)
|
||||
- **Decision**:
|
||||
- Observability Integration: Spans, generations, scores, trace attributes, and latency/cost metrics are transmitted directly to Langfuse via SDK (configured with `LANGFUSE_BASE_URL` and `LANGFUSE_PUBLIC_KEY`/`LANGFUSE_SECRET_KEY`).
|
||||
- Metric Catalog Compliance: Metrics strictly and exclusively use only the dimensions defined in the metric catalog of Doc 05 and FR-067:
|
||||
- `input_validation_failure_total`: `reason`, `schema_version`
|
||||
- `llm_request_total`: `logical_call`, `provider`, `model`, `status`
|
||||
- `llm_retry_total`: `reason`, `provider`, `model`
|
||||
- `llm_fallback_total`: `logical_call`, `reason`
|
||||
- `llm_output_validation_failure_total`: `logical_call`, `reason`, `prompt_version`
|
||||
- `prompt_review_signal_total`: dimensions defined in FR-068
|
||||
*(All other metrics follow strictly and solely the definitions of Doc 05 without unapproved custom dimensions).*
|
||||
- Dashboards: The 3 required dashboards (Runtime Health, Quality, Future Review Signals) are configured and viewed directly in Langfuse via Langfuse Custom Dashboards (imported operationally as JSON template configurations, avoiding unstable programmatic APIs).
|
||||
- Outage Fallback: If Langfuse is unreachable, events are recorded in SQLite table `pending_telemetry` and flushed via operational CLI `src.cli.telemetry_flush`.
|
||||
- Secret Redaction: Structural sanitization of authorization headers and exact token replacement of environment secrets (`GROQ_API_KEY`, `DEEPSEEK_API_KEY`, `LANGFUSE_SECRET_KEY`), without semantic text manipulation.
|
||||
|
||||
### Decision 9: Empirically Calibrated Staging SLOs (FR-076, FR-077)
|
||||
- **Decision**:
|
||||
- Latency (p50/p95/p99), cost per article, cost per approved Markdown, timeout limits, and SQLite lock wait thresholds are measured and calibrated during staging load testing (100 articles/hour) and approved prior to go-live.
|
||||
@@ -0,0 +1,549 @@
|
||||
# Feature Specification: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature Branch**: `006-article-consolidation-runtime`
|
||||
**Created**: 2026-08-23
|
||||
**Status**: Draft
|
||||
**Input**: User description: "baseado em todos esses arquivos docs/structured_extraction (01_PRD_Runtime_Consolidacao_Artigos.md, 02_Arquitetura_Runtime_Consolidacao_Artigos.md, 03_ADRs_Runtime_Consolidacao_Artigos.md, 04_Plano_Testes_Evals_Runtime.md, 05_Metricas_KPIs_Runtime.md, 06_Runbook_Producao_Runtime.md, 07_Especificacao_Prompt_Contexto_Harness_Runtime.md). Nao deve ser negligenciado nada! Tem que implementar 100% do que esta previsto nessas documentacoes, nao pode extrapolar em nada!"
|
||||
|
||||
---
|
||||
|
||||
## Priority & Governance Statement
|
||||
|
||||
> [!IMPORTANT]
|
||||
> **Priority Definitions**: In this specification, priority labels (P1, P2, P3) denote the sequential order of module implementation and testing. **Every User Story (US1 through US9), functional requirement (FR-001 to FR-084), metric, operational procedure, and quality gate in this document is strictly mandatory for production go-live.** No requirement is optional.
|
||||
|
||||
---
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Single Article Ingestion, Contract Validation, and Deterministic Candidate Preparation (Priority: P1)
|
||||
|
||||
As a pipeline orchestrator, I want to submit a single extracted news article unit (containing structural crawl fields, extractions from Trafilatura, Newspaper4k, and Readability, and a pre-calculated `selected_extractor`) alongside a versioned canonical ECP snapshot, so that the runtime validates input contracts locally before any remote call, establishes a deterministic idempotency fingerprint over canonical serialization, preserves unknown fields, and structures candidate elements using standard parsers without regular expressions or keyword dictionaries.
|
||||
|
||||
**Why this priority**: Foundational entry point of the pipeline. Enforces strict input validation, protects against premature remote calls, and establishes candidate provenance.
|
||||
|
||||
**Independent Test**: Can be tested by providing individual article JSON objects and ECP snapshots (valid, corrupted, missing fields, or batch wrappers), verifying schema validation, deterministic fingerprint generation over canonical serialization, SQLite state initialization (`received`, `validated`), candidate ID assignment (opaque, without quality judgment, stable within execution), and immediate failure with explicit error codes before any remote call when preconditions fail.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a valid single-article JSON input containing structural fields (`crawled_url`, `error_message`, `extraction_status`, `http_status`, `input_meta`, `page_title`, `selected_extractor`, `trafilatura`, `newspaper4k`, `readability`), a usable `selected_extractor`, at least one valid HTTP/HTTPS source URL, at least one non-empty candidate title, at least one processable text block, and a valid ECP snapshot matching the canonical ECP schema, **When** ingestion executes, **Then** the runtime validates all contracts locally without making remote provider, remote Langfuse, or ECP classifier calls, computes a deterministic content hash over canonical serialization, records `received` and `validated` states in SQLite, preserves unknown fields in recorded input while ignoring them in processing, and extracts structural candidates using DOM, CommonMark AST, JSON parser, URL parser, Unicode normalizers, and multilingual tokenizers/segmenters without any regular expressions.
|
||||
2. **Given** an invalid input payload (a batch JSON containing a root `articles` array, unparseable JSON, missing or unrecognized `selected_extractor`, `selected_extractor` pointing to an extraction with no usable content, missing source URL, missing title candidate, or missing textual content), **When** validation executes, **Then** the runtime terminates immediately prior to any remote call, returns a structured JSON error result, and records the specific error code:
|
||||
- `INVALID_ARTICLE_SCHEMA` for malformed JSON or batch payload;
|
||||
- `MISSING_SELECTED_EXTRACTOR` when `selected_extractor` is absent;
|
||||
- `INVALID_SELECTED_EXTRACTOR` when `selected_extractor` is unrecognized;
|
||||
- `SELECTED_EXTRACTOR_UNAVAILABLE` when the selected extractor has no usable content;
|
||||
- `MISSING_SOURCE_URL` when no valid source URL can be resolved;
|
||||
- `MISSING_TITLE_CANDIDATE` when no valid candidate title exists;
|
||||
- `MISSING_CONTENT` when no processable text content exists.
|
||||
The runtime MUST NOT recalculate or silently substitute `selected_extractor`.
|
||||
3. **Given** an absent, unparseable, or schema-incompatible ECP snapshot, **When** validation executes, **Then** the runtime terminates immediately with `INVALID_ECP_SCHEMA` before candidate preparation or any remote invocation.
|
||||
4. **Given** an input whose deterministic fingerprint matches a previously completed execution under identical functional versions and configurations, **When** ingestion executes, **Then** the runtime recognizes the completed state and returns the existing persisted result and artifacts without redundant LLM invocations.
|
||||
5. **Given** deterministic metadata resolution, **When** metadata candidates are resolved, **Then**:
|
||||
- **Source URL** is resolved strictly in priority order (`crawled_url` → `input_meta.url` → canonical URL of `selected_extractor` → `newspaper4k.canonical_link` → `trafilatura.canonical_url`) and is never chosen or modified by the LLM.
|
||||
- **Published Date** is resolved via date parser and normalized to ISO 8601, prioritizing source consensus or fallback priority (`newspaper4k.publish_date` → `input_meta.quando_publicado` → `trafilatura.date`), omitting `published_at` if invalid, and is never chosen or modified by the LLM.
|
||||
- **Candidate Title** sources include `input_meta.titulo`, `page_title`, `trafilatura.title`, `newspaper4k.title`, `readability.title`, `readability.short_title`.
|
||||
- **Candidate Subtitle** sources include `input_meta.subtitulo`, `trafilatura.description`, `newspaper4k.meta_description`.
|
||||
- **Candidate Author** sources include `trafilatura.author`, items of `newspaper4k.authors` (preserving structured list items and order without regex/delimiter splitting), `readability.author`.
|
||||
- **Candidate Types** include title, subtitle, author, date, block, heading, list, quote, link, and image.
|
||||
- **Candidate IDs** are opaque, contain no quality judgment, are unique within the execution, and duplicate IDs result in internal construction failure.
|
||||
- **Cross-Extractor Equivalence** is evaluated via Unicode normalization, whitespace normalization by library, tokenization, and sequence similarity without regex or keyword dictionaries. Low similarity retains distinct candidate IDs.
|
||||
- **Language** is detected using an appropriate multilingual NLP library.
|
||||
- **Structural Backbone**: `selected_extractor` defines the base ordering backbone; the LLM is not forced to select all its blocks, and blocks exclusive to secondary extractors enter only via explicit LLM selection and grounding validation.
|
||||
- **HTML/JSON-LD Parsing**: Malformed HTML is parsed tolerantly into a safe DOM or controlled failure; valid JSON-LD `Article`/`NewsArticle` yields structural metadata; invalid JSON-LD is ignored with a warning logged.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Mandatory LLM Extractive Content Hygiene and Controlled Text Repairs (Priority: P1)
|
||||
|
||||
As an editorial consumer, I want every article to undergo LLM extractive hygiene to select genuine editorial candidates and eliminate noise (advertisements, cross-promotions, player chrome, navigation, duplicate snippets, and newsletters), while permitting only auditable, bounded micro-repairs for unmistakable encoding and typographical flaws.
|
||||
|
||||
**Why this priority**: Eliminates editorial noise and guarantees 100% grounded content selection without autonomous rewriting.
|
||||
|
||||
**Independent Test**: Can be tested with single articles exhibiting consensus or divergence across extractors, containing textual ads, navigation links, and controlled encoding flaws, verifying that the LLM returns only candidate IDs and explicit repair operations, the harness validates grounding, invalid repairs are discarded with originals preserved, and ungrounded responses trigger fallback.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an article where all three extractors agree 100% on content, **When** content hygiene executes, **Then** the LLM hygiene step is executed unconditionally (consensus improves evidence but never bypasses LLM hygiene).
|
||||
2. **Given** the `article_content_hygiene` prompt and candidate payload, **When** the LLM responds, **Then** the response conforms strictly to the schema returning `title_candidate_id`, `subtitle_candidate_id` (or null), `author_candidate_id` (or null), `kept_block_ids` (in order), `kept_link_ids`, `kept_image_ids`, `repairs`, and categorical removal reasons (when enabled for observability), with an absolute absence of free-form body text or free Markdown fields.
|
||||
3. **Given** the hygiene harness validation pipeline, **When** an LLM hygiene response is evaluated, **Then** the harness executes the strict 10-step sequence:
|
||||
1. Parse JSON;
|
||||
2. Validate schema;
|
||||
3. Validate metadata IDs and their expected types;
|
||||
4. Validate block IDs exist in candidate storage;
|
||||
5. Validate ordering compatibility with canonical representation;
|
||||
6. Validate links and images against input candidates;
|
||||
7. Validate proposed repairs individually;
|
||||
8. Assemble intermediate structure by retrieving candidate text from internal maps;
|
||||
9. Validate grounding of assembled Markdown;
|
||||
10. Validate minimum content requirements (rejecting selections that omit material editorial content).
|
||||
4. **Given** an LLM hygiene response, **When** the harness validates the response, **Then** an ungrounded candidate ID, ungrounded link/image URL, or ungrounded content triggers a `GROUNDING_VIOLATION` (invalidating the entire response) and routes to fallback; an invalid schema triggers semantic fallback without being labeled as a grounding violation; and if all fallback options are exhausted, processing terminates with `HYGIENE_FAILED`.
|
||||
5. **Given** proposed text repairs within closed allowable categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`) containing target ID, exact original fragment, replacement fragment, category, and rationale, **When** the harness validates the repair, **Then** it normalizes and tokenizes original and replacement with Unicode/NLP libraries without regex, computes diffs without regex, verifies closed category, and records original, replacement, decision, and rationale.
|
||||
6. **Given** a proposed repair targeting sensitive entities (names, numbers, dates, scores, quotes, or facts), **When** the harness validates the repair, **Then** the repair is rejected (`INVALID_TEXT_REPAIR`) UNLESS the difference is strictly caused by an unmistakable, verifiable encoding/Unicode defect (such as mojibake in a proper name). If in doubt, the original text is preserved.
|
||||
7. **Given** a proposed repair that attempts stylistic improvement, synonym replacement, paraphrasing, tone change, or whose original fragment is ambiguous or missing, **When** the harness validates the repair, **Then** the invalid repair is discarded (`INVALID_TEXT_REPAIR`), the exact original candidate text is preserved, and processing continues without invalidating the rest of the valid selection.
|
||||
8. **Given** editorial assembly of intermediate Markdown, **When** the assembler constructs the document, **Then** it follows the `selected_extractor` backbone, resolves candidate content from internal maps, applies validated repairs, eliminates exact structural title/subtitle duplicates, grounds image positions structurally with alt/caption derived strictly from input text (images without editorial position are omitted), verifies link URLs and anchor text exist in input (rejecting malformed links), and formats Markdown without underline or inline HTML.
|
||||
9. **Given** editorial rules, **When** hygiene executes, **Then** the system strictly enforces: no translation, no summarization, no narrative reorganization, no transition creation, no information completion, no factual correction, no linguistic variant changes, no repetition of author/date/sentiment/tags/ECP in the body. The LLM never controls the renderer or filesystem.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Mandatory ECP Gate and Relevance Enforcement (Priority: P1)
|
||||
|
||||
As a content governance stakeholder, I want the intermediate sanitized Markdown to be evaluated by the ECP classifier adapter before any final editorial output is generated, ensuring that only articles with direct or contextual inherent relevance are permitted to produce published Markdown.
|
||||
|
||||
**Why this priority**: Enforces business relevance and editorial boundary constraints; prevents non-inherent content from being published while maintaining an audit trail.
|
||||
|
||||
**Independent Test**: Can be tested by passing intermediate sanitized Markdown documents to the ECP classifier adapter with associated ECP profiles across all four classification outcomes, verifying that only `DIRECT_INHERENT` and `CONTEXTUAL_INHERENT` progress to enrichment and Markdown output, while `TANGENTIAL` and `NOT_RELATED` generate a persisted `rejected_ecp` manifest and zero Markdown files.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** intermediate sanitized Markdown submitted to the ECP adapter, **When** the ECP classifier returns `DIRECT_INHERENT` or `CONTEXTUAL_INHERENT`, **Then** the state machine transitions from `content_cleaned` to `ecp_approved` and proceeds to enrichment.
|
||||
2. **Given** intermediate sanitized Markdown submitted to the ECP adapter, **When** the ECP classifier returns `TANGENTIAL` or `NOT_RELATED`, **Then** the state machine transitions from `content_cleaned` to `ecp_rejected` (a terminal state), writes a persistent `<fingerprint>.result.json` manifest with status `rejected_ecp` (`ECP_REJECTED`) and classification metadata, and guarantees no Markdown file is created.
|
||||
3. **Given** intermediate sanitized Markdown submitted to the ECP adapter, **When** the ECP adapter validates the classifier output, **Then** it requires category, `is_inherent`, confidence, rationale, and evidences, and verifies that all evidence fragments belong to the intermediate sanitized Markdown.
|
||||
4. **Given** an ECP classifier technical failure, invalid enum, or evidence fragment not present in the intermediate document, **When** the adapter processes the output, **Then** the execution terminates with `ECP_CLASSIFICATION_FAILED`, transitions to `failed`, and outputs no Markdown file.
|
||||
5. **Given** any LLM tier utilized by the ECP classifier, **When** the adapter executes, **Then** the adapter verifies that only certified cheap models are configured, records the generation telemetry, and rejects any configuration specifying uncertified powerful models.
|
||||
|
||||
---
|
||||
|
||||
### User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2)
|
||||
|
||||
As a content consumer, I want ECP-approved articles to receive structured metadata enrichment consisting of entity-relative sentiment (`positive`, `negative`, `neutral`) and 3 to 8 topic tags in the article's native language, strictly decoupled from the editorial body text.
|
||||
|
||||
**Why this priority**: Required for mandatory front matter metadata; must remain isolated from body text to prevent content modification.
|
||||
|
||||
**Independent Test**: Can be tested by providing approved intermediate Markdown, language, and minimal ECP entity identity to the `article_sentiment_tags` prompt, verifying that sentiment is evaluated strictly relative to the ECP entity, tags are in the article's language and bounded between 3 and 8 with evidence IDs, duplicate tags are rejected via NLP libraries, and the body text is not modified.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an approved article and target ECP entity, **When** the `article_sentiment_tags` prompt executes, **Then** it returns sentiment (`positive`, `negative`, or `neutral` evaluated strictly relative to the ECP entity), between 3 and 8 unique tags in the article's language supported by textual evidence, and evidence candidate IDs, without returning or altering body text.
|
||||
2. **Given** an enrichment output, **When** the harness validates the response, **Then** it verifies tag count (3 to 8), validates tag uniqueness using Unicode/NLP libraries without regex, verifies all evidence IDs against the intermediate document, and rejects responses containing duplicate tags, ungrounded tags, or attempts to output/modify body content.
|
||||
3. **Given** an invalid enrichment response on `runtime_primary`, **When** the harness handles the failure, **Then** it routes to `runtime_fallback`.
|
||||
4. **Given** a scenario where both primary and fallback enrichment calls fail, **When** the stage concludes, **Then** processing terminates with status `failed_processing` (`ENRICHMENT_FAILED`), transitions to `failed`, and no Markdown file is created (sentiment and tags are mandatory front matter fields).
|
||||
|
||||
---
|
||||
|
||||
### User Story 5 - Canonical Markdown Rendering and Atomic Persistence (Priority: P2)
|
||||
|
||||
As a system integrator, I want the runtime to render a standardized Markdown document with structured YAML front matter and an accompanying machine-readable JSON result manifest, persisted atomically using fingerprint-based filenames and tracked in SQLite, so that downstream consumers never encounter partial, corrupt, or inconsistent artifacts.
|
||||
|
||||
**Why this priority**: Guarantees file-system atomicity, data integrity, and deterministic contract adherence under normal and failure conditions.
|
||||
|
||||
**Independent Test**: Can be tested by verifying generated `.result.json` manifests and `.md` files, validating YAML front matter structure, verifying atomic rename lifecycles, content hash verification, and SQLite state consistency across normal completions, concurrent executions, and interrupted writes.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** an approved, enriched article, **When** final Markdown rendering executes, **Then** it produces a document with:
|
||||
- **Mandatory YAML front matter**: `title`, `source_url`, `sentiment`, `tags`, `ecp_qid`, `ecp_canonical_name`, `ecp_category`, `ecp_confidence`.
|
||||
- **Optional YAML front matter (omitted when empty/absent)**: `subtitle`, `author`, `published_at`.
|
||||
- **Body Structure**: H1 `# Title`, italic `*Subtitle*` (when present, with NO artificial blank line generated if absent), and editorial content in canonical order with grounded links and images, omitting author, date, sentiment, tags, and ECP data from the body text.
|
||||
2. **Given** output artifacts (`<fingerprint>.result.json` and optional `<fingerprint>.md`), **When** persistence executes, **Then** artifacts are written to temporary files in the same filesystem, flushed, closed, verified against content hashes, atomically renamed to their final destination, and the SQLite state is persisted in the same logical completion unit, without exposing an inconsistent completed state. Manifest and Markdown are never presented as completed while the pair is inconsistent.
|
||||
3. **Given** any execution (success, validation failure, ECP rejection, or processing failure), **When** persistence completes, **Then**:
|
||||
- A structured JSON result is returned.
|
||||
- When a deterministic fingerprint exists, `<fingerprint>.result.json` is persisted containing:
|
||||
- `fingerprint`: deterministic hash of the execution;
|
||||
- `source_url`: resolved source URL;
|
||||
- `selected_extractor`: extractor received;
|
||||
- `final_status`: `completed_text`, `rejected_ecp`, `failed_validation`, or `failed_processing`;
|
||||
- `generate_markdown`: boolean decision indicating if Markdown was produced;
|
||||
- `markdown_path`: file path of Markdown output, or null;
|
||||
- `ecp_classification`: category and confidence, or null if pre-ECP failure;
|
||||
- `provider_versions`: configured provider identifiers and versions, or null if pre-LLM failure;
|
||||
- `model_versions`: configured model versions, or null if pre-LLM failure;
|
||||
- `prompt_versions`: prompt names, versions, and hashes, or null if pre-LLM failure;
|
||||
- `config_version`: runtime functional configuration version;
|
||||
- `trace_id`: Langfuse trace identifier, or null if local failure;
|
||||
- `error_codes`: specific error or rejection codes (`INVALID_ARTICLE_SCHEMA`, `INVALID_ECP_SCHEMA`, `MISSING_SELECTED_EXTRACTOR`, `INVALID_SELECTED_EXTRACTOR`, `SELECTED_EXTRACTOR_UNAVAILABLE`, `MISSING_SOURCE_URL`, `MISSING_TITLE_CANDIDATE`, `MISSING_CONTENT`, `HYGIENE_FAILED`, `GROUNDING_VIOLATION`, `INVALID_TEXT_REPAIR`, `ECP_CLASSIFICATION_FAILED`, `ECP_REJECTED`, `ENRICHMENT_FAILED`, `PERSISTENCE_FAILED`, `TELEMETRY_PENDING`).
|
||||
4. **Given** a storage write failure or hash mismatch, **When** persistence fails, **Then** processing terminates with `PERSISTENCE_FAILED` and leaves no incomplete final output.
|
||||
5. **Given** an interrupted execution where Markdown was written to disk but SQLite state was not updated before interruption, **When** recovery or replay executes, **Then** hash-based reconciliation identifies the existing valid output file, updates the SQLite state to `completed_text`, and avoids duplicate calls or duplicate output creation.
|
||||
6. **Given** two identical concurrent executions with the same fingerprint, **When** both run simultaneously, **Then** exactly one execution completes the full flow and the other execution safely resumes or reuses the persisted result, resulting in zero duplicate published outputs.
|
||||
|
||||
---
|
||||
|
||||
### User Story 6 - Agnostic Model Gateway, Cheap Models, and Controlled Fallbacks (Priority: P2)
|
||||
|
||||
As a cloud operations manager, I want all runtime LLM calls to be managed by an agnostic Model Gateway configured exclusively with certified cheap models (logical roles `runtime_primary` and `runtime_fallback`), supporting limited technical retries for transient errors, immediate fallback on schema/grounding failures, and conservative deterministic fallback, so that operational costs remain low and powerful models are never invoked.
|
||||
|
||||
**Why this priority**: Enforces strict cost discipline and architectural decoupling from specific LLM providers.
|
||||
|
||||
**Independent Test**: Can be tested by configuring primary and fallback providers, simulating transient network errors (timeouts, connection resets, HTTP 429, HTTP 5xx, empty responses) and semantic failures (schema violations, grounding failures), verifying technical retries with backoff, fallback routing, deterministic hygiene fallback, and preflight rejection of uncertified/powerful models.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** the Model Gateway interface, **When** invoked by the runtime, **Then** it accepts logical role, messages/context, structured schema, timeout, and trace metadata, and returns structured output, effective provider, effective model, tokens (input, output, cached), cost, latency, technical status, attempt number, and fallback indication.
|
||||
2. **Given** certified role configuration, **When** mapped in the gateway, **Then** each role maps to provider, model, parameters, compatible prompt, expected schema, timeout, and version.
|
||||
3. **Given** a transient technical error (timeout, connection interruption / connection reset, HTTP 429 with backoff up to configured limit, HTTP 5xx, empty response due to technical failure), **When** the Model Gateway executes a call, **Then** it performs limited technical retries on the same provider according to certified role limits.
|
||||
4. **Given** a semantic failure (invalid JSON schema, grounding violation) or exhausted technical retries on `runtime_primary`, **When** the gateway handles the failure, **Then** it transitions immediately to `runtime_fallback` without retrying semantic errors on the same model and without entering semantic retry loops.
|
||||
5. **Given** the gateway architecture, **When** routing calls, **Then** the gateway uses two adapters/configurations (`runtime_primary` and `runtime_fallback`) without implementing smart routers, autonomous model selectors, or dynamic model scoring.
|
||||
6. **Given** a failure of both `runtime_primary` and `runtime_fallback` during hygiene, **When** fallback evaluates, **Then** the runtime applies a conservative deterministic fallback based on the `selected_extractor` backbone, which eliminates only structurally invalid elements, does NOT use regex, does NOT attempt semantic ad/recommendation filtering, and fails with `HYGIENE_FAILED` if grounding or minimum content requirements cannot be guaranteed.
|
||||
7. **Given** any gateway configuration specifying a model uncertified for the runtime role (including any powerful model tier), **When** preflight validation runs, **Then** the runtime fails preflight checks and blocks execution.
|
||||
|
||||
---
|
||||
|
||||
### User Story 7 - Production Observability, Log Sanitization, and Deferred Telemetry Resend (Priority: P3)
|
||||
|
||||
As an SRE, I want every validated execution to record structured JSON logs and end-to-end tracing in Langfuse (with stable spans covering validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, and persistence, and generations capturing token counts, latency, calculated costs, and versions), with automatic redaction of secrets and graceful degradation to local SQLite queuing (`TELEMETRY_PENDING`) during observability outages, so that telemetry is complete and operations remain resilient.
|
||||
|
||||
**Why this priority**: Production traceability, cost tracking, and incident diagnosis without risking article pipeline blockage.
|
||||
|
||||
**Independent Test**: Can be tested by running executions with Langfuse available and unavailable, checking structured JSON logs for sanitized fields and correct event codes, inspecting Langfuse spans and generations, and executing the operational telemetry resend routine to flush pending records from SQLite.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a validated execution, **When** telemetry is recorded, **Then** a Langfuse trace is created under a stable `run_id` with stable spans covering the normative stages (validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, persistence) devoid of URLs, model names, or dynamic IDs, and each LLM attempt is recorded as a separate generation capturing prompt versions/hashes, model/provider, schema version, cached tokens (when available), token counts, calculated cost, latency, timeout status, attempts, technical status, applicable scores, schema results, fallback status, logical role, normalized context sent, structured response, grounding validation result, applied repairs, and rejected repairs (subject to trace content policy), without duplicating raw HTML or full JSON payloads.
|
||||
2. **Given** log outputs, Langfuse traces, and manifest files, **When** payloads are generated, **Then** all API keys, authorization headers, and environment secrets are redacted (`trace_redaction_failure_total` = 0). Full ECP, full HTML, and full article text are omitted from standard logs. Repair diffs reside in controlled traces only, never in metric labels.
|
||||
3. **Given** a network failure or outage reaching Langfuse, **When** an article is processed, **Then** the article processing completes normally, a `TELEMETRY_PENDING` record is saved in SQLite, and execution exits cleanly without stalling.
|
||||
4. **Given** pending telemetry records in SQLite, **When** the operational telemetry resend procedure is executed, **Then** pending events are delivered to Langfuse, deduplicated by event ID, and `telemetry_pending_total` returns to zero.
|
||||
5. **Given** metrics collection, **When** metrics are recorded, **Then** high-cardinality values (URLs, fingerprints, run IDs, trace IDs, titles, authors, full text, free tags) are strictly forbidden as metric labels and restricted to traces and logs.
|
||||
6. **Given** runtime operations, **When** prompt review signals occur (schema failures, grounding violations, rejected repairs, fallback invocations, terminal failures), **Then** the runtime emits `prompt_review_signal_total` metrics capturing `logical_call`, `prompt_version`, `provider`, `model`, `language`, `source_domain_group`, and `reason`, without initiating any autonomous self-healing.
|
||||
7. **Given** operations dashboards, **When** monitored, **Then** the system provides:
|
||||
- **Runtime Health Dashboard**: received, completed, rejected, failed, throughput, latency, cost, providers, fallback, persistence, pending telemetry.
|
||||
- **Quality Dashboard**: schema, grounding, repairs, ECP, sentiment, tags, and results by prompt/model/language/domain.
|
||||
- **Future Review Signals Dashboard**: `prompt_review_signal_total`, failures after fallback, concentration by prompt version, and associated potential cost.
|
||||
|
||||
---
|
||||
|
||||
### User Story 8 - Automated Quality Evaluation, Staging Baselines, and CI Quality Gates (Priority: P3)
|
||||
|
||||
As a release engineer, I want the CI pipeline and staging environments to execute AST static analysis for regex prohibition in text pipeline modules and assertions, Promptfoo evaluations executed in development/CI over the reference fixture (20 cases) and golden dataset (with holdout), fault injection, and 100 articles/hour load tests, so that cost/latency SLOs are empirically calibrated and zero-tolerance invariants prevent flawed releases.
|
||||
|
||||
**Why this priority**: Guarantees production readiness, validates multilingual performance across slices, and prevents architectural degradation.
|
||||
|
||||
**Independent Test**: Can be tested by running AST static checks on text pipeline modules and assertions, executing Promptfoo in dev/CI against prompt files without online runtime coupling, simulating fault injection (primary down, both down, Langfuse unavailable, SQLite lock, disk full, process crash during write, ECP down, truncated LLM response, orphan temp, telemetry flush failure), running 100 articles/hour load tests, and verifying critical invariant gates.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** text processing pipeline modules and Promptfoo assertion configurations, **When** AST static analysis runs in CI, **Then** it validates that no Python `re` module imports or regex functions are called in those modules/assertions, failing the build if any are detected.
|
||||
2. **Given** Promptfoo evaluation suites executed in development/CI, **When** evals execute, **Then** Promptfoo loads the identical versioned prompt files used in production, evaluating JSON schemas, candidate ID grounding, and repair diffs without using regex assertions, semantic keyword dictionaries, or LLM-as-a-judge for grounding.
|
||||
3. **Given** critical release gates, **When** a build is evaluated, **Then** promotion is blocked if any of the 11 critical metrics is greater than zero: `ungrounded_text_total`, `ungrounded_url_total`, `ungrounded_image_total`, `unauthorized_rewrite_total`, `critical_fact_change_total`, `duplicate_output_total`, `lost_article_total`, `secret_exposure_total`, `powerful_runtime_model_call_total`, `online_promptfoo_call_total`, `text_regex_usage_total`.
|
||||
4. **Given** quality evaluations across golden dataset and holdout, **When** metrics are computed, **Then** pass rate is evaluated per slice (language, domain, extractor, prompt version, model version), requiring at least 95% pass rate per slice without allowing a global average to mask localized failures, and measuring precision/recall/F1 of kept blocks, metadata accuracy, correct vs unauthorized repair rates, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, and tag acceptance.
|
||||
5. **Given** staging load testing, **When** 100 articles per hour are processed under realistic concurrency using the identical SQLite and output directory planned for production, with a representative mix of languages/domains/extractors/structures and realistic fallback rates, **Then** the run completes with zero lost articles, zero duplicate outputs, zero partial files exposed, zero database corruption, stable memory/disk usage, 100% traces delivered or preserved as pending telemetry, and produces empirical p50/p95/p99 latency and cost baselines for operational approval prior to go-live.
|
||||
6. **Given** the CI/CD pipeline, **When** changes are integrated, **Then** execution follows the minimum sequence:
|
||||
1. Format validation;
|
||||
2. Lint & static analysis;
|
||||
3. AST regex check;
|
||||
4. Unit tests;
|
||||
5. Contract tests;
|
||||
6. Simulated integration tests (with simulated/mock providers);
|
||||
7. Reduced Promptfoo eval;
|
||||
8. Package build;
|
||||
9. Full golden set eval (pre-promotion, using authorized offline providers);
|
||||
10. Staging 100 art/h load test;
|
||||
11. Cost/latency limits approval;
|
||||
12. Controlled promotion.
|
||||
|
||||
---
|
||||
|
||||
### User Story 9 - Production Runbook Operations and Lifecycle Management (Priority: P3)
|
||||
|
||||
As a production operations engineer, I want standardized operational routines for preflight checks, smoke tests, consistent SQLite backups/restores, graceful shutdown, reconciliation, manual rollbacks, credential rotations, and certified model/provider rotations, so that production can be reliably maintained, diagnosed, and recovered without manual file tampering.
|
||||
|
||||
**Why this priority**: Production operational readiness requirement ensuring all failure modes, deployments, and rollbacks have verified, auditable procedures.
|
||||
|
||||
**Independent Test**: Can be tested by executing preflight validation, running smoke tests against versioned fixtures, performing consistent SQLite backup/restore cycles, triggering graceful shutdown signals, running reconciliation reports, executing manual rollbacks to previous certified configurations, and testing credential rotations.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a new deployment or environment startup, **When** preflight validation runs, **Then** it validates system clock synchronization, prompts and configs belonging strictly to the same release, local configuration reading, prompt file existence & hashes, schema compatibility, SQLite access, filesystem atomic write permissions & directory permissions, minimum disk space, validated active credentials (not just presence), absence of powerful models in runtime roles, Langfuse local configuration (environment, content policy), and ECP classifier/schema access before accepting live traffic. Remote Langfuse connectivity failure does NOT block preflight.
|
||||
2. **Given** a preflight pass, **When** smoke testing executes, **Then** it processes a versioned reference fixture, verifying fingerprint generation, state transitions, LLM call, ECP gate, manifest, Markdown output, Langfuse trace, cost and latency within approved staging ranges, and idempotent re-execution.
|
||||
3. **Given** deployment procedures, **When** a release is deployed, **Then** it follows the complete 11-step sequence:
|
||||
1. Pause new executions in orchestrator;
|
||||
2. Await or gracefully drain existing executions;
|
||||
3. Preserve consistent backup of SQLite and active configuration;
|
||||
4. Deploy package, prompts, and schemas of the release;
|
||||
5. Test state migration on a copy before applying to production;
|
||||
6. Run preflight checks;
|
||||
7. Run smoke test with approved fixture;
|
||||
8. Confirm manifest, Markdown, state, and trace;
|
||||
9. Release with reduced concurrency;
|
||||
10. Verify errors, fallback rate, cost, and latency;
|
||||
11. Release full volume.
|
||||
4. **Given** operational lifecycle procedures, **When** maintenance tasks execute, **Then**:
|
||||
- **Responsibilities Matrix**: Operational roles reproduce the normative matrix:
|
||||
- *Orchestrator*: provide article/ECP, control concurrency, and consume the manifest.
|
||||
- *Operations*: deploy, monitor, recover, and execute rollback.
|
||||
- *Engineering*: correct code, prompts, schemas, or integrations through the normal release process.
|
||||
- *Curator/Eval*: maintain the golden set and approve quality.
|
||||
- **Backup**: SQLite is backed up using consistent database backup mechanisms (not raw file copies during active writes).
|
||||
- **Shutdown**: Upon receiving a shutdown signal, the runtime stops accepting new units, completes or persists the active unit safely, closes open transactions, flushes open files/telemetry, preserves pending telemetry, and exits cleanly with a coherent status.
|
||||
- **Reconciliation**: Periodic routines detect and report discrepancies between SQLite states, manifest files, Markdown files, orphan temp files, and pending telemetry without manual editing.
|
||||
- **Rollback**: Triggered by critical invariant violations, contract incompatibilities, or unmitigated failures, manual rollback restores the previous certified package/configuration/prompt versions and reprocesses affected units.
|
||||
- **Rotation**: Model/provider changes follow certification via Promptfoo, golden set evals, and staging baseline calibration before promotion, maintaining the previous configuration for rollback.
|
||||
- **Credentials**: Credential rotations create least-privilege credentials, update environment secrets, run preflight/smoke tests, verify absence of exposure, revoke old credentials, and avoid altering functional fingerprints when only operational secrets change.
|
||||
- **Incidents**: Failures in providers, ECP, enrichment, Langfuse, SQLite, disk, grounding, cost, latency, or input validation / producer schema mismatches (checking producer version, comparing with release schema, confirming unit payload, never calling LLM manually, fixing producer or contract via standard release), preserving incident evidence and creating mandatory regression test cases after critical incidents. Production prompts, sentiments, tags, Markdown files, manifests, and SQLite MUST NEVER be edited manually. Reprocessing MUST locate state/fingerprint, verify versions, reuse existing outputs or resume incomplete states under the same configuration, generate distinct fingerprints for new configurations, avoid manual file edits, and never reprocess articles merely to recreate traces.
|
||||
|
||||
---
|
||||
|
||||
## Edge Cases
|
||||
|
||||
- **Batch Array Payload**: Input containing a root `articles` array is rejected immediately with error code `INVALID_ARTICLE_SCHEMA`.
|
||||
- **Selected Extractor Handling**: If `selected_extractor` is absent, invalid, or lacks usable content, execution terminates with `MISSING_SELECTED_EXTRACTOR`, `INVALID_SELECTED_EXTRACTOR`, or `SELECTED_EXTRACTOR_UNAVAILABLE`. The runtime never recalculates or silently substitutes the extractor.
|
||||
- **Extractor Full Consensus**: LLM hygiene executes unconditionally even when all three extractors agree 100%.
|
||||
- **Textual Noise**: The LLM hygiene step removes textual ads, cross-promotions, player chrome, navigation, duplicate snippets, and newsletters using semantic understanding, without relying on hardcoded keyword lists per language.
|
||||
- **Unauthorized Textual Repairs**: Repairs attempting semantic paraphrasing, style improvements, synonym replacement, or alterations to facts, names, dates, numbers, scores, or quotes are rejected by the harness (`INVALID_TEXT_REPAIR`); the original candidate text is preserved, UNLESS the difference is strictly an unmistakable, verifiable encoding/Unicode defect (such as mojibake in a proper name).
|
||||
- **Ambiguous or Non-Existent Repair Target**: Repairs with ungrounded target IDs or ambiguous original fragments are rejected; original text is preserved.
|
||||
- **ECP Rejection (`TANGENTIAL` or `NOT_RELATED`)**: The runtime terminates editorial processing cleanly, persists a `<fingerprint>.result.json` manifest with status `rejected_ecp` (`ECP_REJECTED`), and produces NO Markdown file.
|
||||
- **ECP Outage or Invalid Output**: Terminate with `ECP_CLASSIFICATION_FAILED` and generate no Markdown.
|
||||
- **Primary & Fallback Model Outage**: If both primary and fallback LLMs fail during hygiene, conservative deterministic fallback is used only if baseline structural integrity is satisfied; otherwise, `HYGIENE_FAILED` is returned. If both fail during enrichment, `ENRICHMENT_FAILED` is returned and no Markdown is output.
|
||||
- **Prompt Injection in Article Content**: Article text and metadata are strictly treated as untrusted data candidates. Structural parsing, candidate ID referencing, and output schemas prevent prompt injection from executing instructions.
|
||||
- **Observability Outage**: Langfuse network failures do not halt processing; telemetry records are stored in SQLite as `TELEMETRY_PENDING` for deferred resending.
|
||||
- **SQLite Concurrency & Lock Wait**: Concurrent CLI executions on the same SQLite state store use WAL mode, busy timeout, and short transactions to prevent lock corruption under 100 articles/hour load.
|
||||
- **Process Termination During Write**: Staged writing to temporary files and atomic rename prevent partial files from being exposed as final outputs.
|
||||
- **Interrupted Write Reconciliation**: Hash-based reconciliation resolves crashes between Markdown write and SQLite final state update without duplicate processing.
|
||||
- **Future Self-Healing Boundary**: The runtime exclusively logs telemetry, versions, and review signals (`prompt_review_signal_total`). It does NOT contain prompt optimizers, judges, candidate generation, canaries, or automatic rollback mechanisms.
|
||||
|
||||
---
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
#### Master Simplicity and Architecture Principles
|
||||
- **FR-001**: System MUST satisfy all production requirements using the minimum necessary code, abstractions, dependencies, and components. No complexity MAY be added without satisfying an explicit requirement, documented risk, or proven operational need. Small responsibilities MAY share modules; the architectural component list does NOT mandate a separate class or package per item. No anticipatory implementation of future components (such as self-healing) is permitted.
|
||||
- **FR-002**: System MUST adhere to the strict dependency preference hierarchy: 1. standard library; 2. existing monorepo dependencies; 3. consolidated libraries eliminating significant custom implementation; 4. custom code for product-specific rules only. Every new dependency MUST document its requirement, stdlib alternative, security impact, maintenance impact, license, size impact, and startup impact. Agent frameworks, libraries for trivial single functions, secondary ECP schemas, manual HTML/Markdown/URL parsers, and anticipatory self-healing dependencies are strictly prohibited.
|
||||
- **FR-003**: System MUST execute as an ephemeral Python CLI processing exactly one article unit per invocation, without providing an API, without internal batch loops, and without internal worker pools. Concurrency is managed externally by the orchestrator, sharing the SQLite state store and filesystem safely.
|
||||
- **FR-004**: System MUST maintain independent semantic versions for all system contracts: input article contract, ECP Snapshot reference, runtime configuration, candidates payload, hygiene response schema, repair operations schema, enrichment response schema, output manifest schema, and prompts. Any incompatible change in any contract MUST require a new version and complete evaluation.
|
||||
|
||||
#### Input, Validation, and Fingerprinting Contracts
|
||||
- **FR-005**: System MUST require a valid ECP snapshot matching the canonical ECP schema per execution, failing immediately with `INVALID_ECP_SCHEMA` before any remote call if absent or invalid. The runtime MUST NOT duplicate or redefine the ECP schema.
|
||||
- **FR-006**: System MUST require and validate `selected_extractor` to be one of `trafilatura`, `newspaper4k`, or `readability`, and MUST NOT calculate, recalculate, or silently substitute the selected extractor.
|
||||
- **FR-007**: System MUST validate minimum editorial validity locally before any remote call (provider, remote Langfuse, or ECP classifier), failing with specific error codes:
|
||||
- `MISSING_SOURCE_URL` if no valid HTTP/HTTPS source URL exists;
|
||||
- `MISSING_TITLE_CANDIDATE` if no non-empty candidate title exists;
|
||||
- `MISSING_CONTENT` if no processable text block exists;
|
||||
- `SELECTED_EXTRACTOR_UNAVAILABLE` if the selected extractor lacks usable content.
|
||||
- **FR-008**: System MUST compute a deterministic hash over the canonical serialization of article identity, extraction content hashes, `selected_extractor`, ECP snapshot identity/version, prompt versions, functional runtime configuration, and configured model identifiers (excluding runtime timestamps and trace IDs).
|
||||
- **FR-009**: System MUST enforce idempotency via SQLite state store, returning existing completed outputs without re-running LLMs when an identical fingerprint and configuration are submitted. When identical concurrent executions occur, exactly one completes effectively and the other safely resumes or reuses the persisted result, ensuring zero duplicate published outputs.
|
||||
- **FR-010**: System MUST preserve unknown fields in recorded input while ignoring them during pipeline processing.
|
||||
- **FR-011**: System MUST consume available structural fields (`crawled_url`, `error_message`, `extraction_status`, `http_status`, `input_meta`, `page_title`, `selected_extractor`) and extraction fields:
|
||||
- *Trafilatura*: `title`, `author`, `date`, `description`, `text`, `markdown`, `canonical_url`, `image`, `language`, `sitename`, `categories`, `tags`, `raw_json`, `pagetype`, `error`.
|
||||
- *Newspaper4k*: `title`, `authors`, `publish_date`, `meta_description`, `text`, `article_html`, `canonical_link`, `top_image`, `images`, `meta_data`, `meta_lang`, `meta_site_name`, `tags`, `keywords`, `error`.
|
||||
- *Readability*: `title`, `short_title`, `author`, `cleaned_text`, `cleaned_html`, `error`.
|
||||
- **FR-012**: System MUST NOT perform any remote calls (provider, remote Langfuse, or ECP classifier) before completing all local validations capable of terminating execution.
|
||||
- **FR-013**: System MUST load environment configuration containing paths (input, output, state), provider endpoints/credentials, role configurations (`runtime_primary`, `runtime_fallback`), timeouts, retry limits, prompt versions, Langfuse config, trace content policy, and state store concurrency limits. Versioned functional configs MUST enter the fingerprint; operational secrets MUST NOT. System MUST support reproducible packaging and builds.
|
||||
|
||||
#### Deterministic Structural Candidate Model
|
||||
- **FR-014**: System MUST parse HTML (DOM parser), Markdown (CommonMark AST), JSON-LD (JSON parser), URLs (URL parser), and text (Unicode normalization, multilingual tokenizer/segmenter), creating identified candidate objects with opaque IDs (without embedded quality judgment, stable within execution), extractor origin, source field, original text/URL, structural type (title, subtitle, author, date, block, heading, list, quote, link, image), content hash, and cross-extractor equivalences. Duplicate candidate IDs MUST cause an internal construction failure.
|
||||
- **FR-015**: System MUST NOT use regular expressions (`re` or regex engines) anywhere in the text processing pipeline, classification, hygiene, parsing, or test assertions.
|
||||
- **FR-016**: System MUST NOT use manual keyword dictionaries or hardcoded word lists per language to make semantic decisions regarding advertising, recommendations, or editorial value.
|
||||
- **FR-017**: System MUST deterministically resolve the source URL using URL parsers in priority order: `crawled_url` → `input_meta.url` → canonical URL of `selected_extractor` → `newspaper4k.canonical_link` → `trafilatura.canonical_url`. The LLM MUST NOT select or modify the source URL.
|
||||
- **FR-018**: System MUST deterministically resolve the publication date normalized to ISO 8601, prioritizing consensus across sources, or fallback priority `newspaper4k.publish_date` → `input_meta.quando_publicado` → `trafilatura.date`, omitting `published_at` if invalid. The LLM MUST NOT select or modify the publication date.
|
||||
- **FR-019**: System MUST resolve candidate metadata sources:
|
||||
- *Title*: `input_meta.titulo`, `page_title`, `trafilatura.title`, `newspaper4k.title`, `readability.title`, `readability.short_title`.
|
||||
- *Subtitle*: `input_meta.subtitulo`, `trafilatura.description`, `newspaper4k.meta_description`.
|
||||
- *Author*: `trafilatura.author`, items of `newspaper4k.authors` (preserving structured list items and order without regex/delimiter splitting), `readability.author`.
|
||||
- **FR-020**: System MUST use `selected_extractor` as the structural backbone for ordering, using secondary extractors as consensus evidence and alternative candidate sources, without adopting similarity thresholds not calibrated on the golden set. `selected_extractor` does not force the LLM to keep all its blocks; secondary blocks enter only by explicit LLM selection and grounding validation. Low similarity MUST keep candidates distinct.
|
||||
- **FR-021**: System MUST parse malformed HTML into a safe DOM or controlled failure, register structural metadata from JSON-LD `Article`/`NewsArticle`, and ignore invalid JSON-LD with a logged warning.
|
||||
|
||||
#### LLM Extractive Content Hygiene and Controlled Repairs
|
||||
- **FR-022**: System MUST invoke the LLM content hygiene step for every article, even when all three extractors agree 100%.
|
||||
- **FR-023**: System MUST provide the LLM with candidate IDs, metadata candidates, block candidates, link candidates, image candidates, detected language, and strict JSON output schema, and the LLM MUST return only selected IDs and repair diffs without outputting free-form full article text or free Markdown.
|
||||
- **FR-024**: System MUST require the hygiene output schema to return: `title_candidate_id`, `subtitle_candidate_id` (or null), `author_candidate_id` (or null), `kept_block_ids` (in order), `kept_link_ids`, `kept_image_ids`, `repairs`, and categorical removal reasons when enabled for observability.
|
||||
- **FR-025**: System MUST execute the strict 10-step hygiene harness validation: 1. parse JSON; 2. validate schema; 3. validate metadata IDs and types; 4. validate block IDs; 5. validate order compatibility with canonical representation; 6. validate links and images against input candidates; 7. validate repairs individually; 8. assemble intermediate structure by retrieving candidates from internal maps; 9. validate grounding of assembled Markdown; 10. validate minimum content requirements (rejecting selections omitting material content).
|
||||
- **FR-026**: System MUST reject any hygiene response containing ungrounded candidate IDs, ungrounded URLs, ungrounded images, or candidate text not originating from input candidates as a `GROUNDING_VIOLATION`, invalidating the entire response and triggering fallback. Schema violations MUST trigger semantic fallback without being labeled as grounding violations, and if all fallback options are exhausted, processing MUST terminate with `HYGIENE_FAILED`.
|
||||
- **FR-027**: System MUST validate proposed text repairs against closed categories (`encoding`, `unicode`, `spacing`, `punctuation_corruption`, `obvious_typo`), requiring target candidate ID, exact original fragment, replacement fragment, category, and short rationale. Harness MUST normalize/tokenize original and replacement via Unicode/NLP libraries without regex, compute diffs without regex, and record original, replacement, decision, and rationale.
|
||||
- **FR-028**: System MUST reject any repair modifying facts, names, dates, numbers, scores, quotes, tone, or style, or referencing ambiguous/missing fragments (`INVALID_TEXT_REPAIR`), UNLESS the difference is strictly an unmistakable, verifiable encoding/Unicode defect (e.g. mojibake in a proper name). If in doubt, original text MUST be preserved.
|
||||
- **FR-029**: System MUST discard invalid repairs while preserving the exact original candidate text, continuing processing without authorizing unconstrained regeneration.
|
||||
- **FR-030**: System MUST enforce editorial rules: no translation, no summarization, no narrative reorganization, no transition creation, no information completion, no factual correction, no linguistic variant changes, removal of exact structural title/subtitle duplicates, no repetition of author/date/sentiment/tags/ECP in the body, image alt/captions derived strictly from input text, structurally grounded image positioning (omitting images without editorial position), and link validation (link URLs and anchor texts must exist in input, malformed links structurally rejected). The LLM MUST NEVER control the renderer or filesystem.
|
||||
|
||||
#### ECP Gate and Relevance Enforcement
|
||||
- **FR-031**: System MUST submit intermediate sanitized Markdown to the ECP classifier adapter prior to generating final Markdown.
|
||||
- **FR-032**: System MUST require the ECP adapter output to provide category, `is_inherent`, confidence, rationale, and evidences, and MUST verify that all evidence fragments belong to the intermediate sanitized Markdown.
|
||||
- **FR-033**: System MUST proceed to enrichment only when ECP returns `DIRECT_INHERENT` or `CONTEXTUAL_INHERENT`.
|
||||
- **FR-034**: System MUST block Markdown output and generate a persisted `rejected_ecp` manifest when ECP returns `TANGENTIAL` or `NOT_RELATED` (`ECP_REJECTED`).
|
||||
- **FR-035**: System MUST terminate with `ECP_CLASSIFICATION_FAILED` and block Markdown output if the ECP classifier fails, returns invalid enums, or contains evidence not belonging to the document.
|
||||
- **FR-036**: System MUST verify that any LLM tier used by the ECP classifier uses certified cheap models and records generation telemetry.
|
||||
|
||||
#### Post-ECP Enrichment
|
||||
- **FR-037**: System MUST invoke enrichment only after ECP approval, submitting final title, subtitle (if any), body Markdown, language, and minimal ECP identity.
|
||||
- **FR-038**: System MUST classify entity sentiment as strictly `positive`, `negative`, or `neutral` relative specifically to the ECP entity.
|
||||
- **FR-039**: System MUST generate between 3 and 8 unique tags in the article's native language, grounded in content evidence with candidate evidence IDs. Tag uniqueness MUST be verified via Unicode/NLP libraries without regex. Responses attempting to output or alter body text MUST be rejected.
|
||||
- **FR-040**: System MUST terminate processing with `ENRICHMENT_FAILED` and generate no Markdown if primary and fallback enrichment calls fail.
|
||||
|
||||
#### Model Gateway and Fallback Strategy
|
||||
- **FR-041**: System MUST route runtime LLM invocations through an agnostic Model Gateway supporting logical roles `runtime_primary` and `runtime_fallback`. The gateway MUST accept role, messages/context, schema, timeout, and trace metadata, and return structured output, effective provider, effective model, tokens (input, output, cached), cost, latency, technical status, attempt number, and fallback indicator.
|
||||
- **FR-042**: System MUST configure each role with provider, model, parameters, compatible prompt, expected schema, timeout, and version. Only certified cheap models MAY be configured in runtime roles; powerful/expensive models MUST NOT be configured or called in the runtime.
|
||||
- **FR-043**: System MUST perform limited technical retries on the same provider for: timeout, connection interruption (connection reset), HTTP 429 (with backoff up to configured limit), HTTP 5xx, and empty technical response. System MUST switch immediately to `runtime_fallback` upon semantic failure (invalid schema, grounding violation) or retry exhaustion, without entering semantic retry loops on the same model.
|
||||
- **FR-044**: System MUST use two adapters/configurations (`runtime_primary`, `runtime_fallback`) without implementing smart routers, autonomous model selectors, or dynamic model scoring.
|
||||
- **FR-045**: System MUST apply a conservative deterministic fallback for hygiene only if baseline structural integrity is satisfied (using `selected_extractor` backbone, eliminating structurally invalid elements, without regex and without semantic ad filtering); otherwise, it must terminate with `HYGIENE_FAILED`.
|
||||
|
||||
#### State Machine, Persistence, and Rendering
|
||||
- **FR-046**: System MUST implement orchestration as an explicit Python state machine (`received` → `validated` → `content_cleaned`; `content_cleaned` → `ecp_approved` → `enriched` → `completed_text`; `content_cleaned` → `ecp_rejected`; valid terminal failures → `failed`) persisted in SQLite (WAL mode, short transactions, configurable lock timeout), without using LangChain, LangGraph, agents, planners, workflow frameworks, Postgres, external message queues, or object storage within the runtime. Each state transition MUST record `start_time`, `end_time`, `duration`, and `result`.
|
||||
- **FR-047**: System MUST output a structured JSON result on every invocation and persist a `<fingerprint>.result.json` manifest whenever a deterministic fingerprint is established, including fingerprint, source URL, selected extractor, final status, generate_markdown decision, markdown_path (or null), ECP classification/confidence (or null), provider versions (or null), model versions (or null), prompt versions/hashes (or null), config_version, trace_id (or null), and error/rejection codes.
|
||||
- **FR-048**: System MUST generate `<fingerprint>.md` conditionally (only upon ECP approval and valid enrichment) containing mandatory YAML front matter (`title`, `source_url`, `sentiment`, `tags`, `ecp_qid`, `ecp_canonical_name`, `ecp_category`, `ecp_confidence`) and optional fields (`subtitle`, `author`, `published_at`) omitted when absent.
|
||||
- **FR-049**: System MUST render Markdown body with H1 title, italic subtitle (when present, with NO artificial blank line generated if absent), and editorial blocks in canonical order (headings, paragraphs, bold, italic, blockquotes, lists, links, images), without underline or inline HTML.
|
||||
- **FR-050**: System MUST persist Markdown and manifest through temporary files and atomic renames, and persist SQLite state in the same logical completion unit, verifying content hashes before renaming. Manifest and Markdown MUST NOT be presented as completed while inconsistent. Hash-based reconciliation MUST resolve crashes between file write and SQLite update without duplicate processing. Persistence failures MUST terminate with `PERSISTENCE_FAILED`.
|
||||
|
||||
#### Security
|
||||
- **FR-051**: System MUST treat HTML and article text as untrusted data candidates, never executing scripts embedded in inputs.
|
||||
- **FR-052**: System MUST treat input URLs as data, never accessing or crawling them during runtime execution.
|
||||
- **FR-053**: System MUST construct output filenames strictly from deterministic fingerprints to prevent path traversal attacks.
|
||||
- **FR-054**: System MUST NOT overwrite pre-existing output files from other executions without an exact idempotent fingerprint match.
|
||||
- **FR-055**: System MUST manage credentials strictly via secure environment mechanisms (never CLI parameters, committed configs, or log outputs) and require minimal directory permissions.
|
||||
- **FR-056**: System MUST enforce maximum input size limits, failing in a controlled manner before the provider or applying a previously approved context strategy if exceeded.
|
||||
|
||||
#### Prompts and Context Engineering
|
||||
- **FR-057**: System MUST maintain exactly two atomic prompt responsibilities: `article_content_hygiene` and `article_sentiment_tags`. Prompts MUST be versioned in repository files with semantic version, file hash, expected schemas, and associated Promptfoo test cases. Production and Promptfoo MUST load identical prompt files. Langfuse receives prompt references but is not the primary source. Promptfoo suites MUST explicitly test:
|
||||
- *Accepted Repairs when unmistakable*: mojibake, broken Unicode, accidental spacing, corrupted punctuation, small typos.
|
||||
- *Rejected Repairs*: synonyms, paraphrasing, title improvements, entity name changes without verified encoding defects, dates/numbers/scores alterations, factual corrections, tone changes, quote rewriting.
|
||||
- *Mandatory Assertions*: JSON Schema validation, custom Python validators without regex, candidate IDs belonging to context, expected sets and ordering, URLs belonging to input candidates, precision and recall metrics, enum and cardinality validation, diffs computed with sequence/Unicode libraries without regex, comparison with ground truth reference data, and cost/latency threshold metrics.
|
||||
- **FR-058**: System MUST structure LLM context in strict normative order: 1. system rules; 2. call responsibility; 3. schema and enums; 4. structural context; 5. candidates and evidence; 6. final structured response request. Article content MUST be delimited as data and separated from instructions.
|
||||
- **FR-059**: System MUST strictly exclude from LLM context: full raw JSON when selected fields suffice, full raw HTML when reduced DOM/AST suffices, data from other articles, logs, rejected prior responses (except technical fallback metadata), full ECP when minimal identity suffices, secrets, self-healing instructions, and language-specific semantic keyword examples.
|
||||
|
||||
#### Observability, Logging, and Metrics
|
||||
- **FR-060**: System MUST record Langfuse traces under stable `run_id` with stable spans covering validation, candidate preparation, hygiene, grounding validation, ECP gate, enrichment, rendering, and persistence, without embedding URLs, model names, or dynamic IDs in span names.
|
||||
- **FR-061**: System MUST record each LLM attempt as a separate generation capturing prompt versions/hashes, model/provider, schema version, cached tokens (when available), token counts, calculated cost, latency, timeout status, attempts, technical status, applicable scores, schema results, fallback status, logical role, normalized context sent, structured response, grounding validation result, applied repairs, and rejected repairs (subject to trace content policy), without duplicating raw HTML or full JSON payloads.
|
||||
- **FR-062**: System MUST provide configuration to disable textual content in Langfuse traces while preserving hashes, metrics, and execution status.
|
||||
- **FR-063**: System MUST redact all API keys, authorization headers, and environment secrets from logs, traces, and output files (`trace_redaction_failure_total` = 0). Standard logs MUST omit full ECP, full HTML, and full article text. Repair diffs MUST reside in controlled traces only, never in metric labels.
|
||||
- **FR-064**: System MUST degrade gracefully when Langfuse is unavailable by storing pending telemetry in SQLite (`TELEMETRY_PENDING`), attempting a flush on shutdown, and providing an operational resend routine that returns `telemetry_pending_total` to zero without blocking article processing.
|
||||
- **FR-065**: System MUST emit structured JSON logs capturing timestamp, severity, environment, run_id, fingerprint, state, event, error codes, logical call, provider/model, prompt version, duration, retry/fallback, trace ID, and final status. Stack traces for unexpected errors MUST be logged locally and sanitized.
|
||||
- **FR-066**: System MUST enforce metric label cardinality, forbidding URLs, fingerprints, run IDs, trace IDs, titles, authors, full text, and free tags as metric labels. Metric labels MUST be restricted to: `environment`, `state`, `error_code`, `logical_call`, `provider`, `model`, `prompt_version`, `schema_version`, `language`, `extractor`, `ecp_category`, `reason`, and `source_domain_group`.
|
||||
- **FR-067**: System MUST instrument the normative metric groups with each metric strictly using its specific dimensions defined in the metric catalog:
|
||||
- *Volume/Result*: `article_received_total`, `article_validated_total`, `article_duplicate_total`, `article_completed_text_total`, `article_rejected_ecp_total`, `article_failed_validation_total`, `article_failed_processing_total`.
|
||||
- *Derived Rates*: text completion rate per validated article, ECP rejection rate per validated article, validation failure rate per received article, processing failure rate per validated article, duplication rate per received article.
|
||||
- *Input*: `input_validation_failure_total` (dimensions: `reason`, `schema_version`), `selected_extractor_total`, `selected_extractor_unavailable_total`, `source_language_total`, `source_domain_group_total`, `ecp_version_total`.
|
||||
- *Hygiene*: `hygiene_call_total`, `hygiene_schema_failure_total`, `hygiene_grounding_failure_total`, `hygiene_fallback_total`, `hygiene_deterministic_fallback_total`, `hygiene_terminal_failure_total`, `block_candidate_total`, `block_kept_total`, `block_removed_total`, `link_kept_total`, `image_kept_total`.
|
||||
- *Repairs*: `text_repair_proposed_total`, `text_repair_applied_total`, `text_repair_rejected_total`, `text_repair_category_total`, `text_repair_ambiguous_target_total`, `text_repair_sensitive_change_total`.
|
||||
- *ECP*: `ecp_classification_total`, `ecp_classification_failure_total`, `ecp_pass_total`, `ecp_reject_total`, `ecp_latency_seconds`, `ecp_fallback_tier_total`.
|
||||
- *Enrichment*: `sentiment_total`, `tag_count`, `enrichment_schema_failure_total`, `enrichment_grounding_failure_total`, `enrichment_fallback_total`, `enrichment_terminal_failure_total`.
|
||||
- *LLM*: `llm_request_total` (dimensions: `logical_call`, `provider`, `model`, `status`), `llm_input_tokens_total`, `llm_output_tokens_total`, `llm_cost_total`, `llm_latency_seconds`, `llm_retry_total` (dimensions: `reason`, `provider`, `model`), `llm_fallback_total` (dimensions: `logical_call`, `reason`), `llm_output_validation_failure_total`.
|
||||
- *State/Persistence*: `state_transition_total`, `state_transition_failure_total`, `sqlite_lock_wait_seconds`, `sqlite_busy_failure_total`, `atomic_write_failure_total`, `resume_total`, `idempotent_hit_total`, `orphan_temp_file_total`.
|
||||
- *Observability*: `trace_created_total`, `telemetry_send_failure_total`, `telemetry_pending_total`, `telemetry_flush_failure_total`, `trace_content_disabled_total`, `trace_redaction_failure_total`.
|
||||
- *Capacity*: throughput, concurrency, total duration p50/p95/p99, duration per state p50/p95/p99, CPU, memory, SQLite growth, disk usage separated by outputs and temporary files, lock wait seconds, saturation failures, tokens/cost per received article, tokens/cost per approved Markdown.
|
||||
- **FR-068**: System MUST record objective prompt review signals (`prompt_review_signal_total`) capturing `logical_call`, `prompt_version`, `provider`, `model`, `language`, `source_domain_group`, and `reason` without executing autonomous self-healing.
|
||||
- **FR-069**: System MUST provide minimum dashboards for Runtime Health, Quality, and Future Review Signals.
|
||||
|
||||
#### Testing, CI Quality Gates, and Staging Baselines
|
||||
- **FR-070**: System MUST enforce static AST verification in CI to prevent regex imports or calls in text pipeline modules and test assertions.
|
||||
- **FR-071**: System MUST execute unit tests, contract tests, simulated integration tests (with simulated/mock providers; real remote providers are permitted strictly in authorized offline evaluations and calibration suites), and basic security tests on every pull request.
|
||||
- **FR-072**: System MUST execute Promptfoo evals (in dev/CI), regression of the 20 reference cases, and cost/latency comparisons on every change to prompts, context, schema, or models.
|
||||
- **FR-073**: System MUST execute full golden set evaluation across stratified language/domain/extractor slices (with golden set size justified by observed stability and confidence intervals, holdout never used for few-shot examples, and ground truth covering: full raw input, ECP and version, expected status, title/subtitle/author/date expected or candidates, kept/removed blocks, expected links/images, allowed/forbidden repairs, expected ECP, expected sentiment, accepted tags or a closed evaluation rubric, expected Markdown/structure, discard reason), full security tests on every release, reprocessing tests on every release, fault injection (10 scenarios: primary down, both down, Langfuse unavailable during processing, SQLite lock timeout, disk full, process terminated during write, ECP down, truncated LLM response, orphan temp, telemetry flush failure), and 100 articles/hour load test before promotion.
|
||||
- **FR-074**: System MUST evaluate quality metrics (precision, recall, F1 of kept blocks; metadata accuracy; correct vs unauthorized repair rates; material loss; residual noise; link/image precision; ECP accuracy; sentiment accuracy; tag acceptance) across slices (language, domain, extractor, prompt version, model version) requiring at least 95% pass rate per slice without allowing global averages to hide slice failures. 100% of accepted LLM responses MUST have valid schema.
|
||||
- **FR-075**: System MUST enforce zero-tolerance release gates blocking promotion if any of the 11 critical invariants occurs: `ungrounded_text_total` > 0, `ungrounded_url_total` > 0, `ungrounded_image_total` > 0, `unauthorized_rewrite_total` > 0, `critical_fact_change_total` > 0, `duplicate_output_total` > 0, `lost_article_total` > 0, `secret_exposure_total` > 0, `powerful_runtime_model_call_total` > 0, `online_promptfoo_call_total` > 0, `text_regex_usage_total` > 0.
|
||||
- **FR-076**: System MUST sustain 100 articles/hour load in staging using production-equivalent SQLite and filesystem configurations with zero lost articles, zero duplicate outputs, zero partial files exposed, zero database corruption, stable memory/disk, 100% traces sent or queued, establishing empirical cost and latency baselines (max cost per article, max cost per approved Markdown, p50/p95/p99 latency, provider timeouts, fallback limits, storage limits) for approval prior to go-live.
|
||||
- **FR-077**: System MUST preserve mandatory release evidence artifacts and a complete staging report containing: code version, prompt versions/hashes, providers/models config, Promptfoo config, golden set hashes, per-case and per-slice results, critical violation reports, load test report, formal release sign-off, corpus size and composition (languages, domains, extractors, structures), code/prompt/model/ECP versions, throughput, latency p50/p95/p99 per step and total, cost p50/p95/p99 per article, cost per approved Markdown, fallback rate, resource utilization (CPU, memory, disk, SQLite growth), failures, and outliers.
|
||||
|
||||
#### Production Operations and Runbook Procedures
|
||||
- **FR-078**: System MUST implement preflight checks validating system clock synchronization, prompts and configs belonging strictly to the same release, local configuration reading, prompt file existence & hashes, schema compatibility, SQLite access, filesystem atomic write permissions & directory permissions, minimum disk space, validated active credentials (not just presence), absence of powerful models in runtime roles, Langfuse local configuration (environment, content policy), and ECP classifier/schema access before accepting live traffic. Remote Langfuse connectivity failure MUST NOT block preflight.
|
||||
- **FR-079**: System MUST implement smoke tests executing a versioned fixture to verify end-to-end processing, artifacts, Langfuse trace, cost and latency within approved staging ranges, and idempotent re-execution.
|
||||
- **FR-080**: System MUST follow the complete 11-step deployment sequence (pause new runs, drain active runs, backup SQLite/config, deploy package/prompts/schemas, test migration on copy, preflight, smoke test, verify artifacts/state/trace, release with reduced concurrency, verify metrics, release full volume).
|
||||
- **FR-081**: System MUST support consistent SQLite backup/restore mechanisms, graceful shutdown upon receiving a shutdown signal (stopping new units, completing active unit safely, closing transactions, flushing files/telemetry, preserving pending telemetry, exiting cleanly), periodic reconciliation reports, safe orphan temporary-file cleanup through approved operational/reconciliation routines, manual rollback procedures to previous certified configurations, and certified model/provider rotation.
|
||||
- **FR-082**: System MUST enforce operational retention policies for manifests, Markdown, logs, and traces.
|
||||
- **FR-083**: System MUST provide credential rotation procedures updating environment secrets, running preflight/smoke tests, revoking old credentials, and avoiding altering functional fingerprints when only operational secrets change.
|
||||
- **FR-084**: System MUST provide documented incident procedures for providers, ECP, enrichment, Langfuse, SQLite, disk, grounding, cost, latency, and input validation / producer schema mismatches (checking producer version, comparing with release schema, confirming unit payload, never calling LLM manually, fixing producer or contract via standard release), preserving incident evidence and creating mandatory regression test cases after critical incidents. Production prompts, sentiments, tags, Markdown files, manifests, and SQLite MUST NEVER be edited manually. Reprocessing MUST locate state/fingerprint, verify versions, reuse existing outputs or resume incomplete states under the same configuration, generate distinct fingerprints for new configurations, avoid manual file edits, and never reprocess articles merely to recreate traces.
|
||||
|
||||
---
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Article Input Unit**: Single news article JSON containing `crawled_url`, `error_message`, `extraction_status`, `http_status`, `input_meta`, `page_title`, `selected_extractor` (`trafilatura` | `newspaper4k` | `readability`), individual extractor payloads (`trafilatura`, `newspaper4k`, `readability`), and preserved unknown fields.
|
||||
- **Entity Context Profile (ECP) Canonical Schema Reference**: Canonical versioned profile managed exclusively by the ECP module; referenced by the runtime without duplicating schema definitions.
|
||||
- **Candidate Object**: Identifiable structural unit (title, subtitle, author, date, block, heading, list, quote, link, or image) with an opaque ID without quality judgment (stable within execution), extractor origin, source field, content hash, and cross-extractor equivalences.
|
||||
- **Text Repair Operation**: Controlled micro-edit specifying target candidate ID, exact original fragment, replacement fragment, category (`encoding` | `unicode` | `spacing` | `punctuation_corruption` | `obvious_typo`), and short rationale.
|
||||
- **State Machine Record**: SQLite-persisted lifecycle state containing `fingerprint`, `current_state` (`received` → `validated` → `content_cleaned`; `content_cleaned` → `ecp_approved` → `enriched` → `completed_text`; `content_cleaned` → `ecp_rejected` as terminal state without Markdown; valid terminal failures → `failed`), `timestamps` (`start_time`, `end_time`, `duration`), `result`, `output_paths`, `file_hashes`, `functional_versions`, `terminal_error`, and `pending_telemetry`.
|
||||
- **Output Manifest (`<fingerprint>.result.json`)**: Machine-readable summary containing fingerprint, source URL, selected extractor, final status (`completed_text` | `rejected_ecp` | `failed_validation` | `failed_processing`), generate_markdown decision, markdown_path (or null), ECP classification/confidence (or null), provider versions (or null), model versions (or null), prompt versions/hashes (or null), config_version, trace_id (or null), and error/rejection codes.
|
||||
- **Published Markdown (`<fingerprint>.md`)**: Markdown document with structured YAML front matter and clean, grounded editorial body.
|
||||
- **Telemetry Event**: Queued observability payload stored in SQLite (`TELEMETRY_PENDING`) for deferred transmission when Langfuse is unavailable.
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001 (Zero Hallucination)**: 100% of published text, URLs, and images are traceable to input candidates or approved repairs (`ungrounded_text_total` = 0, `ungrounded_url_total` = 0, `ungrounded_image_total` = 0).
|
||||
- **SC-002 (Zero Critical Fact Corruption)**: 0 unauthorized rewrites or alterations to names, dates, numbers, scores, quotes, or facts across all evaluations (`unauthorized_rewrite_total` = 0, `critical_fact_change_total` = 0).
|
||||
- **SC-003 (Zero Duplication & Data Integrity)**: 0 duplicate outputs generated for identical input fingerprints and functional configurations; 0 lost articles (`duplicate_output_total` = 0, `lost_article_total` = 0).
|
||||
- **SC-004 (End-to-End Quality Pass Rate by Slice)**: At least 95% end-to-end pass rate across the reference golden dataset and across every stratified language, domain, extractor, prompt version, and model version slice without global average masking. 100% of accepted LLM responses have valid schema.
|
||||
- **SC-005 (Telemetry Completeness)**: 100% of validated executions have complete traces either delivered to Langfuse or preserved in SQLite as pending telemetry (`trace_created_total` / `article_validated_total` = 100%); `trace_redaction_failure_total` = 0; `telemetry_pending_total` returns to zero after recovery.
|
||||
- **SC-006 (Throughput and Concurrency)**: Sustained throughput of at least 100 articles per hour under realistic concurrency without data loss, partial file exposure, or database corruption, with controlled handling of lock waits and measured `sqlite_lock_wait_seconds` and `sqlite_busy_failure_total`.
|
||||
- **SC-007 (Empirical SLO Approval)**: Measured p50, p95, and p99 latency (per step and total), cost per received article, cost per approved Markdown, provider timeouts, acceptable fallback threshold, and storage limits established in staging and approved prior to production go-live.
|
||||
- **SC-008 (Strict Regex & Dependency Discipline)**: 0 occurrences of regular expression imports/calls within text pipeline modules and test assertions (`text_regex_usage_total` = 0); all dependencies strictly justified and locked in project lockfile.
|
||||
- **SC-009 (Model Cost Control)**: 0 invocations of powerful/expensive LLMs within runtime roles (`powerful_runtime_model_call_total` = 0).
|
||||
- **SC-010 (Operational Readiness)**: 100% completion of preflight checks, smoke tests, 11-step deployment sequence verification, backup/restore verifications, graceful shutdown handling, reconciliation routines, and manual rollback drills.
|
||||
|
||||
---
|
||||
|
||||
## Traceability Matrix
|
||||
|
||||
| Documento Fonte | Cláusula / Tópico Normativo | Cobertura Específica na Spec |
|
||||
| :--- | :--- | :--- |
|
||||
| **01_PRD** | 1–2: Contexto e Objetivo (Artigo único, ECP, pico 100 art/h) | US1, US6, FR-003, FR-005, FR-076, SC-006 |
|
||||
| **01_PRD** | 3: Princípio Mestre (Zero complexidade supérflua, orquestração Python direta) | FR-001, FR-046, SC-008, Assumptions |
|
||||
| **01_PRD** | 4.1: Fundamentação e Proibição de Invenção/Alucinação | US2, FR-023, FR-026, SC-001, SC-002 |
|
||||
| **01_PRD** | 4.2: Proibição Estrita de Regex e Listas Manuais de Palavras | US1, FR-015, FR-016, FR-070, SC-008 |
|
||||
| **01_PRD** | 4.3: Higienização LLM Obrigatória Mesmo em Consenso 100% | US2, FR-022, Edge Cases |
|
||||
| **01_PRD** | 4.4: Modelos Baratos no Runtime (Zero modelos caros) | US6, FR-042, SC-009, Assumptions |
|
||||
| **01_PRD** | 5: Escopo Incluído / Fora de Escopo (Vídeo/galeria upstream, sem self-healing) | FR-003, FR-068, Assumptions, Edge Cases |
|
||||
| **01_PRD** | 6–7: Atores e Unidade de Processamento (1 artigo + ECP snapshot) | US1, FR-003, FR-005, Key Entities |
|
||||
| **01_PRD** | 8: Contrato do Artigo (`selected_extractor`, campos dos extratores) | US1, FR-006, FR-007, FR-010, FR-011 |
|
||||
| **01_PRD** | 9: Contrato do ECP (Snapshot canônico, Gate, 4 classificações) | US1, US3, FR-005, FR-031, FR-032, FR-033, FR-034 |
|
||||
| **01_PRD** | 10: Contrato de Saída (JSON, 4 status, Markdown YAML front matter) | US5, FR-047, FR-048, FR-049, Key Entities |
|
||||
| **01_PRD** | 11–13: Fluxo, Preparação Determinística, Resolução URL/Data, Candidatos | US1, FR-008, FR-014, FR-017, FR-018, FR-019, FR-020, FR-021 |
|
||||
| **01_PRD** | 14: Higienização Extrativa LLM (Entrada, Saída por IDs, Regras Editoriais) | US2, FR-023, FR-024, FR-025, FR-026, FR-030 |
|
||||
| **01_PRD** | 15: Pequenos Reparos Textuais (5 categorias, diffs, reversibilidade, exceção encoding) | US2, FR-027, FR-028, FR-029, Key Entities |
|
||||
| **01_PRD** | 16: Imagens e Links Editoriais (Grounding estrutural, alt/caption, links válidos) | US2, FR-026, FR-030 |
|
||||
| **01_PRD** | 17–18: Gate ECP e Enriquecimento (Sentimento relativo, 3-8 tags nativas) | US3, US4, FR-031, FR-032, FR-037, FR-038, FR-039, FR-040 |
|
||||
| **01_PRD** | 19–20: Model Gateway, Idempotência e Persistência Atômica | US5, US6, FR-008, FR-009, FR-041, FR-043, FR-050 |
|
||||
| **01_PRD** | 21–22: Observabilidade Langfuse e Promptfoo fora do runtime | US7, US8, FR-057, FR-060, FR-064, FR-072 |
|
||||
| **01_PRD** | 23–25: Requisitos Funcionais, NFRs e 16 Códigos Mínimos de Erro | FR-001 a FR-084, Acceptance Scenarios, Edge Cases |
|
||||
| **01_PRD** | 26–29: Critérios de Aceite do Produto, Métricas e DoD | SC-001 a SC-010, FR-067, FR-075, FR-076 |
|
||||
| **02_Arquitetura** | 1–6: Propósito, Direcionadores, Limites e Componentes | US1 a US9, FR-001 a FR-084 |
|
||||
| **02_Arquitetura** | 7: Orquestração (Máquina de estados Python, SQLite WAL, sem frameworks/agentes) | FR-046, Key Entities |
|
||||
| **02_Arquitetura** | 8–10: Módulos, Política de Dependências, Contratos Versionados | FR-001, FR-002, FR-004, FR-015, SC-008, Assumptions |
|
||||
| **02_Arquitetura** | 11–14: Validação sem chamadas remotas, Fingerprint, Parsing, Candidatos | US1, FR-007, FR-008, FR-012, FR-014, FR-019, FR-020, FR-021 |
|
||||
| **02_Arquitetura** | 15–19: Higienização, Reparos, Assembler, ECP Adapter, Enriquecimento | US2, US3, US4, FR-023, FR-024, FR-027, FR-031, FR-032, FR-037 |
|
||||
| **02_Arquitetura** | 20–22: Model Gateway (2 adapters, sem router), Prompts no repo, Escrita Atômica | US5, US6, FR-041, FR-042, FR-044, FR-048, FR-050, FR-057 |
|
||||
| **02_Arquitetura** | 23–26: Observabilidade, Logs JSON, Segurança, Concorrência 100 art/h | US7, FR-051 a FR-056, FR-060 a FR-066, FR-076 |
|
||||
| **02_Arquitetura** | 27–33: Falhas, Deploy, CI/CD, Simplicidade, Riscos | US8, US9, FR-045, FR-070 a FR-077, FR-078 a FR-084 |
|
||||
| **03_ADRs** | ADR-001: Separação de Runtime e Self-Healing | FR-068, Assumptions |
|
||||
| **03_ADRs** | ADR-002: Início após seleção do extrator (Sem recálculo) | FR-006, Edge Cases |
|
||||
| **03_ADRs** | ADR-003: Orquestração direta em Python (Sem LangChain/LangGraph/agentes) | FR-046 |
|
||||
| **03_ADRs** | ADR-004: Gateway agnóstico com apenas modelos baratos no runtime | US6, FR-041, FR-042, SC-009 |
|
||||
| **03_ADRs** | ADR-005: Proibição de regex e palavras-chave manuais em decisões textuais | US1, FR-015, FR-016, FR-070, SC-008 |
|
||||
| **03_ADRs** | ADR-006: LLM seleciona IDs e propõe reparos (Não regenera o artigo) | US2, FR-023, FR-024, FR-027, FR-030 |
|
||||
| **03_ADRs** | ADR-007: ECP obrigatório antes de toda saída editorial | US3, FR-005, FR-031, FR-033, FR-034 |
|
||||
| **03_ADRs** | ADR-008: Langfuse no runtime e Promptfoo no CI | US7, US8, FR-057, FR-060, FR-064, FR-072 |
|
||||
| **03_ADRs** | ADR-009: Persistência em SQLite (WAL) e saídas no filesystem | US5, FR-046, FR-050 |
|
||||
| **03_ADRs** | ADR-010: Resultado estruturado sempre e Markdown condicional | US5, FR-047, FR-048 |
|
||||
| **03_ADRs** | ADR-011: Definição de SLOs de custo e latência a partir de staging | US8, FR-076, SC-007 |
|
||||
| **04_Plano_Testes** | 1–7: Objetivo, Princípios, Camadas de Teste, Dados, Golden Set (Contrato 4.2 e 4.4), Holdout, Gates | US8, FR-071, FR-072, FR-073, FR-074, FR-075, SC-004 |
|
||||
| **04_Plano_Testes** | 8: Matriz IN (IN-001 a IN-015: Contratos de entrada e validações locais) | US1, FR-003, FR-005, FR-006, FR-007, FR-010, FR-011, FR-012 |
|
||||
| **04_Plano_Testes** | 9: Matriz ID (ID-001 a ID-010: Fingerprint, idempotência, concorrência e reconciliação) | US1, US5, FR-008, FR-009, FR-050 |
|
||||
| **04_Plano_Testes** | 10: Matriz PAR (PAR-001 a PAR-010: Parsing estrutural, malformed HTML, JSON-LD, sem regex) | US1, US8, FR-014, FR-015, FR-021, FR-070 |
|
||||
| **04_Plano_Testes** | 11: Matriz CAN (CAN-001 a CAN-010: Candidatos, IDs únicos, similaridade, sem imagem auto) | US1, FR-014, FR-019, FR-020, FR-030 |
|
||||
| **04_Plano_Testes** | 12: Matriz HYG (HYG-001 a HYG-021: 10 passos do harness, grounding, fallback determinístico) | US2, US6, FR-022 a FR-030, FR-043, FR-045 |
|
||||
| **04_Plano_Testes** | 13: Matriz REP (REP-001 a REP-016: Reparos permitidos, proibidos, sensíveis e reversão) | US2, FR-027, FR-028, FR-029, FR-030 |
|
||||
| **04_Plano_Testes** | 14: Matriz ECP (ECP-001 a ECP-009: Gate ECP, evidências no doc e cheap tiers) | US3, FR-031 a FR-036 |
|
||||
| **04_Plano_Testes** | 15: Matriz ENR (ENR-001 a ENR-009: Sentimento relativo, tags nativas sem duplicatas, falhas) | US4, FR-037 a FR-040 |
|
||||
| **04_Plano_Testes** | 16: Matriz OUT (OUT-001 a OUT-012: Markdown YAML, sem linha artificial sem subtítulo, escrita atômica) | US5, FR-047, FR-048, FR-049, FR-050 |
|
||||
| **04_Plano_Testes** | 17: Matriz LLM (LLM-001 a LLM-012: Retries técnicos autorizados, fallbacks e gates de modelos) | US6, FR-041, FR-042, FR-043, FR-044, FR-045 |
|
||||
| **04_Plano_Testes** | 18: Matriz OBS (OBS-001 a OBS-011: Traces, spans, degradação graciosa e logs) | US7, FR-060 a FR-065 |
|
||||
| **04_Plano_Testes** | 19: Matriz SEC (SEC-001 a SEC-008: Prompt injection, traversal, pre-existing files, secrets, limites) | FR-051 a FR-056 |
|
||||
| **04_Plano_Testes** | 20: Teste de Carga (100 art/h, sem perda, estabilidade de memória e disco) | US8, FR-076, SC-006 |
|
||||
| **04_Plano_Testes** | 21: Fault Injection (FLT-001 a FLT-010: 10 cenários completos incluindo Langfuse down) | US8, FR-073, FR-084 |
|
||||
| **04_Plano_Testes** | 22–25: Promptfoo em CI, Ordem CI/CD, Evidências Preservadas e Conclusão | US8, FR-057, FR-072, FR-077 |
|
||||
| **05_Metricas** | 1–4: KPIs do Produto e 11 Invariantes Críticas (Meta zero) | US8, SC-001 a SC-010, FR-075 |
|
||||
| **05_Metricas** | 5–11: Métricas de Volume, Entrada, Higienização, Reparos, ECP, Enriquecimento, LLM | US7, FR-067 |
|
||||
| **05_Metricas** | 12: Sinais para Revisão Futura de Prompt (`prompt_review_signal_total` com 7 dimensões) | US7, FR-068 |
|
||||
| **05_Metricas** | 13–15: Métricas de Persistência, Observabilidade e Capacidade (Durações, Disco, Locks) | US7, FR-067, SC-006 |
|
||||
| **05_Metricas** | 16: Baseline de Staging e Relatório Completo Obrigatório | US8, FR-076, FR-077, SC-007 |
|
||||
| **05_Metricas** | 17–20: Logs Estruturados, 3 Dashboards Mínimos e Dimensões por Métrica no Catálogo | US7, FR-063, FR-065, FR-066, FR-067, FR-069 |
|
||||
| **06_Runbook** | 1–6: Princípios Operacionais, Matriz de Responsabilidades (4 papéis) e Pré-requisitos | US9, FR-013, FR-078, FR-084, Key Entities |
|
||||
| **06_Runbook** | 7–11: Checklist de Release, Sequência de Implantação 11 Passos, Preflight e Smoke Test | US9, FR-078, FR-079, FR-080 |
|
||||
| **06_Runbook** | 12–14: Monitoramento, Logs/Correlação e Reprocessamento Idempotente | US7, US9, FR-009, FR-065, FR-081, FR-084 |
|
||||
| **06_Runbook** | 15–25: Diagnósticos de Incidentes (Todos subsistemas + Entrada/Produtor) e Proibição Edição Manual | US9, FR-064, FR-084, Edge Cases |
|
||||
| **06_Runbook** | 26: Rollback Manual (Gatilhos, Procedimento e Retomada) | US9, FR-081 |
|
||||
| **06_Runbook** | 27–29: Troca Certificada de Modelo, Rotação de Credenciais e Backup/Retenção | US9, FR-081, FR-082, FR-083 |
|
||||
| **06_Runbook** | 30–32: Reconciliação, Shutdown Controlado e Prontidão Operacional | US9, FR-081, SC-010 |
|
||||
| **07_Prompt_Harness** | 1–5: 2 Prompts Atômicos, Regras Comuns, Ordem do Contexto (6 blocos) e Exclusões Estritas | US2, US4, FR-057, FR-058, FR-059 |
|
||||
| **07_Prompt_Harness** | 6–8: Prompt `article_content_hygiene`, Regras Normativas de Reparos e 10 Passos do Harness | US2, FR-024, FR-025, FR-027, FR-028, FR-029, FR-030 |
|
||||
| **07_Prompt_Harness** | 9: Prompt `article_sentiment_tags` e Harness de Enriquecimento (Sem corpo, validação tags) | US4, FR-037, FR-038, FR-039, FR-040 |
|
||||
| **07_Prompt_Harness** | 10–13: Política de Provider, Casos Promptfoo de Reparo/10 Assertions, Langfuse e Aceite | US6, US7, US8, FR-041, FR-042, FR-044, FR-057, FR-060, FR-061, FR-072 |
|
||||
|
||||
---
|
||||
|
||||
## Assumptions
|
||||
|
||||
- Upstream crawling, HTML fetching, and multi-extractor execution (`trafilatura`, `newspaper4k`, `readability`) as well as `selected_extractor` calculation and video/gallery filtering are performed by prior pipeline stages and are out of scope.
|
||||
- Self-healing prompt optimization, automated prompt mutation, LLM-as-a-judge for prompt improvement, canary deployments, and auto-rollback belong to a distinct future subproject; the runtime only emits telemetry review signals (`prompt_review_signal_total`).
|
||||
- All LLM providers configured in runtime roles (`runtime_primary`, `runtime_fallback`) are low-cost models certified via Promptfoo and supporting structured JSON schema outputs.
|
||||
- The local filesystem and SQLite (WAL mode, short transactions, configurable lock timeout) provide the persistence and concurrency foundation for the target workload of 100 articles/hour.
|
||||
- Execution occurs in the Python version supported by the repository with standard dependencies specified in the project lockfile.
|
||||
@@ -0,0 +1,341 @@
|
||||
# Implementation Tasks: Article Consolidation and Hygiene Runtime
|
||||
|
||||
**Feature**: Article Consolidation and Hygiene Runtime (`specs/006-article-consolidation-runtime/spec.md`)
|
||||
**Branch**: `006-article-consolidation-runtime` | **Date**: 2026-08-23 | **Plan**: [`plan.md`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/006-article-consolidation-runtime/plan.md)
|
||||
**Status**: Ready for Execution
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure & Tooling)
|
||||
|
||||
**Purpose**: Project initialization, dependency management, and quality verification tooling.
|
||||
|
||||
- [x] T001 Initialize the package structure and lockfile using only dependencies approved by the implementation plan in `pyproject.toml`, recording for every new dependency: requirement served, standard-library alternative, security impact, maintenance impact, license, size impact, and startup impact
|
||||
- [x] T002 [P] Implement multi-parser static policy verification script in `tests/scripts/check_zero_regex.py`: checking Python AST for imports and direct calls of `re` or any regular-expression engine/API in the scoped text-processing modules, including aliases, without inspecting internals of transitive dependencies, rejecting `pattern` keys in JSON schemas via JSON parser, and validating Promptfoo YAML configurations via YAML parser (failing on regex assertions, semantic `contains`/`not-contains` assertions, LLM-as-a-judge for grounding, powerful models as judge, and configurations relying solely on global averages without per-case and per-slice gates)
|
||||
- [x] T003 [P] Configure Promptfoo test environment and suite settings in `evals/promptfoo.config.yaml` strictly following policy constraints (no `contains`/`not-contains` semantic decisions, no LLM-as-a-judge for grounding, no powerful models as judge, and per-case and per-slice assertion gates)
|
||||
- [x] T004 [P] Create initial 20-case reference regression dataset in `evals/reference_20/`, converting each of the 20 reference articles into an individual unit file associated with a valid, versioned canonical ECP snapshot per Test Plan §4.1
|
||||
- [x] T005 [P] Create local validation configuration fixture in `runtime_config.local.json`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites & Shared Core)
|
||||
|
||||
**Purpose**: Core infrastructure, base models, SQLite WAL store, atomic file writer, manifest generator, and configuration engine that MUST be complete before pipeline execution.
|
||||
|
||||
> **CRITICAL**: No user story implementation can begin until this foundational phase is complete.
|
||||
|
||||
- [x] T006 Implement configuration loading, validation, and exact-byte SHA-256 hash verification in `src/core/config.py`
|
||||
- [x] T007 [P] Implement input byte size limiter with fail-before-provider policy in `src/core/limits.py`
|
||||
- [x] T008 [P] Implement deterministic canonical SHA-256 execution fingerprint calculator in `src/core/fingerprint.py`
|
||||
- [x] T009 [P] Implement `CandidateObject` and text repair dataclasses in `src/candidate/models.py`
|
||||
- [x] T010 Implement SQLite WAL store in `src/storage/sqlite_store.py` with short transactions, busy timeout, native backup/restore API, and atomic fingerprint claim logic (checking completed fingerprint before remote calls, returning existing result, and safely resuming/reusing concurrent executions)
|
||||
- [x] T011 Implement Python explicit state machine and SQLite transition logger in `src/core/state_machine.py`
|
||||
- [x] T012 [P] Implement structured JSON logging with all normative fields, structural authorization-header sanitization, and exact replacement of known environment-secret values in `src/observability/structured_logger.py`, strictly omitting full ECP, full HTML, and full article text
|
||||
- [x] T013 [P] Implement the atomic filesystem writer and shared manifest generator in `src/storage/file_store.py`, writing temporary files in the same destination filesystem, flushing, closing, verifying exact SHA-256 hashes, and performing atomic rename (`os.replace`) without cross-filesystem moves, complying with `manifest-output.schema.json` and the 16 normative error codes (shared across `completed_text`, `rejected_ecp`, `failed_validation`, and `failed_processing`)
|
||||
- [x] T014 [P] Implement contract tests for runtime configuration in `tests/contract/test_runtime_config_contract.py`
|
||||
- [x] T015 [P] Implement contract tests for manifest output schema in `tests/contract/test_manifest_output_contract.py`
|
||||
|
||||
**Checkpoint**: Foundation ready — Model Gateway and Pipeline components can now proceed.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 6 - Model Gateway Infrastructure & Cheap Model Enforcement
|
||||
|
||||
**Goal**: Agnostic Model Gateway, provider adapters, technical retries, and preflight cheap model enforcement (MUST exist before any LLM hygiene or enrichment call).
|
||||
|
||||
### Tests for Model Gateway
|
||||
|
||||
- [x] T016 [P] [US6] Implement unit tests for Model Gateway client and adapters in `tests/unit/test_model_gateway.py` covering normative scenarios `LLM-001` to `LLM-012` (logical roles, pricing, token tracking, timeouts)
|
||||
- [x] T017 [P] [US6] Implement fault injection tests for gateway transient errors and failovers in `tests/fault_injection/test_gateway_faults.py` covering provider fault scenarios
|
||||
|
||||
### Implementation for Model Gateway
|
||||
|
||||
- [x] T018 [US6] Implement agnostic Model Gateway client managing logical roles (`runtime_primary`, `runtime_fallback`) and token pricing calculations in `src/gateway/client.py`
|
||||
- [x] T019 [US6] Implement minimal HTTP adapters for Groq and DeepSeek using `httpx` in `src/gateway/adapters.py`
|
||||
- [x] T020 [US6] Implement limited technical retries for timeout, connection interruption/reset, HTTP 429 with configured backoff up to limit, HTTP 5xx, and empty technical responses in `src/gateway/client.py`
|
||||
- [x] T021 [US6] Implement immediate semantic fallback from `runtime_primary` to `runtime_fallback` on schema or grounding failure without retrying on the same model in `src/gateway/client.py`
|
||||
- [x] T022 [US6] Implement a single, reusable certified-configuration validation in `src/core/config.py` rejecting any uncertified or powerful models across all runtime roles (including any internal ECP LLM)
|
||||
|
||||
**Checkpoint**: Model Gateway implementation and configuration enforcement are ready for pipeline integration; production certification occurs only after Phase 12 gates.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 1 - Single Article Ingestion, Contract Validation, and Candidate Preparation (Priority: P1)
|
||||
|
||||
**Goal**: Ingest single article units, validate contracts (Article and ECP) locally before any remote call, compute deterministic fingerprint, reject batch wrappers, and extract structured candidates without regular expressions.
|
||||
|
||||
**Independent Test**: Provide single article JSON objects (valid, corrupt, batch wrapper) and ECP snapshots, verifying schema validation, SQLite state initialization (`received`, `validated`), deterministic candidate ID generation, and immediate pre-remote termination with exact error codes.
|
||||
|
||||
### Tests for User Story 1
|
||||
|
||||
- [x] T023 [P] [US1] Implement contract test for Article Input schema against all 20 real reference units in `tests/contract/test_article_input_contract.py`
|
||||
- [x] T024 [P] [US1] Implement contract test for ECP Snapshot schema and local `referencing.Registry` resolution in `tests/contract/test_ecp_snapshot_contract.py`
|
||||
- [x] T025 [P] [US1] Implement contract test for Candidates Payload schema in `tests/contract/test_candidates_payload_contract.py`
|
||||
- [x] T026 [P] [US1] Implement unit tests for input limits, validation, and error code mapping in `tests/unit/test_input_limits.py` covering scenarios `IN-001` to `IN-015` and proving zero remote provider, remote Langfuse, or classifier calls on local failure
|
||||
- [x] T027 [P] [US1] Implement unit tests for deterministic fingerprint calculation and idempotency claims in `tests/unit/test_fingerprint.py` covering scenarios `ID-001` to `ID-010`
|
||||
- [x] T028 [P] [US1] Implement unit tests for candidate extraction without regex in `tests/unit/test_candidate_parser.py` covering scenarios `PAR-001` to `PAR-010`
|
||||
- [x] T029 [P] [US1] Implement unit tests for cross-extractor sequence equivalence mapping in `tests/unit/test_equivalence_mapping.py` covering scenarios `CAN-001` to `CAN-010`
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [x] T030 [P] [US1] Create executable article and ECP fixtures in `examples/sample_article_valid.json`, `examples/sample_article_tangential.json`, and `examples/sample_ecp_snapshot.json`
|
||||
- [x] T031 [US1] Implement local canonical ECP schema resolution and registration via `referencing.Registry` (disabling HTTP network fetching) in `src/ecp/adapter.py`
|
||||
- [x] T032 [US1] Implement local pre-call input validation and batch wrapper rejection (`"articles": false`) in `src/core/config.py`, preserving unknown fields in the recorded original input while ignoring them during processing
|
||||
- [x] T033 [US1] Implement structural candidate parsing in `src/candidate/parser.py` using DOM for HTML, CommonMark AST for Markdown, JSON parsing for JSON-LD, URL parsing, Unicode normalization, and an appropriate multilingual tokenizer/segmenter and language detector, handling malformed HTML safely and invalid JSON-LD through a controlled warning
|
||||
- [x] T034 [US1] Implement non-destructive candidate equivalence mapping using `difflib.SequenceMatcher` in `src/candidate/equivalence.py`
|
||||
- [x] T035 [US1] Implement deterministic source URL and publication date resolution using the exact normative priorities and date-consensus rule, and prepare title, subtitle, and author candidates using their normative source priorities in `src/candidate/parser.py`, omitting invalid dates and strictly forbidding delimiter-based author splitting
|
||||
- [x] T036 [US1] Connect initial validation to state machine `received → validated` in `src/core/state_machine.py`, ensuring transition only occurs after article, ECP, config, `selected_extractor`, size limit, and minimum content checks pass
|
||||
- [x] T037 [US1] Implement CLI ingestion entrypoint in `src/cli/consolidate.py` with full idempotency checks (querying fingerprint before LLM, returning existing result on match, resuming incomplete runs, claiming atomic execution), emitting structured JSON, complete manifest on stdout when fingerprint exists, technical envelope on unparseable JSON, persisting manifest, and enforcing exact exit codes (`0`: completed/rejected_ecp, `1`: invalid article/ECP, `2`: config/preflight error, `3`: failed processing, `4`: persistence failure)
|
||||
|
||||
**Checkpoint**: User Story 1 is independently functional, validating contracts and preparing candidates locally.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 2 - Mandatory LLM Extractive Hygiene & Controlled Text Repairs (Priority: P1)
|
||||
|
||||
**Goal**: Execute 100% LLM extractive hygiene over candidate payloads, enforce the 10-step validation harness, permit only 5 closed micro-repair categories, and reject ungrounded edits without regex.
|
||||
|
||||
**Independent Test**: Feed candidate payloads with consensus, divergence, and noise into the hygiene harness, verifying that the LLM returns only candidate IDs and repairs, ungrounded IDs trigger `GROUNDING_VIOLATION`, invalid repairs are discarded with originals preserved, and valid intermediate Markdown is assembled.
|
||||
|
||||
### Tests for User Story 2
|
||||
|
||||
- [x] T038 [P] [US2] Implement contract test for Hygiene Response schema in `tests/contract/test_hygiene_response_contract.py`
|
||||
- [x] T039 [P] [US2] Implement contract test for Repair Operations schema in `tests/contract/test_repair_operations_contract.py`
|
||||
- [x] T040 [P] [US2] Implement unit tests for 10-step hygiene validation harness in `tests/unit/test_hygiene_harness.py` covering scenarios `HYG-001` to `HYG-021` (testing context exclusions, grounding enforcement, and candidate ID validation)
|
||||
- [x] T041 [P] [US2] Implement unit tests for controlled text repairs without regex in `tests/unit/test_repairs_validator.py` covering scenarios `REP-001` to `REP-016` (5 closed categories, sensitive entity protection, exact fragment targeting, and Unicode/NLP-based diff validation without uncalibrated numeric thresholds)
|
||||
- [x] T042 [P] [US2] Implement Promptfoo evaluation suite for `article_content_hygiene` prompt in `evals/promptfoo.config.yaml` validating context exclusions (no raw JSON, no full HTML, no logs, no secrets, no self-healing, no semantic `contains` assertions)
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [x] T043 [P] [US2] Author normative versioned prompt in `prompts/article_content_hygiene.v1.txt` following the exact 6-block ordering (Doc 07 §5.3)
|
||||
- [x] T044 [US2] Implement the minimal candidate/context projection builder in `src/hygiene/harness.py`, excluding full raw JSON, full HTML, other-article data, logs, secrets, full ECP when minimal identity is sufficient, rejected prior responses except required technical fallback metadata, self-healing instructions, and language-specific semantic keyword examples, while delimiting article content strictly as untrusted data
|
||||
- [x] T045 [US2] Implement the micro-repair validator in `src/hygiene/repairs.py` enforcing the 5 closed categories, exact-fragment targeting, Unicode/NLP-based comparison, sensitive-entity preservation, and audit decisions without regex (no quantitative similarity threshold unless one is later approved through the golden-set evaluation)
|
||||
- [x] T046 [US2] Implement 10-step hygiene harness in `src/hygiene/harness.py` validating candidate IDs, ordering, links/images, minimum content, and grounding
|
||||
- [x] T047 [US2] Implement grounded intermediate Markdown assembler in `src/hygiene/assembler.py`
|
||||
- [x] T048 [US2] Implement decoupling between schema failures (semantic fallback) and grounding violations (immediate invalidation) in `src/hygiene/harness.py`
|
||||
- [x] T049 [US2] Implement the conservative deterministic hygiene fallback in `src/hygiene/harness.py` using only the `selected_extractor` structural backbone, removing only structurally invalid elements, without regex, keyword dictionaries, semantic advertisement filtering, or content-quality inference; use it only when grounding and minimum-content requirements are satisfied, otherwise terminate with `HYGIENE_FAILED`
|
||||
- [x] T050 [US2] Connect hygiene stage to state machine `validated → content_cleaned` in `src/core/state_machine.py`
|
||||
- [x] T051 [US2] Integrate hygiene harness execution and error handling into `src/cli/consolidate.py`
|
||||
|
||||
**Checkpoint**: User Stories 1 and 2 operate together, performing grounded extractive hygiene and controlled repairs.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: User Story 3 - Mandatory ECP Gate & Relevance Enforcement (Priority: P1)
|
||||
|
||||
**Goal**: Evaluate intermediate sanitized Markdown against canonical ECP Snapshot using `src.classifier.InherenceClassifier`, transitioning inherent articles to `ecp_approved` and non-inherent articles to `ecp_rejected` with zero Markdown generated.
|
||||
|
||||
**Independent Test**: Submit intermediate Markdown to ECP adapter with profiles across all 4 categories (`DIRECT_INHERENT`, `CONTEXTUAL_INHERENT`, `TANGENTIAL`, `NOT_RELATED`), verifying that only inherent articles proceed to enrichment, while non-inherent articles persist `<fingerprint>.result.json` with status `rejected_ecp` and produce no `.md` file.
|
||||
|
||||
### Tests for User Story 3
|
||||
|
||||
- [x] T052 [P] [US3] Implement unit tests for ECP adapter invoking `InherenceClassifier` in `tests/unit/test_ecp_adapter.py` covering scenarios `ECP-001` to `ECP-009` (full output validation, grounded evidence check, tier tracking, cheap model enforcement)
|
||||
- [x] T053 [P] [US3] Implement integration test for ECP rejection producing zero Markdown files in `tests/integration/test_ecp_rejection_flow.py` covering scenario `OUT-008`
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [x] T054 [US3] Implement the ECP classification adapter in `src/ecp/adapter.py` invoking `src.classifier.InherenceClassifier` through its public contract and certified configuration, validating `category`, `is_inherent`, `confidence`, `rationale`, and `evidences`, asserting that all evidence fragments belong to the intermediate Markdown, and recording any classifier tier or LLM generation exposed by the classifier
|
||||
- [x] T055 [US3] Make the ECP adapter consume the shared certified-configuration validation from `src/core/config.py`, verifying the ECP classifier configuration against packaged release metadata without duplicating hash or certification logic
|
||||
- [x] T056 [US3] Connect ECP inherence gate to state machine `content_cleaned → ecp_approved | ecp_rejected` in `src/core/state_machine.py`
|
||||
- [x] T057 [US3] Implement `rejected_ecp` terminal flow writing manifest with status `rejected_ecp` (`ECP_REJECTED`) and strictly omitting Markdown output in `src/storage/file_store.py`
|
||||
- [x] T058 [US3] Integrate ECP gate execution and error handling into `src/cli/consolidate.py`
|
||||
|
||||
**Checkpoint**: Core pipeline evaluates inherence and enforces the strict ECP publishing gate.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: User Story 4 - Post-ECP Enrichment: Entity Sentiment and Native Language Tags (Priority: P2)
|
||||
|
||||
**Goal**: Enrich ECP-approved articles with entity-relative sentiment and 3 to 8 native language tags supported by textual evidence IDs, strictly decoupled from body text.
|
||||
|
||||
**Independent Test**: Submit approved intermediate Markdown and minimal ECP identity (`qid`, `canonical_name`) to enrichment harness, verifying sentiment extraction, tag bounding (3–8), NLP uniqueness without regex, evidence grounding, and failure handling without body modification.
|
||||
|
||||
### Tests for User Story 4
|
||||
|
||||
- [x] T059 [P] [US4] Implement contract test for Enrichment Response schema in `tests/contract/test_enrichment_response_contract.py`
|
||||
- [x] T060 [P] [US4] Implement unit tests for entity sentiment and native tags validator in `tests/unit/test_enrichment_harness.py` covering scenarios `ENR-001` to `ENR-009` (sentiment relative to entity, tag bounding, evidence IDs, context exclusions)
|
||||
- [x] T061 [P] [US4] Implement Promptfoo evaluation suite for `article_sentiment_tags` prompt in `evals/promptfoo.config.yaml` verifying minimal ECP identity context and prohibiting semantic `contains` assertions
|
||||
- [x] T062 [P] [US4] Implement contract tests for both versioned prompts in `tests/contract/test_prompts_contract.py` verifying 6-block sequence, semver parsing without regex, SHA-256 calculation, and Promptfoo parity
|
||||
|
||||
### Implementation for User Story 4
|
||||
|
||||
- [x] T063 [P] [US4] Author normative versioned prompt in `prompts/article_sentiment_tags.v1.txt` following the exact 6-block ordering (Doc 07 §5.3)
|
||||
- [x] T064 [US4] Implement enrichment harness in `src/enrichment/harness.py` validating sentiment enum, 3–8 unique tags via NLP/Unicode, and evidence candidate IDs, strictly limiting ECP context to `qid` and `canonical_name` (no raw ECP snapshot, no keyword lists)
|
||||
- [x] T065 [US4] Connect enrichment stage to state machine `ecp_approved → enriched` in `src/core/state_machine.py`
|
||||
- [x] T066 [US4] Implement fallback routing and terminal `ENRICHMENT_FAILED` handling (blocking Markdown generation on failure) in `src/enrichment/harness.py`
|
||||
- [x] T067 [US4] Integrate enrichment stage execution and error handling into `src/cli/consolidate.py`
|
||||
|
||||
**Checkpoint**: User Story 4 delivers structured sentiment and native tags metadata for inherent articles.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: User Story 5 - Canonical Markdown Rendering & Atomic Persistence (Priority: P2)
|
||||
|
||||
**Goal**: Render canonical Markdown with YAML front matter, generate machine-readable `.result.json` manifests, execute atomic filesystem writes (temp + rename), and maintain strict SQLite state consistency.
|
||||
|
||||
**Independent Test**: Verify generated `.md` and `.result.json` files, validating YAML front matter structure, 64-character SHA-256 content hashes, atomic rename lifecycle, and hash-based reconciliation of interrupted writes.
|
||||
|
||||
### Tests for User Story 5
|
||||
|
||||
- [x] T068 [P] [US5] Implement unit tests for canonical YAML front matter and Markdown body renderer in `tests/unit/test_markdown_renderer.py` covering scenarios `OUT-002` to `OUT-007`, `OUT-009`, and `OUT-010` (grounding, formatting, front matter structure)
|
||||
- [x] T069 [P] [US5] Implement unit tests for atomic file writes, permissions, and 64-character hash verification in `tests/unit/test_file_store.py` covering scenarios `OUT-001`, `OUT-011`, and `OUT-012`
|
||||
- [x] T070 [P] [US5] Implement unit tests for SQLite WAL state persistence and crash reconciliation in `tests/unit/test_sqlite_store.py` covering crash recovery and multi-terminal state consistency
|
||||
|
||||
### Implementation for User Story 5
|
||||
|
||||
- [x] T071 [US5] Implement canonical YAML front matter and Markdown body renderer in `src/storage/markdown_renderer.py` (H1 title, italic subtitle when present with no extra blank line when absent, canonical body order, grounded links/images, omitting author/date/sentiment/tags/ECP from body)
|
||||
- [x] T072 [US5] Integrate the atomic filesystem writer from `src/storage/file_store.py` with SQLite completion state in `src/storage/sqlite_store.py` within the same logical completion unit (without distributed transactions), resolving crash divergence via hash-based reconciliation for ID-009, ensuring no terminal state is exposed as completed while files and hashes disagree, and preserving the last safe state without exposing partial final artifact pairs on `PERSISTENCE_FAILED` (covering `completed_text`, `rejected_ecp`, `failed_validation`, and `failed_processing`)
|
||||
- [x] T073 [US5] Connect final persistence to state machine `enriched → completed_text` in `src/core/state_machine.py`
|
||||
- [x] T074 [US5] Integrate final persistence and reconciliation into `src/cli/consolidate.py` and `src/cli/reconcile.py`
|
||||
|
||||
**Checkpoint**: End-to-end pipeline produces atomic published Markdown and manifests with SQLite consistency.
|
||||
|
||||
---
|
||||
|
||||
## Phase 9: User Story 7 - Direct Langfuse Observability, Log Sanitization & Telemetry Queue (Priority: P3)
|
||||
|
||||
**Goal**: Transmit traces, spans, generations, and metrics directly to Langfuse, enforce secret redaction, degrade gracefully to SQLite `pending_telemetry` on network outage, and provide operational telemetry flush.
|
||||
|
||||
**Independent Test**: Process articles with Langfuse available and blocked, checking trace structure (8 stable spans), secret redaction in stderr logs, SQLite queue insertion on outage, and flush execution via `telemetry_flush` CLI.
|
||||
|
||||
### Tests for User Story 7
|
||||
|
||||
- [x] T075 [P] [US7] Implement unit tests for Langfuse tracer, secret redaction, and offline queue in `tests/unit/test_langfuse_tracer.py` covering scenarios `OBS-001` to `OBS-011` (8 stable spans, generation attributes, metric dimensions, cardinality guards)
|
||||
- [x] T076 [P] [US7] Implement integration tests for telemetry degradation, deduplication, and atomic replay in `tests/integration/test_telemetry_degradation.py`
|
||||
- [x] T077 [P] [US7] Implement specialized security tests for authorization-header and secret redaction in logs and SDK exceptions (SEC-006) in `tests/security/test_secret_redaction.py`
|
||||
|
||||
### Implementation for User Story 7
|
||||
|
||||
- [x] T078 [US7] Implement Langfuse observability integration in `src/observability/langfuse_tracer.py` managing traces with 8 stable spans (`validation`, `candidate_preparation`, `hygiene`, `grounding_validation`, `ecp_gate`, `enrichment`, `rendering`, `persistence`) and generations per LLM attempt; the validation span MUST be buffered or materialized only after successful local validation (no remote Langfuse traffic may occur while terminating local validations are running)
|
||||
- [x] T079 [US7] Implement local SQLite queue insertion for telemetry events during Langfuse network outages and atomic flush procedure in `src/observability/langfuse_tracer.py`
|
||||
- [x] T080 [US7] Implement runtime-observable metric emission in `src/observability/langfuse_tracer.py` using exactly each metric and its dimensions from Doc 05 / FR-067, enforcing cardinality restrictions and emitting `prompt_review_signal_total` without self-healing or a parallel metrics store (release-wide aggregation of the 11 critical invariants remains the responsibility of T090)
|
||||
- [x] T081 [US7] Implement operational telemetry flush command in `src/cli/telemetry_flush.py`, ensuring events are marked flushed only upon confirmed delivery and `telemetry_pending_total` returns to zero
|
||||
- [x] T082 [US7] Configure the 3 mandatory Langfuse dashboards (Runtime Health, Quality, Future Review Signals) and preserve reproducible setup evidence without creating a parallel metrics system
|
||||
- [x] T083 [US7] Integrate observability lifecycle and secret redaction into `src/cli/consolidate.py`
|
||||
|
||||
**Checkpoint**: Observability is complete, compliant with the metric catalog, and resilient against outages.
|
||||
|
||||
---
|
||||
|
||||
## Phase 10: User Story 8 - Quality Gates, Zero Regex Verification & Promptfoo Evaluation (Priority: P3)
|
||||
|
||||
**Goal**: Execute comprehensive automated quality gates, Promptfoo offline evaluations, multi-extractor golden regression across 20 reference cases, and zero-regex verification.
|
||||
|
||||
**Independent Test**: Run `pytest tests/quality/`, `pytest evals/`, and `tests/scripts/check_zero_regex.py`, asserting that all 11 critical quality assertions pass without manual intervention.
|
||||
|
||||
### Tests for User Story 8
|
||||
|
||||
- [x] T084 [P] [US8] Implement automated test in `tests/quality/test_zero_regex_enforcement.py` executing `tests/scripts/check_zero_regex.py` across codebase, schemas, and Promptfoo YAML
|
||||
- [x] T085 [P] [US8] Implement automated test in `tests/quality/test_no_powerful_models.py` verifying no runtime module, config, or internal ECP classifier references powerful models
|
||||
- [x] T086 [P] [US8] Implement contract parity test in `tests/contract/test_contract_parity.py` checking all schema versions match 1.0.0
|
||||
- [x] T087 [P] [US8] Implement multi-extractor golden-set quality tests across the 20 reference cases in `tests/quality/test_golden_reference_20.py`
|
||||
- [x] T088 [P] [US8] Implement adversarial prompt injection evaluation in `tests/quality/test_prompt_injection_guard.py`
|
||||
- [x] T089 [P] [US8] Implement automated cost budget verification in `tests/quality/test_cost_budget.py` asserting median per-article cost <= $0.0006
|
||||
|
||||
### Implementation for User Story 8
|
||||
|
||||
- [x] T091 [US8] Implement CI multi-tier trigger runner in `scripts/ci_check.py` distinguishing PR gates, prompt/schema/model change gates, and pre-promotion gates per Doc 04 (including reproducible lockfile package build and metadata verification)
|
||||
- [x] T092 [US8] Implement a minimal holdout and slice evaluation aggregator in `evals/eval_runner.py` aggregating Promptfoo output without duplicating prompt execution, computing slice pass rates (≥95%), block precision/recall/F1, metadata accuracy, correct vs unauthorized repairs, material loss, residual noise, link/image precision, ECP accuracy, sentiment accuracy, tag acceptance, and schema validity
|
||||
- [x] T093 [US8] Implement empirical latency and cost calibration recorder for staging gates in `tests/load/test_load_100_art_per_hour.py`
|
||||
- [x] T094 [US8] Implement the release packaging script in `scripts/build_release_metadata.py` generating `src/core/release-metadata.json` with `release_version`, `runtime_config_sha256`, real prompt hashes, schema versions, certified provider/model mappings for both logical roles, and the certified ECP classifier configuration hash bundled inside the distributable package
|
||||
|
||||
**Checkpoint**: Quality gates, security test suites, and CI evaluation infrastructure are verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 11: User Story 9 - Production Runbook Operations & Lifecycle Management (Priority: P3)
|
||||
|
||||
**Goal**: Implement operational commands (`preflight`, `smoke`, `reconcile`, `telemetry_flush`), SQLite native backup/restore, signal handling (`SIGTERM`/`SIGINT`), rollback, and credential/model rotations.
|
||||
|
||||
**Independent Test**: Run preflight checks against `src/core/release-metadata.json`, execute smoke tests with fixtures, perform native SQLite backup and restore, simulate `SIGTERM` graceful shutdown, and verify credential and model rotation procedures.
|
||||
|
||||
### Tests for User Story 9
|
||||
|
||||
- [x] T095 [P] [US9] Implement integration tests for operational resilience in `tests/integration/test_operations_resilience.py` covering native backup/restore, graceful shutdown signals, rollback, credential rotation, certified model rotation, and uncertified rotation rejection
|
||||
- [x] T096 [P] [US9] Implement unit tests for preflight verification against release metadata in `tests/unit/test_preflight_certification.py`
|
||||
- [x] T097 [P] [US9] Implement unit tests for smoke test execution in `tests/unit/test_smoke_cli.py`
|
||||
|
||||
### Implementation for User Story 9
|
||||
|
||||
- [x] T098 [US9] Implement preflight validation CLI in `src/cli/preflight.py` checking clock sync, release metadata hashes, schemas, SQLite access, filesystem permissions / atomic rename, minimum disk space, valid/active credentials, certified cheap models (including ECP), trace content policy, and ECP classifier/schema access (remote Langfuse outage does not block preflight)
|
||||
- [x] T099 [US9] Implement smoke test CLI in `src/cli/smoke.py` processing fixture and verifying end-to-end pipeline health (fingerprint, state transitions, LLM call, ECP gate, manifest, Markdown, Langfuse trace, cost/latency within approved staging baseline, and idempotent re-execution)
|
||||
- [x] T100 [US9] Implement state and artifact reconciliation CLI in `src/cli/reconcile.py` (completed states vs files/hashes, final files without state, orphan temp files, pending telemetry, duplicate fingerprints, reconciliation report, safe cleanup without altering editorial content)
|
||||
- [x] T101 [US9] Implement full graceful shutdown signal handling (`SIGTERM`, `SIGINT`) in `src/cli/consolidate.py` (stop accepting new units, complete or safely preserve active unit state, close transactions, flush files, attempt telemetry flush, preserve unsent events in SQLite, and exit with coherent exit code)
|
||||
- [x] T102 [US9] Validate and update operational procedures and structure in the normative runtime runbook (`docs/structured_extraction/06_Runbook_Producao_Runtime.md`) for the 11 deployment steps, preflight, smoke, backup/restore, retention, reconciliation, rollback, and credential/model rotations (ready for staging limit incorporation)
|
||||
|
||||
**Checkpoint**: All operational procedures and lifecycle commands are testable and functional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 12: Polish, Verification Gates & Release Sign-Off
|
||||
|
||||
**Purpose**: Final end-to-end execution, evaluation runs, release evidence report generation, and formal sign-off.
|
||||
|
||||
- [x] T103 [P] Execute multi-parser static policy verification across text-processing runtime modules, content tests/assertions, JSON schemas, and Promptfoo YAML configurations of this feature via `python -m tests.scripts.check_zero_regex`
|
||||
- [x] T104 [P] Execute all 9 contract test suites; the Article Input contract MUST validate all 20 real reference units via `pytest tests/contract -v`
|
||||
- [x] T105 Execute the complete unit and mock integration suites, including idempotency, concurrent replay, and full CLI contract validation (`completed_text`, `rejected_ecp`, `failed_validation`, `failed_processing`, idempotent result, config/preflight errors with stdout/stderr and exit codes) via `pytest tests/unit tests/integration -v`
|
||||
- [x] T106 Execute all 8 security scenario tests (`SEC-001` to `SEC-008`) via `pytest tests/security -v`
|
||||
- [x] T107 Execute all 10 fault injection scenario tests (`FLT-001` to `FLT-010`) via `pytest tests/fault_injection -v`
|
||||
- [x] T108 Run end-to-end quickstart validation scenarios A, B, C, D per `quickstart.md`
|
||||
- [x] T109 Execute Promptfoo over the 20-case regression set, production golden set, and protected holdout using the exact production prompts and schemas, preserving per-case and per-slice results
|
||||
- [x] T110 Execute the production-equivalent 100 articles/hour staging run using the installed lockfile-built package, generate the complete normative staging and release-evidence report, verify all 11 zero-tolerance invariants, verify formal absence of forbidden architectural patterns across this feature's runtime codebase, dependencies/lockfile, prompts, schemas, functional configs, and packaging artifacts (verifying no LangChain, LangGraph, agents, API, internal batch/worker pool, Postgres, external queue, object storage, keyword dictionaries, self-healing, powerful models in runtime, online Promptfoo), and obtain approval for cost/latency/storage/fallback limits
|
||||
- [x] T111 Incorporate the approved staging limits into the normative runbook (`docs/structured_extraction/06_Runbook_Producao_Runtime.md`), update README/CLI documentation, and record formal release sign-off
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — starts immediately.
|
||||
- **Foundational (Phase 2)**: Depends on Setup completion — **BLOCKS all user stories**.
|
||||
- **Model Gateway (Phase 3, US6)**: Depends on Foundational — **BLOCKS Hygiene & Enrichment**.
|
||||
- **User Story 1 (Phase 4, P1)**: Depends on Foundational — delivers initial contract validation (Article + ECP), candidate extraction & idempotency.
|
||||
- **User Story 2 (Phase 5, P1)**: Depends on US1 and US6 (Model Gateway) — delivers 10-step LLM extractive hygiene & repairs.
|
||||
- **User Story 3 (Phase 6, P1)**: Depends on US2 — delivers mandatory ECP inherence gate and zero-Markdown rejection.
|
||||
- **User Story 4 (Phase 7, P2)**: Depends on US3 and US6 (Model Gateway) — delivers post-ECP sentiment & tag enrichment and prompts contract testing.
|
||||
- **User Story 5 (Phase 8, P2)**: Depends on US4 — delivers canonical Markdown & atomic persistence across all terminal outcomes.
|
||||
- **User Story 7 (Phase 9, P3)**: Depends on US5 & US6 — delivers Langfuse observability & telemetry queue.
|
||||
- **User Story 8 (Phase 10, P3)**: Depends on US6 and US7 — delivers quality gates, CI multi-tier evals, golden set & load benchmark.
|
||||
- **User Story 9 (Phase 11, P3)**: Depends on US5, US7, and US8 — delivers operational runbooks & resilience commands.
|
||||
- **Polish & Gates (Phase 12)**: Depends on all user stories being complete (T110 executes staging and calibrates limits; T111 incorporates limits into documentation and signs off release).
|
||||
|
||||
---
|
||||
|
||||
## Parallel Execution Opportunities
|
||||
|
||||
```bash
|
||||
# Launch Foundational independent tasks in parallel:
|
||||
Task: "T007 Implement input byte size limiter in src/core/limits.py"
|
||||
Task: "T008 Implement deterministic fingerprint calculator in src/core/fingerprint.py"
|
||||
Task: "T009 Implement CandidateObject models in src/candidate/models.py"
|
||||
Task: "T012 Implement structured JSON logging in src/observability/structured_logger.py"
|
||||
Task: "T013 Implement atomic writer and manifest generator in src/storage/file_store.py"
|
||||
|
||||
# Launch User Story 1 test tasks in parallel:
|
||||
Task: "T023 Contract test for Article Input in tests/contract/test_article_input_contract.py"
|
||||
Task: "T024 Contract test for ECP Snapshot in tests/contract/test_ecp_snapshot_contract.py"
|
||||
Task: "T025 Contract test for Candidates Payload in tests/contract/test_candidates_payload_contract.py"
|
||||
Task: "T026 Unit tests for input limits in tests/unit/test_input_limits.py"
|
||||
Task: "T027 Unit tests for fingerprint in tests/unit/test_fingerprint.py"
|
||||
Task: "T028 Unit tests for candidate extraction in tests/unit/test_candidate_parser.py"
|
||||
Task: "T029 Unit tests for equivalence mapping in tests/unit/test_equivalence_mapping.py"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### Foundation & Incremental Pipeline Flow
|
||||
1. Complete Phase 1: Setup
|
||||
2. Complete Phase 2: Foundational (blocking prerequisites & shared stores)
|
||||
3. Complete Phase 3: User Story 6 (Model Gateway infrastructure & cheap model enforcement)
|
||||
4. Complete Phase 4: User Story 1 (Ingestion, Article/ECP contract validation, candidates, idempotency)
|
||||
5. Complete Phase 5: User Story 2 (Extractive hygiene & micro-repairs)
|
||||
6. Complete Phase 6: User Story 3 (Mandatory ECP gate & relevance enforcement)
|
||||
7. Complete Phase 7: User Story 4 (Post-ECP sentiment & native tags enrichment, prompt contracts)
|
||||
8. Complete Phase 8: User Story 5 (Canonical Markdown & atomic persistence across all terminal outcomes)
|
||||
9. Complete Phase 9: User Story 7 (Direct Langfuse observability & telemetry queue)
|
||||
10. Complete Phase 10: User Story 8 (Automated quality evaluation, golden set & CI gates)
|
||||
11. Complete Phase 11: User Story 9 (Production runbook operations & lifecycle management)
|
||||
12. Complete Phase 12: Polish, Verification Gates, Staging Calibration & Release Sign-Off
|
||||
Reference in New Issue
Block a user