12 KiB
Data Model: Article Consolidation and Hygiene Runtime
Feature Branch: 006-article-consolidation-runtime
Date: 2026-08-23
Status: Complete
1. Domain Entities & Relationships
erDiagram
ArticleInputUnit ||--o{ CandidateObject : extracts
ArticleInputUnit ||--|| ECPSnapshot : references
ArticleInputUnit ||--|| StateMachineRecord : tracks
CandidateObject ||--o{ TextRepairOperation : receives
StateMachineRecord ||--o| OutputManifest : persists
StateMachineRecord ||--o| PublishedMarkdown : renders
StateMachineRecord ||--o{ StateTransitionLog : logs
StateMachineRecord ||--o{ TelemetryEvent : queues
2. Entity Definitions
2.1 Article Input Unit (article_input)
Represents the incoming single article JSON payload. Rejects explicit batch wrappers via "articles": false. Unknown fields in the input are preserved in the original object without alteration.
| Field | Type | Description | Required |
|---|---|---|---|
selected_extractor |
enum | trafilatura | newspaper4k | readability |
Yes |
crawled_url |
string (URL) | null | Crawled URL | No |
error_message |
string | null | Upstream error message | No |
extraction_status |
string | null | Upstream status | No |
http_status |
integer | null | HTTP response status code | No |
input_meta |
object | null | Metadata map (url, titulo, subtitulo, quando_publicado) |
No |
page_title |
string | null | Raw HTML page title | No |
trafilatura |
object | null | Trafilatura extraction output (accepts raw_json as object, string, or null) |
No |
newspaper4k |
object | null | Newspaper4k extraction output | No |
readability |
object | null | Readability extraction output | No |
articles |
false | Explicitly forbidden (batch wrapper rejection) | No |
Note on Validation: The runtime validates that the collective extraction sources provide at least one resolvable source URL, at least one non-empty candidate title, processable text, and usable content in selected_extractor. Total byte size is checked against limits.max_input_bytes before invoking remote providers (failing with INVALID_ARTICLE_SCHEMA if exceeded).
2.2 Entity Context Profile Snapshot (ecp_snapshot)
Represents the complete, integral ECP Snapshot received by the runtime. The runtime validates the payload locally against the monorepo's canonical ECP schema (src/adapters/ecp/schemas/ecp-profile.schema.json) registered in referencing.Registry matching $ref: "https://schemas.aftech.internal/ecp/v1/ecp-profile.schema.json" without HTTP lookups. It extracts identity and version metadata (qid, canonical_name, version) for manifest and trace recording.
| Field | Type | Description | Required |
|---|---|---|---|
| (opaque payload) | object | Complete canonical ECP snapshot validated via referencing.Registry |
Yes |
qid |
string | Extracted canonical Wikidata / Entity QID (e.g. Q148) |
Extracted |
canonical_name |
string | Extracted entity canonical name | Extracted |
version |
string | Extracted semantic version of the referenced ECP profile | Extracted |
2.3 Candidate Object (candidate_object)
Single structural element extracted from the article payloads. Preserves all candidates without destructive deduplication.
| Field | Type | Description | Required |
|---|---|---|---|
candidate_id |
string | Opaque unique ID (e.g. cand_blk_001, cand_title_001) |
Yes |
type |
enum | title | subtitle | author | date | paragraph | heading | list_item | quote | link | image |
Yes |
extractor_source |
enum | trafilatura | newspaper4k | readability | input_meta | page_title |
Yes |
source_field |
string | Origin field (e.g. text, article_html, title) |
Yes |
original_text_or_url |
string | Exact original text or URL content | Yes |
structural_representation |
string | Markdown/HTML/AST structural snippet | Yes |
order_index |
integer | Position index in extractor backbone | Yes |
parent_candidate_id |
string | null | ID of parent element (for nested list items, blockquotes, etc.) | No |
equivalences |
array[string] | List of candidate IDs representing equivalent content from other extractors | Yes |
content_hash |
string (64-char hex) | Deterministic content hash | Yes |
structural_flags |
object | Purely structural flags (e.g. {"heading_level": 2}) |
Yes |
Note on Projection: The LLM prompt receives CandidatesPayload, which is a clean, normalized projection of these internal CandidateObject instances.
2.4 Text Repair Operation (text_repair)
Micro-repair proposed by the LLM and validated by the harness.
| Field | Type | Description | Required |
|---|---|---|---|
target_candidate_id |
string | ID of the target block or metadata candidate | Yes |
original_fragment |
string | Exact substring in candidate to replace | Yes |
replacement_fragment |
string | Validated replacement text | Yes |
category |
enum | encoding | unicode | spacing | punctuation_corruption | obvious_typo |
Yes |
rationale |
string | Short explanation | Yes |
is_accepted |
boolean | Validation outcome from harness | Yes |
rejection_reason |
string | null | Code if rejected (SENSITIVE_ENTITY, AMBIGUOUS_TARGET, OUT_OF_CATEGORY, etc.) |
No |
2.5 State Machine Record (state_record)
SQLite table article_states storing runtime execution status.
| Column | SQLite Type | Description |
|---|---|---|
fingerprint |
TEXT (PK, 64-char hex) | Deterministic content hash of the execution |
source_url |
TEXT | Resolved source URL |
selected_extractor |
TEXT | Extractor used as backbone |
current_state |
TEXT | received | validated | content_cleaned | ecp_approved | ecp_rejected | enriched | completed_text | failed |
final_status |
TEXT | completed_text | rejected_ecp | failed_validation | failed_processing | NULL |
generate_markdown |
INTEGER | 1 if Markdown generated, 0 otherwise |
markdown_path |
TEXT | Path to generated .md file (or NULL) |
manifest_path |
TEXT | Path to generated .result.json file |
markdown_hash |
TEXT (64-char hex) | Content hash of generated .md file (or NULL) |
manifest_hash |
TEXT (64-char hex) | Content hash of generated .result.json file |
ecp_category |
TEXT | DIRECT_INHERENT | CONTEXTUAL_INHERENT | TANGENTIAL | NOT_RELATED | NULL |
ecp_confidence |
REAL | Confidence score (0.0 to 1.0) |
functional_versions_json |
TEXT (JSON) | Consolidated versions of contracts, ECP reference, config, prompts, and models |
trace_id |
TEXT | Langfuse trace identifier |
terminal_error_code |
TEXT | Normative error code if failed |
error_metadata_json |
TEXT (JSON) | Sanitized minimal error metadata (stack trace in technical log only) |
created_at |
TEXT (ISO 8601) | Timestamp of ingestion |
updated_at |
TEXT (ISO 8601) | Timestamp of last transition |
2.6 State Transition Log (state_transitions)
SQLite table state_transitions tracking execution lifecycle.
| Column | SQLite Type | Description |
|---|---|---|
id |
INTEGER (PK AUTO) | Unique transition ID |
fingerprint |
TEXT (FK) | Reference to article_states.fingerprint |
from_state |
TEXT | Starting state |
to_state |
TEXT | Destination state |
start_time |
TEXT (ISO 8601) | Transition start timestamp |
end_time |
TEXT (ISO 8601) | Transition end timestamp |
duration_ms |
REAL | Elapsed milliseconds |
result |
TEXT | success | failure | skipped |
metadata_json |
TEXT (JSON) | Transition context metadata |
2.7 Telemetry Event (pending_telemetry)
SQLite table pending_telemetry for resilient deferred delivery to Langfuse when the network/service is unreachable.
| Column | SQLite Type | Description |
|---|---|---|
event_id |
TEXT (PK) | UUID / Unique event ID |
fingerprint |
TEXT | Associated article fingerprint |
trace_id |
TEXT | Associated trace ID |
event_type |
TEXT | span | generation | score |
payload_json |
TEXT (JSON) | Sanitized telemetry event payload |
created_at |
TEXT (ISO 8601) | Creation timestamp |
retry_count |
INTEGER | Number of transmission attempts |
last_error |
TEXT | Last error message |
2.8 Release Metadata Contract (src/core/release-metadata.json)
Immutable packaged artifact recording certified configurations for preflight verification. runtime_config_sha256 is strictly calculated as the exact file byte SHA-256 hash (hashlib.sha256(Path(config_path).read_bytes()).hexdigest()).
{
"release_version": "1.0.0",
"runtime_config_sha256": "06a2769f7aa3a15a1e61880f171ecebc0d29094ab2499616243e59c0aecf340f",
"prompts_hashes": {
"article_content_hygiene": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0",
"article_sentiment_tags": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0"
},
"schemas_versions": {
"article_input": "1.0.0",
"ecp_snapshot": "1.0.0",
"runtime_config": "1.0.0",
"candidates_payload": "1.0.0",
"hygiene_response": "1.0.0",
"repair_operations": "1.0.0",
"enrichment_response": "1.0.0",
"manifest_output": "1.0.0"
},
"certified_models": {
"runtime_primary": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
"runtime_fallback": { "provider": "deepseek", "model": "deepseek-v4-flash", "role_config_version": "1.0.0" }
},
"ecp_classifier_config_hash": "a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890"
}
2.9 Output Manifest (<fingerprint>.result.json)
Structure of the machine-readable output manifest complying with manifest-output.schema.json.
{
"schema_version": "1.0.0",
"fingerprint": "a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890",
"source_url": "https://example.com/noticia-123",
"selected_extractor": "trafilatura",
"final_status": "completed_text",
"generate_markdown": true,
"markdown_path": "out/a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890a1b2c3d4e5f67890.md",
"markdown_hash": "06a2769f7aa3a15a1e61880f171ecebc0d29094ab2499616243e59c0aecf340f",
"ecp_classification": {
"category": "DIRECT_INHERENT",
"confidence": 0.95,
"rationale": "Article directly analyzes the economic policy of the entity.",
"evidences": ["trecho textual fundamentado 1", "trecho textual fundamentado 2"]
},
"enrichment": {
"sentiment": "positive",
"tags": ["Economia", "Política Monetária", "Inflação"]
},
"provider_versions": {
"hygiene": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
"enrichment": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" }
},
"model_versions": {
"runtime_primary": { "provider": "groq", "model": "openai/gpt-oss-20b", "role_config_version": "1.0.0" },
"runtime_fallback": { "provider": "deepseek", "model": "deepseek-v4-flash", "role_config_version": "1.0.0" }
},
"prompt_versions": {
"article_content_hygiene": { "version": "1.0.0", "hash": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0" },
"article_sentiment_tags": { "version": "1.0.0", "hash": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0" }
},
"config_version": "1.0.0",
"trace_id": "trace_run_20260823_001",
"error_codes": []
}
2.10 Published Markdown (<fingerprint>.md)
Format of the rendered Markdown document with YAML front matter.
---
title: "Título Principal do Artigo Publicado"
subtitle: "Subtítulo editorial detalhado"
author: "Nome do Autor"
published_at: "2026-08-23T14:00:00Z"
source_url: "https://example.com/noticia-123"
sentiment: positive
tags:
- Economia
- Política Monetária
- Inflação
ecp_qid: "Q148"
ecp_canonical_name: "Entidade Alvo"
ecp_category: DIRECT_INHERENT
ecp_confidence: 0.95
---
# Título Principal do Artigo Publicado
*Subtítulo editorial detalhado*
Primeiro parágrafo do artigo com [link grounded](https://example.com/referencia) e texto limpo.
## Intertítulo Editorial
Segundo parágrafo contendo citação textual sem alterações indevidas.

Parágrafo de encerramento sem notas de rodapé publicitárias ou chamadas de redes sociais.