# Data Model: Convert Article JSON to Markdown **Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21 ## 1. Domain Entities & Schemas ### Entity 1: `ArticleInput` (Source JSON) Represents the raw parsed JSON structure of a single news article. ```text ArticleInput ├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"] ├── crawled_url: string [OPTIONAL: URL string] ├── page_title: string [OPTIONAL: Page title string] ├── input_meta: object [OPTIONAL] │ ├── url: string [OPTIONAL] │ ├── titulo: string [OPTIONAL] │ ├── subtitulo: string [OPTIONAL] │ └── quando_publicado: string [OPTIONAL] ├── trafilatura: ExtractorBlock [CONDITIONAL] ├── newspaper4k: ExtractorBlock [CONDITIONAL] └── readability: ExtractorBlock [CONDITIONAL] ``` #### Validation Rules: - Root must be a JSON object (dict). - Root must NOT contain an `articles` key (batch JSON is rejected). - `selected_extractor` must be exactly one of `"trafilatura"`, `"newspaper4k"`, `"readability"`. - The object corresponding to `selected_extractor` must exist in `ArticleInput` and contain usable body content. --- ### Entity 2: `ExtractorBlock` (Per-Extractor Data) Represents the extraction results produced by each individual extractor library. | Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available | |---|---|---|---| | `trafilatura` | `markdown` (string) | `text` (string) | `title`, `description`, `author`, `date`, `sitename`, `hostname`, `categories`, `tags`, `language`, `image`, `canonical_url` | | `newspaper4k` | `article_html` (string) | `text` (string) | `title`, `meta_description`, `authors`, `publish_date`, `meta_site_name`, `tags`, `keywords`, `meta_keywords`, `meta_lang`, `top_image`, `canonical_link` | | `readability` | `cleaned_html` (string) | `cleaned_text` (string) | `title`, `author` | --- ### Entity 3: `ResolvedArticleMetadata` The normalized, validated, and prioritized metadata extracted from candidate sources. | Attribute | Type | Mandatory? | Normalization / Validation Rule | |---|---|:---:|---| | `title` | `str` | Yes | Unescaped, trimmed, single spaces. Rejection if empty. | | `original_url` | `str` | Yes | Valid absolute URL with `http://` or `https://` and valid hostname. | | `subtitle` | `Optional[str]` | No | Omitted if empty, placeholder, or equal to `title` (case-insensitive). | | `authors` | `List[str]` | No | Deduplicated, case-preserved, no URL entries. Omitted if empty. | | `publish_date` | `Optional[str]` | No | ISO 8601 string (with timezone) or `YYYY-MM-DD`. Omitted if unparseable. | | `site_name` | `Optional[str]` | No | Normalized site string or fallback to original URL hostname. | | `categories` | `List[str]` | No | Deduplicated, non-empty category strings. | | `tags` | `List[str]` | No | Deduplicated, non-empty tag strings. | | `keywords` | `List[str]` | No | Deduplicated, non-empty keyword strings. | | `language` | `Optional[str]` | No | Language code string (e.g. `es`, `pt`, `en`). | | `top_image` | `Optional[str]` | No | Valid absolute URL with `http://` or `https://`. | --- ### Entity 4: `MarkdownDocument` The structured representation of the output Markdown file. ```text MarkdownDocument ├── title_h1: "# " + ResolvedArticleMetadata.title ├── subtitle_block: Optional paragraph ├── metadata_block: Key-value list of bold labels and values │ ├── **Autor:** {authors joined by ", "} │ ├── **Publicado em:** {publish_date} │ ├── **Site:** {site_name} │ ├── **Categoria:** {categories joined by ", "} │ ├── **Tags:** {tags joined by ", "} │ ├── **Palavras-chave:** {keywords joined by ", "} │ ├── **Idioma:** {language} │ └── **Fonte original:** [{original_url}]({original_url}) ├── top_image_block: Optional "![Imagem principal]({top_image})" ├── separator: "---" └── body_content: Converted Markdown text (LF line endings, normalized whitespace) ``` --- ## 2. Priority Resolution Matrix ```text Title: 1. SELECIONADO.title 2. input_meta.titulo 3. page_title 4. newspaper4k.title 5. trafilatura.title 6. readability.title Original URL: 1. input_meta.url 2. crawled_url 3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link) 4. trafilatura.canonical_url 5. newspaper4k.canonical_link Subtitle / Description: 1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description) 2. trafilatura.description 3. newspaper4k.meta_description 4. input_meta.subtitulo Authors: 1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author) 2. newspaper4k.authors 3. trafilatura.author 4. readability.author Publication Date: 1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date) 2. newspaper4k.publish_date 3. trafilatura.date 4. input_meta.quando_publicado Site Name: 1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name) 2. trafilatura.sitename 3. newspaper4k.meta_site_name 4. trafilatura.hostname 5. Hostname of resolved Original URL Categories: 1. Categories of SELECIONADO (trafilatura.categories) 2. trafilatura.categories Tags: 1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags) 2. trafilatura.tags 3. newspaper4k.tags 4. newspaper4k.meta_keywords Keywords: 1. newspaper4k.keywords 2. newspaper4k.meta_keywords Language: 1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang) 2. trafilatura.language 3. newspaper4k.meta_lang Top Image: 1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image) 2. newspaper4k.top_image 3. trafilatura.image ```