5.7 KiB
5.7 KiB
Data Model: Convert Article JSON to Markdown
Feature: 005-convert-json-markdown | Date: 2026-08-21
1. Domain Entities & Schemas
Entity 1: ArticleInput (Source JSON)
Represents the raw parsed JSON structure of a single news article.
ArticleInput
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
├── crawled_url: string [OPTIONAL: URL string]
├── page_title: string [OPTIONAL: Page title string]
├── input_meta: object [OPTIONAL]
│ ├── url: string [OPTIONAL]
│ ├── titulo: string [OPTIONAL]
│ ├── subtitulo: string [OPTIONAL]
│ └── quando_publicado: string [OPTIONAL]
├── trafilatura: ExtractorBlock [CONDITIONAL]
├── newspaper4k: ExtractorBlock [CONDITIONAL]
└── readability: ExtractorBlock [CONDITIONAL]
Validation Rules:
- Root must be a JSON object (dict).
- Root must NOT contain an
articleskey (batch JSON is rejected). selected_extractormust be exactly one of"trafilatura","newspaper4k","readability".- The object corresponding to
selected_extractormust exist inArticleInputand contain usable body content.
Entity 2: ExtractorBlock (Per-Extractor Data)
Represents the extraction results produced by each individual extractor library.
| Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available |
|---|---|---|---|
trafilatura |
markdown (string) |
text (string) |
title, description, author, date, sitename, hostname, categories, tags, language, image, canonical_url |
newspaper4k |
article_html (string) |
text (string) |
title, meta_description, authors, publish_date, meta_site_name, tags, keywords, meta_keywords, meta_lang, top_image, canonical_link |
readability |
cleaned_html (string) |
cleaned_text (string) |
title, author |
Entity 3: ResolvedArticleMetadata
The normalized, validated, and prioritized metadata extracted from candidate sources.
| Attribute | Type | Mandatory? | Normalization / Validation Rule |
|---|---|---|---|
title |
str |
Yes | Unescaped, trimmed, single spaces. Rejection if empty. |
original_url |
str |
Yes | Valid absolute URL with http:// or https:// and valid hostname. |
subtitle |
Optional[str] |
No | Omitted if empty, placeholder, or equal to title (case-insensitive). |
authors |
List[str] |
No | Deduplicated, case-preserved, no URL entries. Omitted if empty. |
publish_date |
Optional[str] |
No | ISO 8601 string (with timezone) or YYYY-MM-DD. Omitted if unparseable. |
site_name |
Optional[str] |
No | Normalized site string or fallback to original URL hostname. |
categories |
List[str] |
No | Deduplicated, non-empty category strings. |
tags |
List[str] |
No | Deduplicated, non-empty tag strings. |
keywords |
List[str] |
No | Deduplicated, non-empty keyword strings. |
language |
Optional[str] |
No | Language code string (e.g. es, pt, en). |
top_image |
Optional[str] |
No | Valid absolute URL with http:// or https://. |
Entity 4: MarkdownDocument
The structured representation of the output Markdown file.
MarkdownDocument
├── title_h1: "# " + ResolvedArticleMetadata.title
├── subtitle_block: Optional paragraph
├── metadata_block: Key-value list of bold labels and values
│ ├── **Autor:** {authors joined by ", "}
│ ├── **Publicado em:** {publish_date}
│ ├── **Site:** {site_name}
│ ├── **Categoria:** {categories joined by ", "}
│ ├── **Tags:** {tags joined by ", "}
│ ├── **Palavras-chave:** {keywords joined by ", "}
│ ├── **Idioma:** {language}
│ └── **Fonte original:** [{original_url}]({original_url})
├── top_image_block: Optional ""
├── separator: "---"
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)
2. Priority Resolution Matrix
Title:
1. SELECIONADO.title
2. input_meta.titulo
3. page_title
4. newspaper4k.title
5. trafilatura.title
6. readability.title
Original URL:
1. input_meta.url
2. crawled_url
3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
4. trafilatura.canonical_url
5. newspaper4k.canonical_link
Subtitle / Description:
1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
2. trafilatura.description
3. newspaper4k.meta_description
4. input_meta.subtitulo
Authors:
1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
2. newspaper4k.authors
3. trafilatura.author
4. readability.author
Publication Date:
1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
2. newspaper4k.publish_date
3. trafilatura.date
4. input_meta.quando_publicado
Site Name:
1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
2. trafilatura.sitename
3. newspaper4k.meta_site_name
4. trafilatura.hostname
5. Hostname of resolved Original URL
Categories:
1. Categories of SELECIONADO (trafilatura.categories)
2. trafilatura.categories
Tags:
1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
2. trafilatura.tags
3. newspaper4k.tags
4. newspaper4k.meta_keywords
Keywords:
1. newspaper4k.keywords
2. newspaper4k.meta_keywords
Language:
1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
2. trafilatura.language
3. newspaper4k.meta_lang
Top Image:
1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
2. newspaper4k.top_image
3. trafilatura.image