Files

5.7 KiB

Data Model: Convert Article JSON to Markdown

Feature: 005-convert-json-markdown | Date: 2026-08-21

1. Domain Entities & Schemas

Entity 1: ArticleInput (Source JSON)

Represents the raw parsed JSON structure of a single news article.

ArticleInput
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
├── crawled_url: string [OPTIONAL: URL string]
├── page_title: string [OPTIONAL: Page title string]
├── input_meta: object [OPTIONAL]
│   ├── url: string [OPTIONAL]
│   ├── titulo: string [OPTIONAL]
│   ├── subtitulo: string [OPTIONAL]
│   └── quando_publicado: string [OPTIONAL]
├── trafilatura: ExtractorBlock [CONDITIONAL]
├── newspaper4k: ExtractorBlock [CONDITIONAL]
└── readability: ExtractorBlock [CONDITIONAL]

Validation Rules:

  • Root must be a JSON object (dict).
  • Root must NOT contain an articles key (batch JSON is rejected).
  • selected_extractor must be exactly one of "trafilatura", "newspaper4k", "readability".
  • The object corresponding to selected_extractor must exist in ArticleInput and contain usable body content.

Entity 2: ExtractorBlock (Per-Extractor Data)

Represents the extraction results produced by each individual extractor library.

Extractor Primary Body Field Fallback Body Field Metadata Fields Available
trafilatura markdown (string) text (string) title, description, author, date, sitename, hostname, categories, tags, language, image, canonical_url
newspaper4k article_html (string) text (string) title, meta_description, authors, publish_date, meta_site_name, tags, keywords, meta_keywords, meta_lang, top_image, canonical_link
readability cleaned_html (string) cleaned_text (string) title, author

Entity 3: ResolvedArticleMetadata

The normalized, validated, and prioritized metadata extracted from candidate sources.

Attribute Type Mandatory? Normalization / Validation Rule
title str Yes Unescaped, trimmed, single spaces. Rejection if empty.
original_url str Yes Valid absolute URL with http:// or https:// and valid hostname.
subtitle Optional[str] No Omitted if empty, placeholder, or equal to title (case-insensitive).
authors List[str] No Deduplicated, case-preserved, no URL entries. Omitted if empty.
publish_date Optional[str] No ISO 8601 string (with timezone) or YYYY-MM-DD. Omitted if unparseable.
site_name Optional[str] No Normalized site string or fallback to original URL hostname.
categories List[str] No Deduplicated, non-empty category strings.
tags List[str] No Deduplicated, non-empty tag strings.
keywords List[str] No Deduplicated, non-empty keyword strings.
language Optional[str] No Language code string (e.g. es, pt, en).
top_image Optional[str] No Valid absolute URL with http:// or https://.

Entity 4: MarkdownDocument

The structured representation of the output Markdown file.

MarkdownDocument
├── title_h1: "# " + ResolvedArticleMetadata.title
├── subtitle_block: Optional paragraph
├── metadata_block: Key-value list of bold labels and values
│   ├── **Autor:** {authors joined by ", "}
│   ├── **Publicado em:** {publish_date}
│   ├── **Site:** {site_name}
│   ├── **Categoria:** {categories joined by ", "}
│   ├── **Tags:** {tags joined by ", "}
│   ├── **Palavras-chave:** {keywords joined by ", "}
│   ├── **Idioma:** {language}
│   └── **Fonte original:** [{original_url}]({original_url})
├── top_image_block: Optional "![Imagem principal]({top_image})"
├── separator: "---"
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)

2. Priority Resolution Matrix

Title:
  1. SELECIONADO.title
  2. input_meta.titulo
  3. page_title
  4. newspaper4k.title
  5. trafilatura.title
  6. readability.title

Original URL:
  1. input_meta.url
  2. crawled_url
  3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
  4. trafilatura.canonical_url
  5. newspaper4k.canonical_link

Subtitle / Description:
  1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
  2. trafilatura.description
  3. newspaper4k.meta_description
  4. input_meta.subtitulo

Authors:
  1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
  2. newspaper4k.authors
  3. trafilatura.author
  4. readability.author

Publication Date:
  1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
  2. newspaper4k.publish_date
  3. trafilatura.date
  4. input_meta.quando_publicado

Site Name:
  1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
  2. trafilatura.sitename
  3. newspaper4k.meta_site_name
  4. trafilatura.hostname
  5. Hostname of resolved Original URL

Categories:
  1. Categories of SELECIONADO (trafilatura.categories)
  2. trafilatura.categories

Tags:
  1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
  2. trafilatura.tags
  3. newspaper4k.tags
  4. newspaper4k.meta_keywords

Keywords:
  1. newspaper4k.keywords
  2. newspaper4k.meta_keywords

Language:
  1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
  2. trafilatura.language
  3. newspaper4k.meta_lang

Top Image:
  1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
  2. newspaper4k.top_image
  3. trafilatura.image