feat(converter): implement deterministic JSON to Markdown article converter (spec 005)

This commit is contained in:
2026-08-21 10:30:14 -03:00
parent 64dfd842de
commit 926a6b8cfc
58 changed files with 20138 additions and 2301 deletions
@@ -0,0 +1,156 @@
# Data Model: Convert Article JSON to Markdown
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
## 1. Domain Entities & Schemas
### Entity 1: `ArticleInput` (Source JSON)
Represents the raw parsed JSON structure of a single news article.
```text
ArticleInput
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
├── crawled_url: string [OPTIONAL: URL string]
├── page_title: string [OPTIONAL: Page title string]
├── input_meta: object [OPTIONAL]
│ ├── url: string [OPTIONAL]
│ ├── titulo: string [OPTIONAL]
│ ├── subtitulo: string [OPTIONAL]
│ └── quando_publicado: string [OPTIONAL]
├── trafilatura: ExtractorBlock [CONDITIONAL]
├── newspaper4k: ExtractorBlock [CONDITIONAL]
└── readability: ExtractorBlock [CONDITIONAL]
```
#### Validation Rules:
- Root must be a JSON object (dict).
- Root must NOT contain an `articles` key (batch JSON is rejected).
- `selected_extractor` must be exactly one of `"trafilatura"`, `"newspaper4k"`, `"readability"`.
- The object corresponding to `selected_extractor` must exist in `ArticleInput` and contain usable body content.
---
### Entity 2: `ExtractorBlock` (Per-Extractor Data)
Represents the extraction results produced by each individual extractor library.
| Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available |
|---|---|---|---|
| `trafilatura` | `markdown` (string) | `text` (string) | `title`, `description`, `author`, `date`, `sitename`, `hostname`, `categories`, `tags`, `language`, `image`, `canonical_url` |
| `newspaper4k` | `article_html` (string) | `text` (string) | `title`, `meta_description`, `authors`, `publish_date`, `meta_site_name`, `tags`, `keywords`, `meta_keywords`, `meta_lang`, `top_image`, `canonical_link` |
| `readability` | `cleaned_html` (string) | `cleaned_text` (string) | `title`, `author` |
---
### Entity 3: `ResolvedArticleMetadata`
The normalized, validated, and prioritized metadata extracted from candidate sources.
| Attribute | Type | Mandatory? | Normalization / Validation Rule |
|---|---|:---:|---|
| `title` | `str` | Yes | Unescaped, trimmed, single spaces. Rejection if empty. |
| `original_url` | `str` | Yes | Valid absolute URL with `http://` or `https://` and valid hostname. |
| `subtitle` | `Optional[str]` | No | Omitted if empty, placeholder, or equal to `title` (case-insensitive). |
| `authors` | `List[str]` | No | Deduplicated, case-preserved, no URL entries. Omitted if empty. |
| `publish_date` | `Optional[str]` | No | ISO 8601 string (with timezone) or `YYYY-MM-DD`. Omitted if unparseable. |
| `site_name` | `Optional[str]` | No | Normalized site string or fallback to original URL hostname. |
| `categories` | `List[str]` | No | Deduplicated, non-empty category strings. |
| `tags` | `List[str]` | No | Deduplicated, non-empty tag strings. |
| `keywords` | `List[str]` | No | Deduplicated, non-empty keyword strings. |
| `language` | `Optional[str]` | No | Language code string (e.g. `es`, `pt`, `en`). |
| `top_image` | `Optional[str]` | No | Valid absolute URL with `http://` or `https://`. |
---
### Entity 4: `MarkdownDocument`
The structured representation of the output Markdown file.
```text
MarkdownDocument
├── title_h1: "# " + ResolvedArticleMetadata.title
├── subtitle_block: Optional paragraph
├── metadata_block: Key-value list of bold labels and values
│ ├── **Autor:** {authors joined by ", "}
│ ├── **Publicado em:** {publish_date}
│ ├── **Site:** {site_name}
│ ├── **Categoria:** {categories joined by ", "}
│ ├── **Tags:** {tags joined by ", "}
│ ├── **Palavras-chave:** {keywords joined by ", "}
│ ├── **Idioma:** {language}
│ └── **Fonte original:** [{original_url}]({original_url})
├── top_image_block: Optional "![Imagem principal]({top_image})"
├── separator: "---"
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)
```
---
## 2. Priority Resolution Matrix
```text
Title:
1. SELECIONADO.title
2. input_meta.titulo
3. page_title
4. newspaper4k.title
5. trafilatura.title
6. readability.title
Original URL:
1. input_meta.url
2. crawled_url
3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
4. trafilatura.canonical_url
5. newspaper4k.canonical_link
Subtitle / Description:
1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
2. trafilatura.description
3. newspaper4k.meta_description
4. input_meta.subtitulo
Authors:
1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
2. newspaper4k.authors
3. trafilatura.author
4. readability.author
Publication Date:
1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
2. newspaper4k.publish_date
3. trafilatura.date
4. input_meta.quando_publicado
Site Name:
1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
2. trafilatura.sitename
3. newspaper4k.meta_site_name
4. trafilatura.hostname
5. Hostname of resolved Original URL
Categories:
1. Categories of SELECIONADO (trafilatura.categories)
2. trafilatura.categories
Tags:
1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
2. trafilatura.tags
3. newspaper4k.tags
4. newspaper4k.meta_keywords
Keywords:
1. newspaper4k.keywords
2. newspaper4k.meta_keywords
Language:
1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
2. trafilatura.language
3. newspaper4k.meta_lang
Top Image:
1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
2. newspaper4k.top_image
3. trafilatura.image
```