feat(converter): implement deterministic JSON to Markdown article converter (spec 005)
This commit is contained in:
@@ -0,0 +1,156 @@
|
||||
# Data Model: Convert Article JSON to Markdown
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Domain Entities & Schemas
|
||||
|
||||
### Entity 1: `ArticleInput` (Source JSON)
|
||||
|
||||
Represents the raw parsed JSON structure of a single news article.
|
||||
|
||||
```text
|
||||
ArticleInput
|
||||
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
|
||||
├── crawled_url: string [OPTIONAL: URL string]
|
||||
├── page_title: string [OPTIONAL: Page title string]
|
||||
├── input_meta: object [OPTIONAL]
|
||||
│ ├── url: string [OPTIONAL]
|
||||
│ ├── titulo: string [OPTIONAL]
|
||||
│ ├── subtitulo: string [OPTIONAL]
|
||||
│ └── quando_publicado: string [OPTIONAL]
|
||||
├── trafilatura: ExtractorBlock [CONDITIONAL]
|
||||
├── newspaper4k: ExtractorBlock [CONDITIONAL]
|
||||
└── readability: ExtractorBlock [CONDITIONAL]
|
||||
```
|
||||
|
||||
#### Validation Rules:
|
||||
- Root must be a JSON object (dict).
|
||||
- Root must NOT contain an `articles` key (batch JSON is rejected).
|
||||
- `selected_extractor` must be exactly one of `"trafilatura"`, `"newspaper4k"`, `"readability"`.
|
||||
- The object corresponding to `selected_extractor` must exist in `ArticleInput` and contain usable body content.
|
||||
|
||||
---
|
||||
|
||||
### Entity 2: `ExtractorBlock` (Per-Extractor Data)
|
||||
|
||||
Represents the extraction results produced by each individual extractor library.
|
||||
|
||||
| Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available |
|
||||
|---|---|---|---|
|
||||
| `trafilatura` | `markdown` (string) | `text` (string) | `title`, `description`, `author`, `date`, `sitename`, `hostname`, `categories`, `tags`, `language`, `image`, `canonical_url` |
|
||||
| `newspaper4k` | `article_html` (string) | `text` (string) | `title`, `meta_description`, `authors`, `publish_date`, `meta_site_name`, `tags`, `keywords`, `meta_keywords`, `meta_lang`, `top_image`, `canonical_link` |
|
||||
| `readability` | `cleaned_html` (string) | `cleaned_text` (string) | `title`, `author` |
|
||||
|
||||
---
|
||||
|
||||
### Entity 3: `ResolvedArticleMetadata`
|
||||
|
||||
The normalized, validated, and prioritized metadata extracted from candidate sources.
|
||||
|
||||
| Attribute | Type | Mandatory? | Normalization / Validation Rule |
|
||||
|---|---|:---:|---|
|
||||
| `title` | `str` | Yes | Unescaped, trimmed, single spaces. Rejection if empty. |
|
||||
| `original_url` | `str` | Yes | Valid absolute URL with `http://` or `https://` and valid hostname. |
|
||||
| `subtitle` | `Optional[str]` | No | Omitted if empty, placeholder, or equal to `title` (case-insensitive). |
|
||||
| `authors` | `List[str]` | No | Deduplicated, case-preserved, no URL entries. Omitted if empty. |
|
||||
| `publish_date` | `Optional[str]` | No | ISO 8601 string (with timezone) or `YYYY-MM-DD`. Omitted if unparseable. |
|
||||
| `site_name` | `Optional[str]` | No | Normalized site string or fallback to original URL hostname. |
|
||||
| `categories` | `List[str]` | No | Deduplicated, non-empty category strings. |
|
||||
| `tags` | `List[str]` | No | Deduplicated, non-empty tag strings. |
|
||||
| `keywords` | `List[str]` | No | Deduplicated, non-empty keyword strings. |
|
||||
| `language` | `Optional[str]` | No | Language code string (e.g. `es`, `pt`, `en`). |
|
||||
| `top_image` | `Optional[str]` | No | Valid absolute URL with `http://` or `https://`. |
|
||||
|
||||
---
|
||||
|
||||
### Entity 4: `MarkdownDocument`
|
||||
|
||||
The structured representation of the output Markdown file.
|
||||
|
||||
```text
|
||||
MarkdownDocument
|
||||
├── title_h1: "# " + ResolvedArticleMetadata.title
|
||||
├── subtitle_block: Optional paragraph
|
||||
├── metadata_block: Key-value list of bold labels and values
|
||||
│ ├── **Autor:** {authors joined by ", "}
|
||||
│ ├── **Publicado em:** {publish_date}
|
||||
│ ├── **Site:** {site_name}
|
||||
│ ├── **Categoria:** {categories joined by ", "}
|
||||
│ ├── **Tags:** {tags joined by ", "}
|
||||
│ ├── **Palavras-chave:** {keywords joined by ", "}
|
||||
│ ├── **Idioma:** {language}
|
||||
│ └── **Fonte original:** [{original_url}]({original_url})
|
||||
├── top_image_block: Optional ""
|
||||
├── separator: "---"
|
||||
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Priority Resolution Matrix
|
||||
|
||||
```text
|
||||
Title:
|
||||
1. SELECIONADO.title
|
||||
2. input_meta.titulo
|
||||
3. page_title
|
||||
4. newspaper4k.title
|
||||
5. trafilatura.title
|
||||
6. readability.title
|
||||
|
||||
Original URL:
|
||||
1. input_meta.url
|
||||
2. crawled_url
|
||||
3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
|
||||
4. trafilatura.canonical_url
|
||||
5. newspaper4k.canonical_link
|
||||
|
||||
Subtitle / Description:
|
||||
1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
|
||||
2. trafilatura.description
|
||||
3. newspaper4k.meta_description
|
||||
4. input_meta.subtitulo
|
||||
|
||||
Authors:
|
||||
1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
|
||||
2. newspaper4k.authors
|
||||
3. trafilatura.author
|
||||
4. readability.author
|
||||
|
||||
Publication Date:
|
||||
1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
|
||||
2. newspaper4k.publish_date
|
||||
3. trafilatura.date
|
||||
4. input_meta.quando_publicado
|
||||
|
||||
Site Name:
|
||||
1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
|
||||
2. trafilatura.sitename
|
||||
3. newspaper4k.meta_site_name
|
||||
4. trafilatura.hostname
|
||||
5. Hostname of resolved Original URL
|
||||
|
||||
Categories:
|
||||
1. Categories of SELECIONADO (trafilatura.categories)
|
||||
2. trafilatura.categories
|
||||
|
||||
Tags:
|
||||
1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
|
||||
2. trafilatura.tags
|
||||
3. newspaper4k.tags
|
||||
4. newspaper4k.meta_keywords
|
||||
|
||||
Keywords:
|
||||
1. newspaper4k.keywords
|
||||
2. newspaper4k.meta_keywords
|
||||
|
||||
Language:
|
||||
1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
|
||||
2. trafilatura.language
|
||||
3. newspaper4k.meta_lang
|
||||
|
||||
Top Image:
|
||||
1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
|
||||
2. newspaper4k.top_image
|
||||
3. trafilatura.image
|
||||
```
|
||||
Reference in New Issue
Block a user