157 lines
5.7 KiB
Markdown
157 lines
5.7 KiB
Markdown
# Data Model: Convert Article JSON to Markdown
|
|
|
|
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
|
|
|
## 1. Domain Entities & Schemas
|
|
|
|
### Entity 1: `ArticleInput` (Source JSON)
|
|
|
|
Represents the raw parsed JSON structure of a single news article.
|
|
|
|
```text
|
|
ArticleInput
|
|
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
|
|
├── crawled_url: string [OPTIONAL: URL string]
|
|
├── page_title: string [OPTIONAL: Page title string]
|
|
├── input_meta: object [OPTIONAL]
|
|
│ ├── url: string [OPTIONAL]
|
|
│ ├── titulo: string [OPTIONAL]
|
|
│ ├── subtitulo: string [OPTIONAL]
|
|
│ └── quando_publicado: string [OPTIONAL]
|
|
├── trafilatura: ExtractorBlock [CONDITIONAL]
|
|
├── newspaper4k: ExtractorBlock [CONDITIONAL]
|
|
└── readability: ExtractorBlock [CONDITIONAL]
|
|
```
|
|
|
|
#### Validation Rules:
|
|
- Root must be a JSON object (dict).
|
|
- Root must NOT contain an `articles` key (batch JSON is rejected).
|
|
- `selected_extractor` must be exactly one of `"trafilatura"`, `"newspaper4k"`, `"readability"`.
|
|
- The object corresponding to `selected_extractor` must exist in `ArticleInput` and contain usable body content.
|
|
|
|
---
|
|
|
|
### Entity 2: `ExtractorBlock` (Per-Extractor Data)
|
|
|
|
Represents the extraction results produced by each individual extractor library.
|
|
|
|
| Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available |
|
|
|---|---|---|---|
|
|
| `trafilatura` | `markdown` (string) | `text` (string) | `title`, `description`, `author`, `date`, `sitename`, `hostname`, `categories`, `tags`, `language`, `image`, `canonical_url` |
|
|
| `newspaper4k` | `article_html` (string) | `text` (string) | `title`, `meta_description`, `authors`, `publish_date`, `meta_site_name`, `tags`, `keywords`, `meta_keywords`, `meta_lang`, `top_image`, `canonical_link` |
|
|
| `readability` | `cleaned_html` (string) | `cleaned_text` (string) | `title`, `author` |
|
|
|
|
---
|
|
|
|
### Entity 3: `ResolvedArticleMetadata`
|
|
|
|
The normalized, validated, and prioritized metadata extracted from candidate sources.
|
|
|
|
| Attribute | Type | Mandatory? | Normalization / Validation Rule |
|
|
|---|---|:---:|---|
|
|
| `title` | `str` | Yes | Unescaped, trimmed, single spaces. Rejection if empty. |
|
|
| `original_url` | `str` | Yes | Valid absolute URL with `http://` or `https://` and valid hostname. |
|
|
| `subtitle` | `Optional[str]` | No | Omitted if empty, placeholder, or equal to `title` (case-insensitive). |
|
|
| `authors` | `List[str]` | No | Deduplicated, case-preserved, no URL entries. Omitted if empty. |
|
|
| `publish_date` | `Optional[str]` | No | ISO 8601 string (with timezone) or `YYYY-MM-DD`. Omitted if unparseable. |
|
|
| `site_name` | `Optional[str]` | No | Normalized site string or fallback to original URL hostname. |
|
|
| `categories` | `List[str]` | No | Deduplicated, non-empty category strings. |
|
|
| `tags` | `List[str]` | No | Deduplicated, non-empty tag strings. |
|
|
| `keywords` | `List[str]` | No | Deduplicated, non-empty keyword strings. |
|
|
| `language` | `Optional[str]` | No | Language code string (e.g. `es`, `pt`, `en`). |
|
|
| `top_image` | `Optional[str]` | No | Valid absolute URL with `http://` or `https://`. |
|
|
|
|
---
|
|
|
|
### Entity 4: `MarkdownDocument`
|
|
|
|
The structured representation of the output Markdown file.
|
|
|
|
```text
|
|
MarkdownDocument
|
|
├── title_h1: "# " + ResolvedArticleMetadata.title
|
|
├── subtitle_block: Optional paragraph
|
|
├── metadata_block: Key-value list of bold labels and values
|
|
│ ├── **Autor:** {authors joined by ", "}
|
|
│ ├── **Publicado em:** {publish_date}
|
|
│ ├── **Site:** {site_name}
|
|
│ ├── **Categoria:** {categories joined by ", "}
|
|
│ ├── **Tags:** {tags joined by ", "}
|
|
│ ├── **Palavras-chave:** {keywords joined by ", "}
|
|
│ ├── **Idioma:** {language}
|
|
│ └── **Fonte original:** [{original_url}]({original_url})
|
|
├── top_image_block: Optional ""
|
|
├── separator: "---"
|
|
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)
|
|
```
|
|
|
|
---
|
|
|
|
## 2. Priority Resolution Matrix
|
|
|
|
```text
|
|
Title:
|
|
1. SELECIONADO.title
|
|
2. input_meta.titulo
|
|
3. page_title
|
|
4. newspaper4k.title
|
|
5. trafilatura.title
|
|
6. readability.title
|
|
|
|
Original URL:
|
|
1. input_meta.url
|
|
2. crawled_url
|
|
3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
|
|
4. trafilatura.canonical_url
|
|
5. newspaper4k.canonical_link
|
|
|
|
Subtitle / Description:
|
|
1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
|
|
2. trafilatura.description
|
|
3. newspaper4k.meta_description
|
|
4. input_meta.subtitulo
|
|
|
|
Authors:
|
|
1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
|
|
2. newspaper4k.authors
|
|
3. trafilatura.author
|
|
4. readability.author
|
|
|
|
Publication Date:
|
|
1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
|
|
2. newspaper4k.publish_date
|
|
3. trafilatura.date
|
|
4. input_meta.quando_publicado
|
|
|
|
Site Name:
|
|
1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
|
|
2. trafilatura.sitename
|
|
3. newspaper4k.meta_site_name
|
|
4. trafilatura.hostname
|
|
5. Hostname of resolved Original URL
|
|
|
|
Categories:
|
|
1. Categories of SELECIONADO (trafilatura.categories)
|
|
2. trafilatura.categories
|
|
|
|
Tags:
|
|
1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
|
|
2. trafilatura.tags
|
|
3. newspaper4k.tags
|
|
4. newspaper4k.meta_keywords
|
|
|
|
Keywords:
|
|
1. newspaper4k.keywords
|
|
2. newspaper4k.meta_keywords
|
|
|
|
Language:
|
|
1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
|
|
2. trafilatura.language
|
|
3. newspaper4k.meta_lang
|
|
|
|
Top Image:
|
|
1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
|
|
2. newspaper4k.top_image
|
|
3. trafilatura.image
|
|
```
|