feat(converter): implement deterministic JSON to Markdown article converter (spec 005)
This commit is contained in:
@@ -0,0 +1,86 @@
|
||||
# Markdown Conversion Checklist: End-to-End Requirements Quality
|
||||
|
||||
**Purpose**: Validate the completeness, clarity, consistency, and measurability of requirements for the single-article JSON to Markdown conversion pipeline, ensuring 100% adherence to PRD `docs/prd_convert_json_markdown.md`.
|
||||
**Created**: 2026-08-21
|
||||
**Feature**: [spec.md](../spec.md) | **Plan**: [plan.md](../plan.md) | **Data Model**: [data-model.md](../data-model.md) | **PRD**: [prd_convert_json_markdown.md](../../../docs/prd_convert_json_markdown.md)
|
||||
|
||||
**Note**: This custom checklist is generated and reviewed for complete requirements quality.
|
||||
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality against the PRD. It does not mean implementation work is complete.
|
||||
|
||||
---
|
||||
|
||||
## 1. Validação de Entrada, Tipagem & Isolamento de Lotes
|
||||
|
||||
- [x] CHK001 Is the requirement to accept only a single-article JSON object (and explicitly reject root structures containing `articles`) unambiguous and testable? [Clarity, Spec §FR-001, §FR-002, PRD §6.1, §8.1]
|
||||
- [x] CHK002 Is input encoding explicitly specified as UTF-8 without BOM with strict JSON parsing validation? [Completeness, Spec §FR-001, PRD §6.1, §10]
|
||||
- [x] CHK003 Is the allowed set of `selected_extractor` values (`trafilatura`, `newspaper4k`, `readability`) strictly bounded, rejecting missing/unknown values? [Clarity, Spec §FR-003, PRD §6.2, §10]
|
||||
- [x] CHK004 Are mandatory resolved output fields (non-empty Title, valid absolute Original URL, non-empty Body) explicitly defined as non-negotiable gates? [Completeness, Spec §FR-006, PRD §6.3, §10]
|
||||
- [x] CHK005 Is the behavior for non-object JSON roots (e.g. lists, primitives) specified to exit with code `1`? [Edge Case, Spec §FR-002, PRD §10]
|
||||
|
||||
## 2. Isolamento Estrito de Extrator & Conversão de Conteúdo (HTML/MD)
|
||||
|
||||
- [x] CHK006 Is the strict isolation rule prohibiting cross-extractor body fallback explicitly defined, terminating with code `1` if the selected extractor has no body? [Consistency, Spec §FR-005, PRD §8.2, §12 CA-005]
|
||||
- [x] CHK007 Are primary and intra-extractor fallback fields unambiguously mapped for all three extractors (`trafilatura.markdown` → `text`, `newspaper4k.article_html` → `text`, `readability.cleaned_html` → `cleaned_text`)? [Completeness, Spec §FR-004, PRD §8.3]
|
||||
- [x] CHK008 Is direct Markdown reuse for Trafilatura specified without redundant HTML re-parsing? [Clarity, Spec §FR-004, PRD §8.3, §12 CA-001]
|
||||
- [x] CHK009 Are HTML-to-Markdown conversion rules via `markdownify` (ATX headings `#`, `##`, `###`, paragraphs, lists, tables, quotes, bold, italic, code blocks, links, body images) fully documented? [Completeness, Spec §FR-012, PRD §8.10, §15]
|
||||
- [x] CHK010 Is it explicitly specified that external collections like `newspaper4k.images` must NOT be injected into the converted body? [Clarity, Spec §FR-013, PRD §8.11]
|
||||
|
||||
## 3. Resolução Determinística de Metadados & Mapeamento de SELECIONADO
|
||||
|
||||
- [x] CHK011 Are candidate priority fallback chains documented for every metadata field without gaps or ambiguity? [Coverage, Spec §FR-008, PRD §8.5]
|
||||
- [x] CHK012 Is the exact field mapping for `SELECIONADO` defined per extractor (including unavailable fields for Readability and Newspaper4k)? [Completeness, Data Model §2, PRD §8.5]
|
||||
- [x] CHK013 Is the fallback hierarchy for `Site Name` (`SELECIONADO` → `trafilatura.sitename` → `newspaper4k.meta_site_name` → `trafilatura.hostname` → Original URL hostname) fully covered? [Completeness, Spec §FR-008, PRD §8.5]
|
||||
- [x] CHK014 Is the first-valid-source rule (picking the first valid source in priority order without merging values across different sources) clearly established? [Consistency, Spec §FR-010, PRD §8.4, §8.7]
|
||||
|
||||
## 4. Normalização de Escalares, Placeholders & Sanitização de Listas
|
||||
|
||||
- [x] CHK015 Are scalar string normalization rules (HTML entity unescaping, trimming leading/trailing whitespace, collapsing internal consecutive whitespace) testable and unambiguous? [Clarity, Spec §FR-009, PRD §8.6]
|
||||
- [x] CHK016 Is the blacklist of ignored placeholder values (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author` - case-insensitive) exhaustively defined? [Completeness, Spec §FR-009, PRD §8.6]
|
||||
- [x] CHK017 Are list normalization rules specified for both array inputs and single strings delimited exclusively by semicolons (`;`)? [Completeness, Spec §FR-010, PRD §8.7]
|
||||
- [x] CHK018 Is the filter discarding author entries starting with `http://`, `https://`, or `www.` explicitly defined? [Edge Case, Spec §FR-010, PRD §8.7]
|
||||
- [x] CHK019 Is case-insensitive deduplication for list fields defined to preserve the original casing and first occurrence order? [Clarity, Spec §FR-010, PRD §8.7]
|
||||
|
||||
## 5. Normalização de Datas, Fusos & Validação Estrita de URLs
|
||||
|
||||
- [x] CHK020 Are date parsing expectations (supporting ISO 8601 and RFC 2822) with timezone preservation and `YYYY-MM-DD` date-only output format explicitly documented? [Clarity, Spec §FR-011, PRD §8.8]
|
||||
- [x] CHK021 Is the behavior for unparseable date candidates specified to discard and advance to the next priority source? [Edge Case, Spec §FR-011, PRD §8.8]
|
||||
- [x] CHK022 Are URL validation criteria (absolute `http`/`https` with non-empty hostname, rejecting `data:`, `javascript:`, and relative paths) defined for Original URL and Top Image without performing network calls? [Clarity, Spec §FR-007, PRD §8.9]
|
||||
|
||||
## 6. Sanitização Editorial, Tratamento de Imagens & Título Duplicado
|
||||
|
||||
- [x] CHK023 Are criteria for stripping the initial H1 heading from the body (exact match with resolved title after entity decoding, whitespace collapsing, and case-insensitive comparison) objectively measurable? [Measurability, Spec §FR-014, PRD §8.12, §12 CA-010]
|
||||
- [x] CHK024 Is subtitle omission behavior when identical to the resolved title (after normalization) clearly specified? [Clarity, Spec §FR-015, PRD §7.2]
|
||||
- [x] CHK025 Are body image sanitation rules (keeping only absolute `http`/`https`, removing relative/empty/`data:` images, deduplicating identical URLs) completely covered? [Coverage, Spec §FR-013, PRD §8.11, §12 CA-009]
|
||||
- [x] CHK026 Is the main top image presentation format `` specified, and omitted when absent or invalid? [Completeness, Spec §FR-015, PRD §7.2]
|
||||
|
||||
## 7. Estrutura, Sintaxe do Markdown de Saída & Restrições
|
||||
|
||||
- [x] CHK027 Is the final Markdown section order (`# Title` → Subtitle → Metadata Block → Main Image → `---` → Body) explicitly defined? [Completeness, Spec §FR-015, Contract §Markdown-Schema, PRD §7.2]
|
||||
- [x] CHK028 Are metadata label formatting rules (`**Autor:**`, `**Publicado em:**`, `**Site:**`, `**Categoria:**`, `**Tags:**`, `**Palavras-chave:**`, `**Idioma:**`, `**Fonte original:** [URL](URL)`) strictly defined, omitting empty labels entirely? [Completeness, Contract §Markdown-Schema, PRD §7.2]
|
||||
- [x] CHK029 Is it explicitly required that `selected_extractor` name and internal JSON debugging metadata MUST NEVER appear in the generated Markdown? [Consistency, Contract §Markdown-Schema, PRD §7.2]
|
||||
- [x] CHK030 Are whitespace and formatting constraints (UNIX `LF` line endings, exactly 1 trailing newline at EOF, no trailing spaces per line, at most 2 consecutive newlines, UTF-8 unicode preservation) measurable? [Measurability, Spec §FR-016, PRD §8.13]
|
||||
|
||||
## 8. Interface CLI, Tratamento de Erros & Atomicidade
|
||||
|
||||
- [x] CHK031 Is the CLI script path `scripts/convert_article_to_markdown.py` and arguments (`-i/--input` required, `-o/--output` optional defaulting to `<input_stem>.md`) defined? [Completeness, Spec §FR-018, Contract §CLI, PRD §9]
|
||||
- [x] CHK032 Are exit codes (`0` for success, `1` for validation/runtime error, `2` for argument syntax error) explicitly documented? [Completeness, Spec §FR-018, Contract §CLI, PRD §9.4]
|
||||
- [x] CHK033 Is stream routing specified (all error and informational diagnostics to `stderr`, no Markdown dumped to `stdout` on file write)? [Clarity, Contract §CLI, PRD §9.4]
|
||||
- [x] CHK034 Is error message sanitization specified to ensure full article contents are never dumped to the terminal during failures? [Security/UX, PRD §10]
|
||||
- [x] CHK035 Are atomic file write requirements (temporary file in same directory + atomic replacement via `os.replace`, with full cleanup on error leaving pre-existing targets untouched) defined? [Non-Functional, Spec §FR-017, PRD §8.14, §11 RNF-004]
|
||||
|
||||
## 9. Estratégia de Testes, Golden Fixtures & Definition of Done
|
||||
|
||||
- [x] CHK036 Are Golden Test Fixtures required for all 3 extractors (`valid_trafilatura.json` → `.md`, `valid_newspaper4k.json` → `.md`, `valid_readability.json` → `.md`) with exact byte-for-byte matching? [Test Quality, PRD §13.3, §14, Spec §SC-002]
|
||||
- [x] CHK037 Are negative test fixtures required for invalid JSON, batch `articles` array, missing/unknown extractor, empty selected body, missing title, and invalid original URL? [Coverage, PRD §13.3]
|
||||
- [x] CHK038 Are non-functional constraints (100% deterministic execution, fully local memory processing, no external network requests, sub-second latency) documented as testable gates? [Non-Functional, Spec §SC-006, PRD §11]
|
||||
- [x] CHK039 Are Quality Gates (Ruff, Mypy, Pytest, SonarQube) and README documentation defined as mandatory completion criteria? [Completeness, Plan §Constitution-Check, PRD §14]
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All 39 items have been rigorously validated against PRD `docs/prd_convert_json_markdown.md`, `specs/005-convert-json-markdown/spec.md`, `plan.md`, and `data-model.md`.
|
||||
- 100% adherence to PRD requirements with zero ambiguity or unhandled edge cases.
|
||||
- All items are marked `[x]` confirming requirements-quality criteria satisfaction.
|
||||
- Ready for `/speckit-tasks` to break down implementation and TDD test tasks.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Specification Quality Checklist: Convert Article JSON to Markdown
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-08-21
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- All requirements were extracted directly from PRD `docs/prd_convert_json_markdown.md`.
|
||||
- No ambiguity remains; strict priority tables, validation rules, normalization procedures, and edge cases are completely defined.
|
||||
- Ready for `/speckit-plan`.
|
||||
@@ -0,0 +1,31 @@
|
||||
# CLI Contract: `convert_article_to_markdown.py`
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Script Signature
|
||||
|
||||
```bash
|
||||
python scripts/convert_article_to_markdown.py -i <input_path> [-o <output_path>]
|
||||
```
|
||||
|
||||
## 2. Command-Line Arguments
|
||||
|
||||
| Flag | Long Option | Type | Required | Default | Description |
|
||||
|---|---|---|:---:|---|---|
|
||||
| `-i` | `--input` | String / Path | Yes | — | Path to the source JSON file containing exactly one article object. |
|
||||
| `-o` | `--output` | String / Path | No | `<input_stem>.md` | Destination path for the generated Markdown file. |
|
||||
|
||||
## 3. Exit Codes
|
||||
|
||||
| Exit Code | Meaning | Standard Streams Behavior |
|
||||
|:---:|---|---|
|
||||
| `0` | **Success**: Article converted and Markdown written atomically. | Diagnostic info on `stderr`, clean execution. |
|
||||
| `1` | **Runtime / Validation Error**: Malformed JSON, root `articles` array, missing mandatory fields (title, original URL, body), invalid selected extractor, or write failure. | Descriptive error message printed to `stderr`. Pre-existing target file unmodified. |
|
||||
| `2` | **Argument Error**: Missing required `-i/--input` argument, unrecognized arguments, or invalid CLI usage. | Standard `argparse` usage and error output printed to `stderr`. |
|
||||
|
||||
## 4. Standard Stream Behavior
|
||||
|
||||
- **`stdout`**: Reserved. No Markdown text is dumped to `stdout` when generating file output.
|
||||
- **`stderr`**: Receives progress/error diagnostics, e.g.:
|
||||
- `[INFO] Converted 'out/article_001.json' -> 'out/article_001.md' (extractor: trafilatura)`
|
||||
- `[ERROR] Invalid input: JSON contains batch 'articles' array. Only single article JSON files are accepted.`
|
||||
@@ -0,0 +1,38 @@
|
||||
# Markdown Schema Contract: Output Article Markdown
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Output Document Specification
|
||||
|
||||
The generated Markdown document MUST strictly adhere to the following template structure:
|
||||
|
||||
```markdown
|
||||
# {Resolved Title}
|
||||
|
||||
{Resolved Subtitle/Description - OMITTED IF ABSENT OR EQUAL TO TITLE}
|
||||
|
||||
**Autor:** {Resolved Authors joined by ", " - OMITTED IF ABSENT}
|
||||
**Publicado em:** {Resolved Publication Date - OMITTED IF ABSENT}
|
||||
**Site:** {Resolved Site Name - OMITTED IF ABSENT}
|
||||
**Categoria:** {Resolved Categories joined by ", " - OMITTED IF ABSENT}
|
||||
**Tags:** {Resolved Tags joined by ", " - OMITTED IF ABSENT}
|
||||
**Palavras-chave:** {Resolved Keywords joined by ", " - OMITTED IF ABSENT}
|
||||
**Idioma:** {Resolved Language Code - OMITTED IF ABSENT}
|
||||
**Fonte original:** [{Resolved Original URL}]({Resolved Original URL})
|
||||
|
||||

|
||||
|
||||
---
|
||||
|
||||
{Converted Markdown Body Content}
|
||||
```
|
||||
|
||||
## 2. Formatting & Syntax Constraints
|
||||
|
||||
1. **Character Encoding**: UTF-8 without BOM.
|
||||
2. **Line Delimiters**: UNIX-style `LF` (`\n`).
|
||||
3. **Trailing Whitespace**: Stripped from every line.
|
||||
4. **Blank Lines**: Maximum of 2 consecutive newline characters (`\n\n`), preventing excessive vertical spacing.
|
||||
5. **EOF Delimiter**: Ends with exactly one trailing newline character (`\n`).
|
||||
6. **No Placeholders**: Never print `null`, `None`, `N/A`, `unknown`, `[no-author]`, or empty metadata labels (e.g. `**Autor:** `).
|
||||
7. **No Internal Leakage**: Never output `selected_extractor` name, scoring metrics, or JSON internals in the final document.
|
||||
@@ -0,0 +1,156 @@
|
||||
# Data Model: Convert Article JSON to Markdown
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Domain Entities & Schemas
|
||||
|
||||
### Entity 1: `ArticleInput` (Source JSON)
|
||||
|
||||
Represents the raw parsed JSON structure of a single news article.
|
||||
|
||||
```text
|
||||
ArticleInput
|
||||
├── selected_extractor: string [REQUIRED: "trafilatura" | "newspaper4k" | "readability"]
|
||||
├── crawled_url: string [OPTIONAL: URL string]
|
||||
├── page_title: string [OPTIONAL: Page title string]
|
||||
├── input_meta: object [OPTIONAL]
|
||||
│ ├── url: string [OPTIONAL]
|
||||
│ ├── titulo: string [OPTIONAL]
|
||||
│ ├── subtitulo: string [OPTIONAL]
|
||||
│ └── quando_publicado: string [OPTIONAL]
|
||||
├── trafilatura: ExtractorBlock [CONDITIONAL]
|
||||
├── newspaper4k: ExtractorBlock [CONDITIONAL]
|
||||
└── readability: ExtractorBlock [CONDITIONAL]
|
||||
```
|
||||
|
||||
#### Validation Rules:
|
||||
- Root must be a JSON object (dict).
|
||||
- Root must NOT contain an `articles` key (batch JSON is rejected).
|
||||
- `selected_extractor` must be exactly one of `"trafilatura"`, `"newspaper4k"`, `"readability"`.
|
||||
- The object corresponding to `selected_extractor` must exist in `ArticleInput` and contain usable body content.
|
||||
|
||||
---
|
||||
|
||||
### Entity 2: `ExtractorBlock` (Per-Extractor Data)
|
||||
|
||||
Represents the extraction results produced by each individual extractor library.
|
||||
|
||||
| Extractor | Primary Body Field | Fallback Body Field | Metadata Fields Available |
|
||||
|---|---|---|---|
|
||||
| `trafilatura` | `markdown` (string) | `text` (string) | `title`, `description`, `author`, `date`, `sitename`, `hostname`, `categories`, `tags`, `language`, `image`, `canonical_url` |
|
||||
| `newspaper4k` | `article_html` (string) | `text` (string) | `title`, `meta_description`, `authors`, `publish_date`, `meta_site_name`, `tags`, `keywords`, `meta_keywords`, `meta_lang`, `top_image`, `canonical_link` |
|
||||
| `readability` | `cleaned_html` (string) | `cleaned_text` (string) | `title`, `author` |
|
||||
|
||||
---
|
||||
|
||||
### Entity 3: `ResolvedArticleMetadata`
|
||||
|
||||
The normalized, validated, and prioritized metadata extracted from candidate sources.
|
||||
|
||||
| Attribute | Type | Mandatory? | Normalization / Validation Rule |
|
||||
|---|---|:---:|---|
|
||||
| `title` | `str` | Yes | Unescaped, trimmed, single spaces. Rejection if empty. |
|
||||
| `original_url` | `str` | Yes | Valid absolute URL with `http://` or `https://` and valid hostname. |
|
||||
| `subtitle` | `Optional[str]` | No | Omitted if empty, placeholder, or equal to `title` (case-insensitive). |
|
||||
| `authors` | `List[str]` | No | Deduplicated, case-preserved, no URL entries. Omitted if empty. |
|
||||
| `publish_date` | `Optional[str]` | No | ISO 8601 string (with timezone) or `YYYY-MM-DD`. Omitted if unparseable. |
|
||||
| `site_name` | `Optional[str]` | No | Normalized site string or fallback to original URL hostname. |
|
||||
| `categories` | `List[str]` | No | Deduplicated, non-empty category strings. |
|
||||
| `tags` | `List[str]` | No | Deduplicated, non-empty tag strings. |
|
||||
| `keywords` | `List[str]` | No | Deduplicated, non-empty keyword strings. |
|
||||
| `language` | `Optional[str]` | No | Language code string (e.g. `es`, `pt`, `en`). |
|
||||
| `top_image` | `Optional[str]` | No | Valid absolute URL with `http://` or `https://`. |
|
||||
|
||||
---
|
||||
|
||||
### Entity 4: `MarkdownDocument`
|
||||
|
||||
The structured representation of the output Markdown file.
|
||||
|
||||
```text
|
||||
MarkdownDocument
|
||||
├── title_h1: "# " + ResolvedArticleMetadata.title
|
||||
├── subtitle_block: Optional paragraph
|
||||
├── metadata_block: Key-value list of bold labels and values
|
||||
│ ├── **Autor:** {authors joined by ", "}
|
||||
│ ├── **Publicado em:** {publish_date}
|
||||
│ ├── **Site:** {site_name}
|
||||
│ ├── **Categoria:** {categories joined by ", "}
|
||||
│ ├── **Tags:** {tags joined by ", "}
|
||||
│ ├── **Palavras-chave:** {keywords joined by ", "}
|
||||
│ ├── **Idioma:** {language}
|
||||
│ └── **Fonte original:** [{original_url}]({original_url})
|
||||
├── top_image_block: Optional ""
|
||||
├── separator: "---"
|
||||
└── body_content: Converted Markdown text (LF line endings, normalized whitespace)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Priority Resolution Matrix
|
||||
|
||||
```text
|
||||
Title:
|
||||
1. SELECIONADO.title
|
||||
2. input_meta.titulo
|
||||
3. page_title
|
||||
4. newspaper4k.title
|
||||
5. trafilatura.title
|
||||
6. readability.title
|
||||
|
||||
Original URL:
|
||||
1. input_meta.url
|
||||
2. crawled_url
|
||||
3. Canonical URL of SELECIONADO (trafilatura.canonical_url or newspaper4k.canonical_link)
|
||||
4. trafilatura.canonical_url
|
||||
5. newspaper4k.canonical_link
|
||||
|
||||
Subtitle / Description:
|
||||
1. Description of SELECIONADO (trafilatura.description or newspaper4k.meta_description)
|
||||
2. trafilatura.description
|
||||
3. newspaper4k.meta_description
|
||||
4. input_meta.subtitulo
|
||||
|
||||
Authors:
|
||||
1. Authors of SELECIONADO (newspaper4k.authors, trafilatura.author, readability.author)
|
||||
2. newspaper4k.authors
|
||||
3. trafilatura.author
|
||||
4. readability.author
|
||||
|
||||
Publication Date:
|
||||
1. Date of SELECIONADO (newspaper4k.publish_date or trafilatura.date)
|
||||
2. newspaper4k.publish_date
|
||||
3. trafilatura.date
|
||||
4. input_meta.quando_publicado
|
||||
|
||||
Site Name:
|
||||
1. Site of SELECIONADO (trafilatura.sitename or newspaper4k.meta_site_name)
|
||||
2. trafilatura.sitename
|
||||
3. newspaper4k.meta_site_name
|
||||
4. trafilatura.hostname
|
||||
5. Hostname of resolved Original URL
|
||||
|
||||
Categories:
|
||||
1. Categories of SELECIONADO (trafilatura.categories)
|
||||
2. trafilatura.categories
|
||||
|
||||
Tags:
|
||||
1. Tags of SELECIONADO (trafilatura.tags or newspaper4k.tags)
|
||||
2. trafilatura.tags
|
||||
3. newspaper4k.tags
|
||||
4. newspaper4k.meta_keywords
|
||||
|
||||
Keywords:
|
||||
1. newspaper4k.keywords
|
||||
2. newspaper4k.meta_keywords
|
||||
|
||||
Language:
|
||||
1. Language of SELECIONADO (trafilatura.language or newspaper4k.meta_lang)
|
||||
2. trafilatura.language
|
||||
3. newspaper4k.meta_lang
|
||||
|
||||
Top Image:
|
||||
1. Image of SELECIONADO (newspaper4k.top_image or trafilatura.image)
|
||||
2. newspaper4k.top_image
|
||||
3. trafilatura.image
|
||||
```
|
||||
@@ -0,0 +1,115 @@
|
||||
# Implementation Plan: Convert Article JSON to Markdown
|
||||
|
||||
**Branch**: `005-convert-json-markdown` | **Date**: 2026-08-21 | **Spec**: [spec.md](spec.md)
|
||||
|
||||
**Input**: Feature specification from `specs/005-convert-json-markdown/spec.md`
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
Implement a standalone Python script and modular library component (`scripts/convert_article_to_markdown.py` and supporting functions) that reads a single news article JSON with a `selected_extractor` attribute, performs deterministic metadata resolution across extractor candidates according to strict priority hierarchies, converts HTML bodies to clean Markdown using `markdownify`, strips duplicate H1 title headings, sanitizes body image links, and writes the assembled Markdown document atomically.
|
||||
|
||||
---
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python `>=3.10` (tested on 3.10, 3.11, 3.12)
|
||||
**Primary Dependencies**: `markdownify>=0.13.0`
|
||||
**Standard Library**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils`
|
||||
**Storage**: Local filesystem (JSON input, Markdown `.md` output)
|
||||
**Testing**: `pytest>=7.0.0` (Unit tests, CLI integration tests, exact byte comparison fixtures)
|
||||
**Quality Gates**: `ruff` (linting/formatting), `mypy` (type checking), `pytest`
|
||||
**Target Platform**: Cross-platform (Windows, Linux, macOS)
|
||||
**Project Type**: CLI Script / Modular Data Pipeline Stage
|
||||
**Performance Goals**: `<200ms` per article on standard hardware; 100% byte-for-byte deterministic output
|
||||
**Constraints**: Fully offline / in-memory execution; no network calls; no LLMs/probabilistic algorithms; transactional atomic file writing
|
||||
|
||||
---
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Principle | Requirement | Compliance Analysis | Status |
|
||||
|---|---|---|:---:|
|
||||
| **I. Library-First** | Self-contained, independently testable modular design | Core parsing, normalization, and assembly logic is modularized into testable pure functions. | ✅ PASS |
|
||||
| **II. CLI Interface** | Clean CLI, text/file I/O, error reporting to `stderr`, standard exit codes (0, 1, 2) | CLI exposes `-i/--input` and `-o/--output`, prints diagnostics to `stderr`, and handles errors cleanly. | ✅ PASS |
|
||||
| **III. Test-First** | TDD mandatory: test fixtures, unit tests, and CLI tests before implementation | Comprehensive test suite planned covering all extractors, fallbacks, and edge cases. | ✅ PASS |
|
||||
| **IV. Integration Testing** | CLI end-to-end integration and fixture contract tests | Golden fixtures for `trafilatura`, `newspaper4k`, and `readability` with exact Markdown match verification. | ✅ PASS |
|
||||
| **V. Simplicity & Observability** | YAGNI, standard library where possible, `markdownify` for HTML conversion | Lightweight dependencies, pure Python standard library for date/URL/normalization routines. | ✅ PASS |
|
||||
|
||||
---
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/005-convert-json-markdown/
|
||||
├── spec.md # Feature specification
|
||||
├── plan.md # This file (/speckit-plan command output)
|
||||
├── research.md # Technical research & decisions
|
||||
├── data-model.md # Entities, normalization rules & priority matrix
|
||||
├── quickstart.md # Quickstart & verification guide
|
||||
├── checklists/
|
||||
│ └── requirements.md # Quality checklist
|
||||
└── contracts/
|
||||
├── cli-contract.md # CLI interface definition
|
||||
└── markdown-schema.md # Output Markdown schema contract
|
||||
```
|
||||
|
||||
### Source Code & Test Layout
|
||||
|
||||
```text
|
||||
TextNLPClassifierApp/
|
||||
├── scripts/
|
||||
│ └── convert_article_to_markdown.py # CLI entry point and conversion logic
|
||||
├── tests/
|
||||
│ ├── fixtures/
|
||||
│ │ ├── markdown_conversion/ # Test fixtures (JSON inputs & expected MD outputs)
|
||||
│ │ │ ├── valid_trafilatura.json
|
||||
│ │ │ ├── valid_trafilatura.md
|
||||
│ │ │ ├── valid_newspaper4k.json
|
||||
│ │ │ ├── valid_newspaper4k.md
|
||||
│ │ │ ├── valid_readability.json
|
||||
│ │ │ ├── valid_readability.md
|
||||
│ │ │ ├── batch_articles_invalid.json
|
||||
│ │ │ └── missing_body_invalid.json
|
||||
│ └── test_convert_article_to_markdown.py # Unit and integration test suite
|
||||
├── requirements.txt # Updated with markdownify>=0.13.0
|
||||
└── README.md # Documenting conversion script usage
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Phases
|
||||
|
||||
### Phase 0: Outline & Research *(Completed)*
|
||||
- Resolved technical decisions in [research.md](research.md).
|
||||
- Confirmed `markdownify` as the HTML-to-Markdown engine and pure Python stdlib for dates/URLs.
|
||||
|
||||
### Phase 1: Design & Contracts *(Completed)*
|
||||
- Defined domain entities and priority resolution matrix in [data-model.md](data-model.md).
|
||||
- Defined CLI interface contract in [contracts/cli-contract.md](contracts/cli-contract.md).
|
||||
- Defined Markdown output document contract in [contracts/markdown-schema.md](contracts/markdown-schema.md).
|
||||
- Created [quickstart.md](quickstart.md) validation instructions.
|
||||
|
||||
### Phase 2: Tasks & Implementation Breakdown *(Next: `/speckit-tasks`)*
|
||||
1. Update `requirements.txt` to include `markdownify>=0.13.0`.
|
||||
2. Build unit test fixtures for each extractor and error condition under `tests/fixtures/markdown_conversion/`.
|
||||
3. Implement metadata extraction, normalization, date parsing, and priority resolution functions.
|
||||
4. Implement HTML-to-Markdown conversion, duplicate H1 heading removal, and image URL sanitation.
|
||||
5. Implement Markdown document assembly and atomic file writing.
|
||||
6. Implement CLI argument parsing and error handling in `scripts/convert_article_to_markdown.py`.
|
||||
7. Write comprehensive test suite in `tests/test_convert_article_to_markdown.py`.
|
||||
8. Run linter (`ruff`), type checker (`mypy`), and test suite (`pytest`).
|
||||
9. Update `README.md` with CLI documentation and pipeline examples.
|
||||
|
||||
---
|
||||
|
||||
## Complexity Tracking
|
||||
|
||||
| Violation | Why Needed | Simpler Alternative Rejected Because |
|
||||
|---|---|---|
|
||||
| *None* | All principles satisfied without architectural violations. | N/A |
|
||||
@@ -0,0 +1,58 @@
|
||||
# Quickstart: Convert Article JSON to Markdown
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Prerequisites & Setup
|
||||
|
||||
Ensure the environment has dependencies installed:
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
Verify that `markdownify` is installed:
|
||||
|
||||
```bash
|
||||
python -c "import markdownify; print(markdownify.__version__)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Running the CLI Tool
|
||||
|
||||
### Basic Conversion (Default Output Path)
|
||||
|
||||
Convert a single article JSON to `<stem>.md` in the same directory:
|
||||
|
||||
```bash
|
||||
python scripts/convert_article_to_markdown.py -i out/river_plate_extracted_selected_001.json
|
||||
```
|
||||
|
||||
Output generated: `out/river_plate_extracted_selected_001.md`.
|
||||
|
||||
### Custom Destination Path
|
||||
|
||||
Specify an explicit output path:
|
||||
|
||||
```bash
|
||||
python scripts/convert_article_to_markdown.py \
|
||||
-i out/river_plate_extracted_selected_001.json \
|
||||
-o out/markdown/river_plate_article_001.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Verification & Testing
|
||||
|
||||
### Run All Unit & Integration Tests
|
||||
|
||||
```bash
|
||||
pytest tests/test_convert_article_to_markdown.py -v
|
||||
```
|
||||
|
||||
### Run Linter & Type Checker
|
||||
|
||||
```bash
|
||||
ruff check scripts/convert_article_to_markdown.py src/ tests/
|
||||
mypy scripts/convert_article_to_markdown.py
|
||||
```
|
||||
@@ -0,0 +1,100 @@
|
||||
# Research: Convert Article JSON to Markdown
|
||||
|
||||
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
|
||||
|
||||
## 1. Executive Summary & Goals
|
||||
|
||||
This research addresses the design decisions for converting a single news article JSON with `selected_extractor` (`trafilatura`, `newspaper4k`, or `readability`) into a standardized, clean, human-readable Markdown file (`.md`) with deterministic metadata extraction and fallback rules.
|
||||
|
||||
---
|
||||
|
||||
## 2. Technical Decisions & Research Findings
|
||||
|
||||
### Decision 1: HTML-to-Markdown Engine Selection
|
||||
|
||||
- **Decision**: Use [`markdownify`](https://github.com/matthewwithanm/python-markdownify) with ATX heading style (`heading_style=ATX`).
|
||||
- **Rationale**:
|
||||
- `markdownify` is a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown.
|
||||
- It natively converts `<h1>`–`<h6>` to `#`–`######` (ATX style), parses tables, code blocks, lists, quotes, and inline styles (`<b>`, `<i>`, `<a>`, `<img>`).
|
||||
- Unlike broader conversion tools like Microsoft MarkItDown or Pandoc, `markdownify` has zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions.
|
||||
- **Alternatives Considered**:
|
||||
- *Microsoft MarkItDown*: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion.
|
||||
- *Custom BeautifulSoup converter*: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting).
|
||||
|
||||
---
|
||||
|
||||
### Decision 2: Direct Markdown Handling for Trafilatura
|
||||
|
||||
- **Decision**: When `selected_extractor == "trafilatura"`, directly use `trafilatura.markdown` (or fallback to `trafilatura.text`) without running HTML-to-Markdown conversion.
|
||||
- **Rationale**:
|
||||
- Trafilatura natively emits high-quality Markdown in its extraction output.
|
||||
- Plain text (`trafilatura.text`) is already valid Markdown without special markup.
|
||||
- **Alternatives Considered**:
|
||||
- *Converting Trafilatura's raw HTML*: Inefficient and degrades Trafilatura's native document structural tree.
|
||||
|
||||
---
|
||||
|
||||
### Decision 3: Metadata Normalization & Priority Resolution Pipeline
|
||||
|
||||
- **Decision**: Implement a pure Python deterministic metadata resolver supporting:
|
||||
- Strict priority tables matching the PRD specification.
|
||||
- Scalar normalization: HTML entity decoding (`html.unescape`), whitespace trimming/collapsing, placeholder discarding (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`).
|
||||
- List normalization: Splitting on `;` if string, trimming elements, filtering out URL-like authors (`http://`, `https://`, `www.`), deduplicating case-insensitively while preserving initial case and original order.
|
||||
- First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging.
|
||||
- **Rationale**:
|
||||
- Guarantees 100% deterministic and reproducible metadata output.
|
||||
- Prevents corrupt or placeholder values from leaking into final editorial documents.
|
||||
|
||||
---
|
||||
|
||||
### Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822)
|
||||
|
||||
- **Decision**: Use Python's standard library `datetime.fromisoformat` and `email.utils.parsedate_to_datetime` / standard datetime parsing routines without heavy external dependencies.
|
||||
- **Rationale**:
|
||||
- All input dates observed from extractors follow ISO 8601 (e.g. `2026-08-20T00:36:33-03:00` or `2026-08-20T03:36:33Z`) or RFC 2822 (e.g. `Thu, 20 Aug 2026 00:36:33 -0300`).
|
||||
- `datetime.fromisoformat()` in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'.
|
||||
- `email.utils.parsedate_to_datetime()` standard library natively handles RFC 2822 dates.
|
||||
- If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order.
|
||||
- **Alternatives Considered**:
|
||||
- *dateparser / python-dateutil*: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors.
|
||||
|
||||
---
|
||||
|
||||
### Decision 5: URL Validation & Media Filtering
|
||||
|
||||
- **Decision**:
|
||||
- Validate all URLs with `urllib.parse.urlparse`: Scheme must be `http` or `https`, and `netloc` (hostname) must be non-empty.
|
||||
- In Markdown body: Filter out images with relative URLs, empty URLs, or `data:` URIs.
|
||||
- Deduplicate identical image URLs in the body, keeping only the first occurrence.
|
||||
- **Rationale**:
|
||||
- Prevents broken local references or bloated base64 data URIs in downstream pipelines.
|
||||
|
||||
---
|
||||
|
||||
### Decision 6: Duplicate Title Heading (H1) Stripping
|
||||
|
||||
- **Decision**:
|
||||
- Check the first top-level ATX heading (`# ...`) in the converted body.
|
||||
- If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace.
|
||||
- Preserve all subsequent H1/H2/H3 headings in the body.
|
||||
- **Rationale**:
|
||||
- Many news articles embed the title in `<h1>` inside the article HTML. Since our Markdown schema places `# <Resolved Title>` at the top of the document, stripping the redundant body H1 avoids awkward repeated headings.
|
||||
|
||||
---
|
||||
|
||||
### Decision 7: Atomic File Writing & Error Resilience
|
||||
|
||||
- **Decision**:
|
||||
- Write Markdown output to a temporary file in the same directory (`.<output_filename>.tmp`) and atomically replace the destination using `os.replace` (or `pathlib.Path.replace`).
|
||||
- If any error or validation exception occurs during processing, clean up the temporary file and exit with code `1`, leaving any pre-existing target file untouched.
|
||||
- **Rationale**:
|
||||
- Guarantees transactional file operations in unattended automated batch pipelines.
|
||||
|
||||
---
|
||||
|
||||
## 3. Technology Stack & Dependencies
|
||||
|
||||
- **Runtime**: Python `>=3.10` (tested on 3.10, 3.11, 3.12)
|
||||
- **New Dependency**: `markdownify>=0.13.0`
|
||||
- **Standard Library Modules**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils`
|
||||
- **Testing & Quality**: `pytest`, `ruff`, `mypy`
|
||||
@@ -0,0 +1,137 @@
|
||||
# Feature Specification: Convert Article JSON to Markdown
|
||||
|
||||
**Feature Branch**: `005-convert-json-markdown`
|
||||
**Created**: 2026-08-21
|
||||
**Status**: Draft
|
||||
**Input**: User description: "a partir do markdown docs/prd_convert_json_markdown.md"
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Single Article JSON to Clean Markdown Conversion (Priority: P1)
|
||||
|
||||
As a content pipeline operator, I want to convert a single validated article JSON file containing a `selected_extractor` into a clean, structured Markdown document so that downstream consumers and publishing pipelines receive consistent, human-readable, and well-formatted editorial content.
|
||||
|
||||
**Why this priority**: This is the core purpose of the feature. Converting an extracted article's primary content and required fields (Title, Original URL, Body) into a standardized Markdown file is the foundational deliverable.
|
||||
|
||||
**Independent Test**: Can be tested independently by providing a single article JSON with `selected_extractor` (for each supported extractor: `trafilatura`, `newspaper4k`, `readability`), running the CLI conversion, and verifying that the generated `.md` file matches expected structure, formatting, and content.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a valid JSON of a single article with `selected_extractor` set to `trafilatura` and `trafilatura.markdown` populated, **When** conversion executes, **Then** the resulting Markdown file uses the Trafilatura Markdown content directly without incorporating text from other extractors.
|
||||
2. **Given** a valid JSON of a single article with `selected_extractor` set to `newspaper4k` and `newspaper4k.article_html` populated, **When** conversion executes, **Then** the HTML content is converted to Markdown with ATX headings and editorial structure preserved.
|
||||
3. **Given** a valid JSON of a single article with `selected_extractor` set to `readability` and `readability.cleaned_html` populated, **When** conversion executes, **Then** the cleaned HTML is converted to Markdown preserving formatting and content hierarchy.
|
||||
4. **Given** a valid JSON of a single article where the structured body of the selected extractor is missing or empty but its raw text is available, **When** conversion executes, **Then** the raw text from the same selected extractor is used as fallback.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Deterministic Metadata Resolution and Fallback (Priority: P2)
|
||||
|
||||
As a pipeline operator, I want the system to deterministically resolve optional and required metadata across all available extractor and input fields according to a strict priority hierarchy, so that missing metadata in the selected extractor is enriched from secondary sources without non-deterministic or probabilistic behavior.
|
||||
|
||||
**Why this priority**: While the body text must strictly come from the selected extractor, metadata (authors, publication date, site name, categories, tags, keywords, language, top image, subtitle) often varies across extractors. Deterministic fallback guarantees maximum metadata completeness while maintaining reproducible output.
|
||||
|
||||
**Independent Test**: Can be tested with synthetic and real JSON fixtures containing missing metadata in the selected extractor but present in secondary extractors or `input_meta`, verifying that the output metadata lines follow the defined hierarchy exactly and omit empty fields/placeholders cleanly.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a selected extractor lacking author or publication date, but secondary extractor blocks or `input_meta` containing valid candidates, **When** conversion executes, **Then** the first valid candidate in the priority order is rendered in the metadata section.
|
||||
2. **Given** metadata candidates with surrounding whitespace, HTML entities, or known placeholder strings (e.g., `null`, `none`, `n/a`, `unknown`, `[no-author]`), **When** conversion executes, **Then** values are sanitized and placeholders are treated as absent, causing fallback to the next candidate or total omission.
|
||||
3. **Given** an article where no valid candidates exist for optional fields (e.g., tags, category, subtitle), **When** conversion executes, **Then** those metadata lines are completely omitted without generating blank lines, empty labels, or placeholder text.
|
||||
4. **Given** a body containing a top-level H1 header identical to the resolved article title, **When** conversion executes, **Then** the duplicate H1 is removed from the body to prevent repeating the title.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - CLI Usability, Validation, and Atomic Output (Priority: P3)
|
||||
|
||||
As a DevOps or system integration engineer, I want the converter tool to operate via a clear command-line interface with custom output path support, clear exit codes, detailed error feedback on `stderr`, and atomic file writing, so that pipeline automation can safely run unattended and never produce corrupt or partial files.
|
||||
|
||||
**Why this priority**: Reliable automation in unattended batch pipelines requires deterministic exit codes, zero corrupt state on failure, and safe atomic file writing.
|
||||
|
||||
**Independent Test**: Can be tested by invoking the CLI with valid arguments, custom `-o` paths, invalid inputs (malformed JSON, batch JSON with `articles` array, missing mandatory fields), verifying stdout/stderr streams, file system state, and process exit codes (0, 1, 2).
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Given** a valid single-article JSON input and no `-o` argument, **When** the CLI runs, **Then** it produces `<input_stem>.md` atomically in the same directory and exits with code `0`.
|
||||
2. **Given** an invalid input JSON containing a root `articles` array (batch file), **When** the CLI runs, **Then** it terminates with exit code `1`, logs a descriptive error message to `stderr`, and writes no output file.
|
||||
3. **Given** a pre-existing output file and an error occurring during validation or conversion, **When** the CLI terminates, **Then** the pre-existing output file remains completely unmodified and no temporary files linger.
|
||||
4. **Given** missing or invalid CLI arguments, **When** the CLI runs, **Then** it exits with code `2` as standard for argument parsing errors.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- **Batch JSON Input**: If the JSON root contains an `articles` key, the process fails immediately with exit code `1` and instructs the user that only single-article objects are accepted.
|
||||
- **Strict Extractor Isolation for Body**: If the `selected_extractor` has no usable body (both HTML/Markdown and plain text are empty), the process fails with exit code `1`. The system MUST NEVER fall back to another extractor for the body.
|
||||
- **Unresolved Mandatory Metadata**: If title, valid absolute original URL (`http`/`https`), or body cannot be resolved from any candidate source, conversion fails with exit code `1`.
|
||||
- **Invalid Body Images**: Images in the converted body with relative paths, empty URLs, or `data:` URIs are stripped. Duplicate image URLs within the body are deduplicated, keeping the first occurrence.
|
||||
- **Date Formatting Variety**: Dates formatted as ISO 8601 or RFC 2822 are parsed and standardized to ISO 8601 preserving time zone offsets, or `YYYY-MM-DD` if date-only. Unparseable dates fall back to the next candidate.
|
||||
- **List and String Flexibility**: Author, tag, keyword, and category fields provided as either lists or semicolon-delimited strings are normalized, deduplicated case-insensitively (preserving original casing of the first instance), and stripped of invalid entries (such as author strings starting with `http://`, `https://`, or `www.`).
|
||||
- **Subtitle Equal to Title**: If the resolved subtitle/description matches the resolved title (case-insensitively after normalization), the subtitle is omitted from the output.
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: System MUST accept an input path to a UTF-8 JSON file representing a single article object.
|
||||
- **FR-002**: System MUST reject root structures containing an `articles` key or root elements that are not JSON objects, exiting with code `1`.
|
||||
- **FR-003**: System MUST require and validate `selected_extractor` to be one of: `trafilatura`, `newspaper4k`, or `readability`. Any other value or missing key MUST terminate with exit code `1`.
|
||||
- **FR-004**: System MUST extract the article body exclusively from the chosen `selected_extractor` using its designated primary content field or fallback text field within that same extractor:
|
||||
- `trafilatura`: primary `trafilatura.markdown`, fallback `trafilatura.text`
|
||||
- `newspaper4k`: primary `newspaper4k.article_html` (converted to Markdown), fallback `newspaper4k.text`
|
||||
- `readability`: primary `readability.cleaned_html` (converted to Markdown), fallback `readability.cleaned_text`
|
||||
- **FR-005**: System MUST NOT perform cross-extractor fallback for article body content; if the selected extractor's content is empty or unusable, processing MUST fail with exit code `1`.
|
||||
- **FR-006**: System MUST resolve mandatory fields (Title, Original URL, Body), failing with exit code `1` if any mandatory field cannot be resolved to a non-empty valid value.
|
||||
- **FR-007**: System MUST validate original URLs and top image URLs to ensure they are absolute URLs with `http` or `https` schemes.
|
||||
- **FR-008**: System MUST resolve metadata fields deterministically using the defined priority order:
|
||||
- **Title**: `SELECIONADO.title` → `input_meta.titulo` → `page_title` → `newspaper4k.title` → `trafilatura.title` → `readability.title`
|
||||
- **Original URL**: `input_meta.url` → `crawled_url` → canonical URL of selected → `trafilatura.canonical_url` → `newspaper4k.canonical_link`
|
||||
- **Subtitle / Description**: description of selected → `trafilatura.description` → `newspaper4k.meta_description` → `input_meta.subtitulo`
|
||||
- **Authors**: author(s) of selected → `newspaper4k.authors` → `trafilatura.author` → `readability.author`
|
||||
- **Publication Date**: date of selected → `newspaper4k.publish_date` → `trafilatura.date` → `input_meta.quando_publicado`
|
||||
- **Site Name**: site name of selected → `trafilatura.sitename` → `newspaper4k.meta_site_name` → `trafilatura.hostname` → hostname from original URL
|
||||
- **Categories**: categories of selected → `trafilatura.categories`
|
||||
- **Tags**: tags of selected → `trafilatura.tags` → `newspaper4k.tags` → `newspaper4k.meta_keywords`
|
||||
- **Keywords**: `newspaper4k.keywords` → `newspaper4k.meta_keywords`
|
||||
- **Language**: language of selected → `trafilatura.language` → `newspaper4k.meta_lang`
|
||||
- **Top Image**: image of selected → `newspaper4k.top_image` → `trafilatura.image`
|
||||
- **FR-009**: System MUST normalize scalar metadata strings by decoding HTML entities, trimming leading/trailing whitespace, collapsing internal consecutive whitespace, and discarding known placeholders (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`).
|
||||
- **FR-010**: System MUST normalize list metadata (authors, categories, tags, keywords) from arrays or semicolon-separated strings, trimming items, discarding author entries that are URLs, deduplicating case-insensitively while preserving original order and initial casing, and picking the first valid source list without merging lists across different sources.
|
||||
- **FR-011**: System MUST parse publication dates in ISO 8601 or RFC 2822 formats and format them as standard ISO 8601 (preserving time zone) or `YYYY-MM-DD` (for date-only values).
|
||||
- **FR-012**: System MUST convert HTML bodies to Markdown using ATX headings (`#`, `##`, `###`), preserving paragraph structure, formatting (bold, italic), lists, blockquotes, code blocks, tables, and valid body images.
|
||||
- **FR-013**: System MUST strip invalid body images (relative paths, `data:` URIs, empty URLs) and deduplicate repeated occurrences of identical image URLs in the body.
|
||||
- **FR-014**: System MUST remove the initial H1 heading from the converted body if it matches the resolved article title (after HTML entity decoding and whitespace/case normalization).
|
||||
- **FR-015**: System MUST assemble the final Markdown document in strict section order:
|
||||
1. `# [Title]`
|
||||
2. Subtitle/Description (omitted if empty or equal to title)
|
||||
3. Metadata key-value block (`**Autor:**`, `**Publicado em:**`, `**Site:**`, `**Categoria:**`, `**Tags:**`, `**Palavras-chave:**`, `**Idioma:**`, `**Fonte original:** [URL](URL)`)
|
||||
4. Main image `` (omitted if invalid or absent)
|
||||
5. Horizontal rule separator `---`
|
||||
6. Converted body content
|
||||
- **FR-016**: System MUST format the final Markdown with `LF` line endings, a single trailing newline, no trailing whitespace per line, and at most two consecutive blank lines.
|
||||
- **FR-017**: System MUST implement atomic file writing (write to temporary file then replace target atomically) and ensure existing target files remain untouched if processing fails.
|
||||
- **FR-018**: System MUST provide a CLI script `scripts/convert_article_to_markdown.py` supporting `-i/--input` (mandatory) and `-o/--output` (optional, defaulting to `<input_stem>.md`), returning exit codes `0` (success), `1` (runtime/validation/conversion error), and `2` (CLI argument error), with diagnostics written to `stderr`.
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **Article JSON Input**: The source data object representing a single crawled and extracted news article containing extractor blocks (`trafilatura`, `newspaper4k`, `readability`), `selected_extractor` tag, and optional crawl metadata (`input_meta`, `crawled_url`, `page_title`).
|
||||
- **Resolved Article Metadata**: The normalized, sanitized, and prioritized editorial properties extracted across candidate sources (Title, Original URL, Subtitle, Authors, Publication Date, Site Name, Categories, Tags, Keywords, Language, Top Image).
|
||||
- **Output Markdown Document**: The final UTF-8 formatted document containing structured header metadata, visual assets, separator, and normalized editorial article body text.
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: 100% of single-article JSON files with valid required fields generate valid Markdown documents adhering to the prescribed section hierarchy.
|
||||
- **SC-002**: 100% byte-for-byte determinism: identical JSON inputs processed across multiple runs produce identical Markdown output files.
|
||||
- **SC-003**: 0% cross-extractor body pollution: in all test cases, body text originates strictly and solely from the specified `selected_extractor`.
|
||||
- **SC-004**: 0% placeholder leakage: no output document contains `null`, `None`, `N/A`, `unknown`, `[no-author]`, or empty metadata labels.
|
||||
- **SC-005**: 100% atomic integrity: failed conversions leave zero leftover temporary files and never corrupt or overwrite existing target files.
|
||||
- **SC-006**: Sub-second execution: conversion of a single standard news article JSON completes in under 200ms in a local execution environment.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- The input JSON is UTF-8 encoded and represents a single article item extracted from upstream crawlers/extractors.
|
||||
- Python 3.10+ standard libraries and `markdownify` library are available in the runtime environment.
|
||||
- No network requests, browser automation, or external AI/LLM services are needed or permitted during conversion.
|
||||
- Extractor-specific JSON schemas match the structures produced by previous pipeline stages (`003-article-content-extractor` and `004-deterministic-content-selection`).
|
||||
- All execution and file manipulation occurs on the local filesystem.
|
||||
@@ -0,0 +1,139 @@
|
||||
# Tasks: Convert Article JSON to Markdown
|
||||
|
||||
**Branch**: `005-convert-json-markdown` | **Feature**: Convert Article JSON to Markdown
|
||||
**Spec**: [spec.md](spec.md) | **Plan**: [plan.md](plan.md) | **Data Model**: [data-model.md](data-model.md)
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup (Shared Infrastructure)
|
||||
|
||||
**Purpose**: Project initialization, dependency management, and test fixture directory setup
|
||||
|
||||
- [X] T001 Add `markdownify>=0.13.0` to `requirements.txt`
|
||||
- [X] T002 [P] Create fixtures directory structure in `tests/fixtures/markdown_conversion/`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Golden fixtures, negative test fixtures, and baseline test infrastructure that MUST be complete before user stories begin
|
||||
|
||||
- [X] T003 [P] Create Golden Test Fixtures for Trafilatura in `tests/fixtures/markdown_conversion/valid_trafilatura.json` and `tests/fixtures/markdown_conversion/valid_trafilatura.md`
|
||||
- [X] T004 [P] Create Golden Test Fixtures for Newspaper4k in `tests/fixtures/markdown_conversion/valid_newspaper4k.json` and `tests/fixtures/markdown_conversion/valid_newspaper4k.md`
|
||||
- [X] T005 [P] Create Golden Test Fixtures for Readability in `tests/fixtures/markdown_conversion/valid_readability.json` and `tests/fixtures/markdown_conversion/valid_readability.md`
|
||||
- [X] T006 [P] Create Negative Test Fixtures (`batch_articles_invalid.json`, `corrupt_json_invalid.json`, `missing_extractor_invalid.json`, `missing_body_invalid.json`, `missing_title_invalid.json`, `invalid_url_invalid.json`) in `tests/fixtures/markdown_conversion/`
|
||||
|
||||
**Checkpoint**: Foundational test fixtures ready. User story implementation can begin in strict TDD order.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 - Single Article JSON to Clean Markdown Conversion (Priority: P1) 🎯 MVP
|
||||
|
||||
**Goal**: Convert a single article JSON into a clean Markdown document using exclusively the selected extractor (`trafilatura`, `newspaper4k`, `readability`), supporting direct Markdown reuse and HTML-to-Markdown conversion with intra-extractor fallback.
|
||||
|
||||
**Independent Test**: Provide single-article JSON fixtures for each extractor and verify that the resulting Markdown body text matches the selected extractor's content without cross-extractor leakage.
|
||||
|
||||
### Tests for User Story 1 (TDD) ⚠️
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementing**
|
||||
|
||||
- [X] T007 [P] [US1] Write unit and integration tests for extractor body resolution, strict extractor isolation, and HTML-to-Markdown conversion in `tests/test_convert_article_to_markdown.py`
|
||||
|
||||
### Implementation for User Story 1
|
||||
|
||||
- [X] T008 [US1] Implement body resolution and strict extractor isolation logic (prohibiting cross-extractor fallback) in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T009 [US1] Implement HTML-to-Markdown conversion using `markdownify` (ATX headings) and intra-extractor fallback (`markdown`/`html` → `text`) in `scripts/convert_article_to_markdown.py`
|
||||
|
||||
**Checkpoint**: At this point, User Story 1 is fully functional and testable independently (MVP ready).
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 - Deterministic Metadata Resolution and Fallback (Priority: P2)
|
||||
|
||||
**Goal**: Resolve mandatory and optional metadata across all extractor candidates and input metadata according to strict priority hierarchies, normalizing strings, lists, dates, and URLs, stripping duplicate H1 headers, and formatting the final Markdown structure.
|
||||
|
||||
**Independent Test**: Pass articles with missing metadata in the selected extractor but present in secondary sources; verify that resolved metadata strictly follows priority chains, discards placeholders, sanitizes dates/URLs/images, and formats output accurately.
|
||||
|
||||
### Tests for User Story 2 (TDD) ⚠️
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementing**
|
||||
|
||||
- [X] T010 [P] [US2] Write unit tests for metadata priority chains, scalar normalization, list deduplication, date parsing (ISO 8601 / RFC 2822), URL validation, duplicate H1 heading removal, and body image sanitation in `tests/test_convert_article_to_markdown.py`
|
||||
|
||||
### Implementation for User Story 2
|
||||
|
||||
- [X] T011 [US2] Implement scalar and list normalizers (HTML unescape, whitespace collapsing, placeholder filtering, author URL filtering, case-insensitive deduplication) in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T012 [US2] Implement ISO 8601 and RFC 2822 date parser (with timezone preservation and `YYYY-MM-DD` date-only output) and URL validator in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T013 [US2] Implement deterministic metadata priority resolution matrix in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T014 [US2] Implement duplicate H1 title heading stripper and body image link sanitizer (removing relative/`data:`/empty URLs and deduplicating repeats) in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T015 [US2] Implement final Markdown document layout assembler (Header, Subtitle, Metadata key-values, Top Image, Separator `---`, Body) with LF line endings in `scripts/convert_article_to_markdown.py`
|
||||
|
||||
**Checkpoint**: At this point, User Stories 1 and 2 work together and pass all unit/integration tests.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 - CLI Usability, Validation, and Atomic Output (Priority: P3)
|
||||
|
||||
**Goal**: Provide a production-ready CLI interface with `-i`/`--input` and `-o`/`--output` flags, standard exit codes (0, 1, 2), error logging to `stderr`, and atomic file replacement with transactional error cleanup.
|
||||
|
||||
**Independent Test**: Execute the CLI against valid JSON, invalid JSON, batch `articles` array JSON, and missing files; verify exit codes, stderr diagnostics, and target file integrity.
|
||||
|
||||
### Tests for User Story 3 (TDD) ⚠️
|
||||
> **NOTE: Write these tests FIRST, ensure they FAIL before implementing**
|
||||
|
||||
- [X] T016 [P] [US3] Write CLI integration tests covering `-i`/`-o` flags, exit codes (`0`, `1`, `2`), stderr logging, rejection of batch `articles` JSON, and atomic replacement rollback on failure in `tests/test_convert_article_to_markdown.py`
|
||||
|
||||
### Implementation for User Story 3
|
||||
|
||||
- [X] T017 [US3] Implement CLI argument parsing, input JSON validation (object root, rejection of `articles` key, `selected_extractor` check), and diagnostic logging to `stderr` in `scripts/convert_article_to_markdown.py`
|
||||
- [X] T018 [US3] Implement transactional atomic file writing (`.<output>.tmp` + `os.replace`) with complete cleanup on failure in `scripts/convert_article_to_markdown.py`
|
||||
|
||||
**Checkpoint**: All user stories are fully implemented, resilient, and verified.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: Polish & Cross-Cutting Concerns
|
||||
|
||||
**Purpose**: Quality gate verification, Golden Fixture byte-for-byte validation, and documentation
|
||||
|
||||
- [X] T019 [P] Run full test suite (`pytest tests/test_convert_article_to_markdown.py -v`) and verify 100% exact Golden Fixtures matching
|
||||
- [X] T020 [P] Run code quality and type checks (`ruff check scripts/ src/ tests/` and `mypy scripts/convert_article_to_markdown.py`)
|
||||
- [X] T021 Update `README.md` with conversion CLI documentation, pipeline usage examples, and argument references
|
||||
- [X] T022 Run quickstart validation scenarios from `specs/005-convert-json-markdown/quickstart.md`
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
### Phase Dependencies
|
||||
|
||||
- **Setup (Phase 1)**: No dependencies — start immediately.
|
||||
- **Foundational (Phase 2)**: Depends on Setup (Phase 1) — BLOCKS all user stories.
|
||||
- **User Story 1 (Phase 3)**: Depends on Foundational (Phase 2) — Core MVP.
|
||||
- **User Story 2 (Phase 4)**: Depends on User Story 1 (Phase 3).
|
||||
- **User Story 3 (Phase 5)**: Depends on User Story 2 (Phase 4).
|
||||
- **Polish (Phase 6)**: Depends on all user stories (Phases 3–5) being complete.
|
||||
|
||||
---
|
||||
|
||||
## Parallel Opportunities
|
||||
|
||||
- **Phase 1**: `T002` [P] can run in parallel with `T001`.
|
||||
- **Phase 2**: `T003` [P], `T004` [P], `T005` [P], and `T006` [P] can all be created in parallel.
|
||||
- **Phase 3**: `T007` [P] (tests) can be created in parallel with fixture setup.
|
||||
- **Phase 4**: `T010` [P] (tests) can be written before implementation.
|
||||
- **Phase 5**: `T016` [P] (CLI tests) can be written before CLI wiring.
|
||||
- **Phase 6**: `T019` [P] (pytest) and `T020` [P] (ruff/mypy) can run in parallel.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Story 1 Only)
|
||||
1. Complete Phase 1 (Setup) and Phase 2 (Foundational Fixtures).
|
||||
2. Complete Phase 3 (User Story 1: TDD tests `T007` → Implementation `T008`, `T009`).
|
||||
3. Validate independent execution of User Story 1.
|
||||
|
||||
### Incremental Delivery
|
||||
1. Foundation Ready → MVP (US1: Body Conversion).
|
||||
2. Deliver US2 (Deterministic Metadata & Sanitization).
|
||||
3. Deliver US3 (CLI, Error Handling & Atomic Output).
|
||||
4. Run Polish & Quality Gates (US1 + US2 + US3 verified with 100% tests passing).
|
||||
Reference in New Issue
Block a user