feat(converter): implement deterministic JSON to Markdown article converter (spec 005)

This commit is contained in:
2026-08-21 10:30:14 -03:00
parent 64dfd842de
commit 926a6b8cfc
58 changed files with 20138 additions and 2301 deletions
+137
View File
@@ -0,0 +1,137 @@
# Feature Specification: Convert Article JSON to Markdown
**Feature Branch**: `005-convert-json-markdown`
**Created**: 2026-08-21
**Status**: Draft
**Input**: User description: "a partir do markdown docs/prd_convert_json_markdown.md"
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Single Article JSON to Clean Markdown Conversion (Priority: P1)
As a content pipeline operator, I want to convert a single validated article JSON file containing a `selected_extractor` into a clean, structured Markdown document so that downstream consumers and publishing pipelines receive consistent, human-readable, and well-formatted editorial content.
**Why this priority**: This is the core purpose of the feature. Converting an extracted article's primary content and required fields (Title, Original URL, Body) into a standardized Markdown file is the foundational deliverable.
**Independent Test**: Can be tested independently by providing a single article JSON with `selected_extractor` (for each supported extractor: `trafilatura`, `newspaper4k`, `readability`), running the CLI conversion, and verifying that the generated `.md` file matches expected structure, formatting, and content.
**Acceptance Scenarios**:
1. **Given** a valid JSON of a single article with `selected_extractor` set to `trafilatura` and `trafilatura.markdown` populated, **When** conversion executes, **Then** the resulting Markdown file uses the Trafilatura Markdown content directly without incorporating text from other extractors.
2. **Given** a valid JSON of a single article with `selected_extractor` set to `newspaper4k` and `newspaper4k.article_html` populated, **When** conversion executes, **Then** the HTML content is converted to Markdown with ATX headings and editorial structure preserved.
3. **Given** a valid JSON of a single article with `selected_extractor` set to `readability` and `readability.cleaned_html` populated, **When** conversion executes, **Then** the cleaned HTML is converted to Markdown preserving formatting and content hierarchy.
4. **Given** a valid JSON of a single article where the structured body of the selected extractor is missing or empty but its raw text is available, **When** conversion executes, **Then** the raw text from the same selected extractor is used as fallback.
---
### User Story 2 - Deterministic Metadata Resolution and Fallback (Priority: P2)
As a pipeline operator, I want the system to deterministically resolve optional and required metadata across all available extractor and input fields according to a strict priority hierarchy, so that missing metadata in the selected extractor is enriched from secondary sources without non-deterministic or probabilistic behavior.
**Why this priority**: While the body text must strictly come from the selected extractor, metadata (authors, publication date, site name, categories, tags, keywords, language, top image, subtitle) often varies across extractors. Deterministic fallback guarantees maximum metadata completeness while maintaining reproducible output.
**Independent Test**: Can be tested with synthetic and real JSON fixtures containing missing metadata in the selected extractor but present in secondary extractors or `input_meta`, verifying that the output metadata lines follow the defined hierarchy exactly and omit empty fields/placeholders cleanly.
**Acceptance Scenarios**:
1. **Given** a selected extractor lacking author or publication date, but secondary extractor blocks or `input_meta` containing valid candidates, **When** conversion executes, **Then** the first valid candidate in the priority order is rendered in the metadata section.
2. **Given** metadata candidates with surrounding whitespace, HTML entities, or known placeholder strings (e.g., `null`, `none`, `n/a`, `unknown`, `[no-author]`), **When** conversion executes, **Then** values are sanitized and placeholders are treated as absent, causing fallback to the next candidate or total omission.
3. **Given** an article where no valid candidates exist for optional fields (e.g., tags, category, subtitle), **When** conversion executes, **Then** those metadata lines are completely omitted without generating blank lines, empty labels, or placeholder text.
4. **Given** a body containing a top-level H1 header identical to the resolved article title, **When** conversion executes, **Then** the duplicate H1 is removed from the body to prevent repeating the title.
---
### User Story 3 - CLI Usability, Validation, and Atomic Output (Priority: P3)
As a DevOps or system integration engineer, I want the converter tool to operate via a clear command-line interface with custom output path support, clear exit codes, detailed error feedback on `stderr`, and atomic file writing, so that pipeline automation can safely run unattended and never produce corrupt or partial files.
**Why this priority**: Reliable automation in unattended batch pipelines requires deterministic exit codes, zero corrupt state on failure, and safe atomic file writing.
**Independent Test**: Can be tested by invoking the CLI with valid arguments, custom `-o` paths, invalid inputs (malformed JSON, batch JSON with `articles` array, missing mandatory fields), verifying stdout/stderr streams, file system state, and process exit codes (0, 1, 2).
**Acceptance Scenarios**:
1. **Given** a valid single-article JSON input and no `-o` argument, **When** the CLI runs, **Then** it produces `<input_stem>.md` atomically in the same directory and exits with code `0`.
2. **Given** an invalid input JSON containing a root `articles` array (batch file), **When** the CLI runs, **Then** it terminates with exit code `1`, logs a descriptive error message to `stderr`, and writes no output file.
3. **Given** a pre-existing output file and an error occurring during validation or conversion, **When** the CLI terminates, **Then** the pre-existing output file remains completely unmodified and no temporary files linger.
4. **Given** missing or invalid CLI arguments, **When** the CLI runs, **Then** it exits with code `2` as standard for argument parsing errors.
---
### Edge Cases
- **Batch JSON Input**: If the JSON root contains an `articles` key, the process fails immediately with exit code `1` and instructs the user that only single-article objects are accepted.
- **Strict Extractor Isolation for Body**: If the `selected_extractor` has no usable body (both HTML/Markdown and plain text are empty), the process fails with exit code `1`. The system MUST NEVER fall back to another extractor for the body.
- **Unresolved Mandatory Metadata**: If title, valid absolute original URL (`http`/`https`), or body cannot be resolved from any candidate source, conversion fails with exit code `1`.
- **Invalid Body Images**: Images in the converted body with relative paths, empty URLs, or `data:` URIs are stripped. Duplicate image URLs within the body are deduplicated, keeping the first occurrence.
- **Date Formatting Variety**: Dates formatted as ISO 8601 or RFC 2822 are parsed and standardized to ISO 8601 preserving time zone offsets, or `YYYY-MM-DD` if date-only. Unparseable dates fall back to the next candidate.
- **List and String Flexibility**: Author, tag, keyword, and category fields provided as either lists or semicolon-delimited strings are normalized, deduplicated case-insensitively (preserving original casing of the first instance), and stripped of invalid entries (such as author strings starting with `http://`, `https://`, or `www.`).
- **Subtitle Equal to Title**: If the resolved subtitle/description matches the resolved title (case-insensitively after normalization), the subtitle is omitted from the output.
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST accept an input path to a UTF-8 JSON file representing a single article object.
- **FR-002**: System MUST reject root structures containing an `articles` key or root elements that are not JSON objects, exiting with code `1`.
- **FR-003**: System MUST require and validate `selected_extractor` to be one of: `trafilatura`, `newspaper4k`, or `readability`. Any other value or missing key MUST terminate with exit code `1`.
- **FR-004**: System MUST extract the article body exclusively from the chosen `selected_extractor` using its designated primary content field or fallback text field within that same extractor:
- `trafilatura`: primary `trafilatura.markdown`, fallback `trafilatura.text`
- `newspaper4k`: primary `newspaper4k.article_html` (converted to Markdown), fallback `newspaper4k.text`
- `readability`: primary `readability.cleaned_html` (converted to Markdown), fallback `readability.cleaned_text`
- **FR-005**: System MUST NOT perform cross-extractor fallback for article body content; if the selected extractor's content is empty or unusable, processing MUST fail with exit code `1`.
- **FR-006**: System MUST resolve mandatory fields (Title, Original URL, Body), failing with exit code `1` if any mandatory field cannot be resolved to a non-empty valid value.
- **FR-007**: System MUST validate original URLs and top image URLs to ensure they are absolute URLs with `http` or `https` schemes.
- **FR-008**: System MUST resolve metadata fields deterministically using the defined priority order:
- **Title**: `SELECIONADO.title` → `input_meta.titulo` → `page_title` → `newspaper4k.title` → `trafilatura.title` → `readability.title`
- **Original URL**: `input_meta.url` → `crawled_url` → canonical URL of selected → `trafilatura.canonical_url` → `newspaper4k.canonical_link`
- **Subtitle / Description**: description of selected → `trafilatura.description` → `newspaper4k.meta_description` → `input_meta.subtitulo`
- **Authors**: author(s) of selected → `newspaper4k.authors` → `trafilatura.author` → `readability.author`
- **Publication Date**: date of selected → `newspaper4k.publish_date` → `trafilatura.date` → `input_meta.quando_publicado`
- **Site Name**: site name of selected → `trafilatura.sitename` → `newspaper4k.meta_site_name` → `trafilatura.hostname` → hostname from original URL
- **Categories**: categories of selected → `trafilatura.categories`
- **Tags**: tags of selected → `trafilatura.tags` → `newspaper4k.tags` → `newspaper4k.meta_keywords`
- **Keywords**: `newspaper4k.keywords` → `newspaper4k.meta_keywords`
- **Language**: language of selected → `trafilatura.language` → `newspaper4k.meta_lang`
- **Top Image**: image of selected → `newspaper4k.top_image` → `trafilatura.image`
- **FR-009**: System MUST normalize scalar metadata strings by decoding HTML entities, trimming leading/trailing whitespace, collapsing internal consecutive whitespace, and discarding known placeholders (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`).
- **FR-010**: System MUST normalize list metadata (authors, categories, tags, keywords) from arrays or semicolon-separated strings, trimming items, discarding author entries that are URLs, deduplicating case-insensitively while preserving original order and initial casing, and picking the first valid source list without merging lists across different sources.
- **FR-011**: System MUST parse publication dates in ISO 8601 or RFC 2822 formats and format them as standard ISO 8601 (preserving time zone) or `YYYY-MM-DD` (for date-only values).
- **FR-012**: System MUST convert HTML bodies to Markdown using ATX headings (`#`, `##`, `###`), preserving paragraph structure, formatting (bold, italic), lists, blockquotes, code blocks, tables, and valid body images.
- **FR-013**: System MUST strip invalid body images (relative paths, `data:` URIs, empty URLs) and deduplicate repeated occurrences of identical image URLs in the body.
- **FR-014**: System MUST remove the initial H1 heading from the converted body if it matches the resolved article title (after HTML entity decoding and whitespace/case normalization).
- **FR-015**: System MUST assemble the final Markdown document in strict section order:
1. `# [Title]`
2. Subtitle/Description (omitted if empty or equal to title)
3. Metadata key-value block (`**Autor:**`, `**Publicado em:**`, `**Site:**`, `**Categoria:**`, `**Tags:**`, `**Palavras-chave:**`, `**Idioma:**`, `**Fonte original:** [URL](URL)`)
4. Main image `![Imagem principal](URL)` (omitted if invalid or absent)
5. Horizontal rule separator `---`
6. Converted body content
- **FR-016**: System MUST format the final Markdown with `LF` line endings, a single trailing newline, no trailing whitespace per line, and at most two consecutive blank lines.
- **FR-017**: System MUST implement atomic file writing (write to temporary file then replace target atomically) and ensure existing target files remain untouched if processing fails.
- **FR-018**: System MUST provide a CLI script `scripts/convert_article_to_markdown.py` supporting `-i/--input` (mandatory) and `-o/--output` (optional, defaulting to `<input_stem>.md`), returning exit codes `0` (success), `1` (runtime/validation/conversion error), and `2` (CLI argument error), with diagnostics written to `stderr`.
### Key Entities
- **Article JSON Input**: The source data object representing a single crawled and extracted news article containing extractor blocks (`trafilatura`, `newspaper4k`, `readability`), `selected_extractor` tag, and optional crawl metadata (`input_meta`, `crawled_url`, `page_title`).
- **Resolved Article Metadata**: The normalized, sanitized, and prioritized editorial properties extracted across candidate sources (Title, Original URL, Subtitle, Authors, Publication Date, Site Name, Categories, Tags, Keywords, Language, Top Image).
- **Output Markdown Document**: The final UTF-8 formatted document containing structured header metadata, visual assets, separator, and normalized editorial article body text.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: 100% of single-article JSON files with valid required fields generate valid Markdown documents adhering to the prescribed section hierarchy.
- **SC-002**: 100% byte-for-byte determinism: identical JSON inputs processed across multiple runs produce identical Markdown output files.
- **SC-003**: 0% cross-extractor body pollution: in all test cases, body text originates strictly and solely from the specified `selected_extractor`.
- **SC-004**: 0% placeholder leakage: no output document contains `null`, `None`, `N/A`, `unknown`, `[no-author]`, or empty metadata labels.
- **SC-005**: 100% atomic integrity: failed conversions leave zero leftover temporary files and never corrupt or overwrite existing target files.
- **SC-006**: Sub-second execution: conversion of a single standard news article JSON completes in under 200ms in a local execution environment.
## Assumptions
- The input JSON is UTF-8 encoded and represents a single article item extracted from upstream crawlers/extractors.
- Python 3.10+ standard libraries and `markdownify` library are available in the runtime environment.
- No network requests, browser automation, or external AI/LLM services are needed or permitted during conversion.
- Extractor-specific JSON schemas match the structures produced by previous pipeline stages (`003-article-content-extractor` and `004-deterministic-content-selection`).
- All execution and file manipulation occurs on the local filesystem.