feat(converter): implement deterministic JSON to Markdown article converter (spec 005)

This commit is contained in:
2026-08-21 10:30:14 -03:00
parent 64dfd842de
commit 926a6b8cfc
58 changed files with 20138 additions and 2301 deletions
+100
View File
@@ -0,0 +1,100 @@
# Research: Convert Article JSON to Markdown
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
## 1. Executive Summary & Goals
This research addresses the design decisions for converting a single news article JSON with `selected_extractor` (`trafilatura`, `newspaper4k`, or `readability`) into a standardized, clean, human-readable Markdown file (`.md`) with deterministic metadata extraction and fallback rules.
---
## 2. Technical Decisions & Research Findings
### Decision 1: HTML-to-Markdown Engine Selection
- **Decision**: Use [`markdownify`](https://github.com/matthewwithanm/python-markdownify) with ATX heading style (`heading_style=ATX`).
- **Rationale**:
- `markdownify` is a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown.
- It natively converts `<h1>`–`<h6>` to `#`–`######` (ATX style), parses tables, code blocks, lists, quotes, and inline styles (`<b>`, `<i>`, `<a>`, `<img>`).
- Unlike broader conversion tools like Microsoft MarkItDown or Pandoc, `markdownify` has zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions.
- **Alternatives Considered**:
- *Microsoft MarkItDown*: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion.
- *Custom BeautifulSoup converter*: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting).
---
### Decision 2: Direct Markdown Handling for Trafilatura
- **Decision**: When `selected_extractor == "trafilatura"`, directly use `trafilatura.markdown` (or fallback to `trafilatura.text`) without running HTML-to-Markdown conversion.
- **Rationale**:
- Trafilatura natively emits high-quality Markdown in its extraction output.
- Plain text (`trafilatura.text`) is already valid Markdown without special markup.
- **Alternatives Considered**:
- *Converting Trafilatura's raw HTML*: Inefficient and degrades Trafilatura's native document structural tree.
---
### Decision 3: Metadata Normalization & Priority Resolution Pipeline
- **Decision**: Implement a pure Python deterministic metadata resolver supporting:
- Strict priority tables matching the PRD specification.
- Scalar normalization: HTML entity decoding (`html.unescape`), whitespace trimming/collapsing, placeholder discarding (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`).
- List normalization: Splitting on `;` if string, trimming elements, filtering out URL-like authors (`http://`, `https://`, `www.`), deduplicating case-insensitively while preserving initial case and original order.
- First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging.
- **Rationale**:
- Guarantees 100% deterministic and reproducible metadata output.
- Prevents corrupt or placeholder values from leaking into final editorial documents.
---
### Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822)
- **Decision**: Use Python's standard library `datetime.fromisoformat` and `email.utils.parsedate_to_datetime` / standard datetime parsing routines without heavy external dependencies.
- **Rationale**:
- All input dates observed from extractors follow ISO 8601 (e.g. `2026-08-20T00:36:33-03:00` or `2026-08-20T03:36:33Z`) or RFC 2822 (e.g. `Thu, 20 Aug 2026 00:36:33 -0300`).
- `datetime.fromisoformat()` in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'.
- `email.utils.parsedate_to_datetime()` standard library natively handles RFC 2822 dates.
- If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order.
- **Alternatives Considered**:
- *dateparser / python-dateutil*: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors.
---
### Decision 5: URL Validation & Media Filtering
- **Decision**:
- Validate all URLs with `urllib.parse.urlparse`: Scheme must be `http` or `https`, and `netloc` (hostname) must be non-empty.
- In Markdown body: Filter out images with relative URLs, empty URLs, or `data:` URIs.
- Deduplicate identical image URLs in the body, keeping only the first occurrence.
- **Rationale**:
- Prevents broken local references or bloated base64 data URIs in downstream pipelines.
---
### Decision 6: Duplicate Title Heading (H1) Stripping
- **Decision**:
- Check the first top-level ATX heading (`# ...`) in the converted body.
- If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace.
- Preserve all subsequent H1/H2/H3 headings in the body.
- **Rationale**:
- Many news articles embed the title in `<h1>` inside the article HTML. Since our Markdown schema places `# <Resolved Title>` at the top of the document, stripping the redundant body H1 avoids awkward repeated headings.
---
### Decision 7: Atomic File Writing & Error Resilience
- **Decision**:
- Write Markdown output to a temporary file in the same directory (`.<output_filename>.tmp`) and atomically replace the destination using `os.replace` (or `pathlib.Path.replace`).
- If any error or validation exception occurs during processing, clean up the temporary file and exit with code `1`, leaving any pre-existing target file untouched.
- **Rationale**:
- Guarantees transactional file operations in unattended automated batch pipelines.
---
## 3. Technology Stack & Dependencies
- **Runtime**: Python `>=3.10` (tested on 3.10, 3.11, 3.12)
- **New Dependency**: `markdownify>=0.13.0`
- **Standard Library Modules**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils`
- **Testing & Quality**: `pytest`, `ruff`, `mypy`