# Research: Convert Article JSON to Markdown **Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21 ## 1. Executive Summary & Goals This research addresses the design decisions for converting a single news article JSON with `selected_extractor` (`trafilatura`, `newspaper4k`, or `readability`) into a standardized, clean, human-readable Markdown file (`.md`) with deterministic metadata extraction and fallback rules. --- ## 2. Technical Decisions & Research Findings ### Decision 1: HTML-to-Markdown Engine Selection - **Decision**: Use [`markdownify`](https://github.com/matthewwithanm/python-markdownify) with ATX heading style (`heading_style=ATX`). - **Rationale**: - `markdownify` is a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown. - It natively converts `

`–`

` to `#`–`######` (ATX style), parses tables, code blocks, lists, quotes, and inline styles (``, ``, ``, ``). - Unlike broader conversion tools like Microsoft MarkItDown or Pandoc, `markdownify` has zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions. - **Alternatives Considered**: - *Microsoft MarkItDown*: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion. - *Custom BeautifulSoup converter*: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting). --- ### Decision 2: Direct Markdown Handling for Trafilatura - **Decision**: When `selected_extractor == "trafilatura"`, directly use `trafilatura.markdown` (or fallback to `trafilatura.text`) without running HTML-to-Markdown conversion. - **Rationale**: - Trafilatura natively emits high-quality Markdown in its extraction output. - Plain text (`trafilatura.text`) is already valid Markdown without special markup. - **Alternatives Considered**: - *Converting Trafilatura's raw HTML*: Inefficient and degrades Trafilatura's native document structural tree. --- ### Decision 3: Metadata Normalization & Priority Resolution Pipeline - **Decision**: Implement a pure Python deterministic metadata resolver supporting: - Strict priority tables matching the PRD specification. - Scalar normalization: HTML entity decoding (`html.unescape`), whitespace trimming/collapsing, placeholder discarding (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`). - List normalization: Splitting on `;` if string, trimming elements, filtering out URL-like authors (`http://`, `https://`, `www.`), deduplicating case-insensitively while preserving initial case and original order. - First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging. - **Rationale**: - Guarantees 100% deterministic and reproducible metadata output. - Prevents corrupt or placeholder values from leaking into final editorial documents. --- ### Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822) - **Decision**: Use Python's standard library `datetime.fromisoformat` and `email.utils.parsedate_to_datetime` / standard datetime parsing routines without heavy external dependencies. - **Rationale**: - All input dates observed from extractors follow ISO 8601 (e.g. `2026-08-20T00:36:33-03:00` or `2026-08-20T03:36:33Z`) or RFC 2822 (e.g. `Thu, 20 Aug 2026 00:36:33 -0300`). - `datetime.fromisoformat()` in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'. - `email.utils.parsedate_to_datetime()` standard library natively handles RFC 2822 dates. - If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order. - **Alternatives Considered**: - *dateparser / python-dateutil*: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors. --- ### Decision 5: URL Validation & Media Filtering - **Decision**: - Validate all URLs with `urllib.parse.urlparse`: Scheme must be `http` or `https`, and `netloc` (hostname) must be non-empty. - In Markdown body: Filter out images with relative URLs, empty URLs, or `data:` URIs. - Deduplicate identical image URLs in the body, keeping only the first occurrence. - **Rationale**: - Prevents broken local references or bloated base64 data URIs in downstream pipelines. --- ### Decision 6: Duplicate Title Heading (H1) Stripping - **Decision**: - Check the first top-level ATX heading (`# ...`) in the converted body. - If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace. - Preserve all subsequent H1/H2/H3 headings in the body. - **Rationale**: - Many news articles embed the title in `

` inside the article HTML. Since our Markdown schema places `# ` at the top of the document, stripping the redundant body H1 avoids awkward repeated headings. --- ### Decision 7: Atomic File Writing & Error Resilience - **Decision**: - Write Markdown output to a temporary file in the same directory (`..tmp`) and atomically replace the destination using `os.replace` (or `pathlib.Path.replace`). - If any error or validation exception occurs during processing, clean up the temporary file and exit with code `1`, leaving any pre-existing target file untouched. - **Rationale**: - Guarantees transactional file operations in unattended automated batch pipelines. --- ## 3. Technology Stack & Dependencies - **Runtime**: Python `>=3.10` (tested on 3.10, 3.11, 3.12) - **New Dependency**: `markdownify>=0.13.0` - **Standard Library Modules**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils` - **Testing & Quality**: `pytest`, `ruff`, `mypy`