5.8 KiB
5.8 KiB
Research: Convert Article JSON to Markdown
Feature: 005-convert-json-markdown | Date: 2026-08-21
1. Executive Summary & Goals
This research addresses the design decisions for converting a single news article JSON with selected_extractor (trafilatura, newspaper4k, or readability) into a standardized, clean, human-readable Markdown file (.md) with deterministic metadata extraction and fallback rules.
2. Technical Decisions & Research Findings
Decision 1: HTML-to-Markdown Engine Selection
- Decision: Use
markdownifywith ATX heading style (heading_style=ATX). - Rationale:
markdownifyis a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown.- It natively converts
<h1>–<h6>to#–######(ATX style), parses tables, code blocks, lists, quotes, and inline styles (<b>,<i>,<a>,<img>). - Unlike broader conversion tools like Microsoft MarkItDown or Pandoc,
markdownifyhas zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions.
- Alternatives Considered:
- Microsoft MarkItDown: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion.
- Custom BeautifulSoup converter: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting).
Decision 2: Direct Markdown Handling for Trafilatura
- Decision: When
selected_extractor == "trafilatura", directly usetrafilatura.markdown(or fallback totrafilatura.text) without running HTML-to-Markdown conversion. - Rationale:
- Trafilatura natively emits high-quality Markdown in its extraction output.
- Plain text (
trafilatura.text) is already valid Markdown without special markup.
- Alternatives Considered:
- Converting Trafilatura's raw HTML: Inefficient and degrades Trafilatura's native document structural tree.
Decision 3: Metadata Normalization & Priority Resolution Pipeline
- Decision: Implement a pure Python deterministic metadata resolver supporting:
- Strict priority tables matching the PRD specification.
- Scalar normalization: HTML entity decoding (
html.unescape), whitespace trimming/collapsing, placeholder discarding (null,none,n/a,unknown,[no-author],no-author). - List normalization: Splitting on
;if string, trimming elements, filtering out URL-like authors (http://,https://,www.), deduplicating case-insensitively while preserving initial case and original order. - First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging.
- Rationale:
- Guarantees 100% deterministic and reproducible metadata output.
- Prevents corrupt or placeholder values from leaking into final editorial documents.
Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822)
- Decision: Use Python's standard library
datetime.fromisoformatandemail.utils.parsedate_to_datetime/ standard datetime parsing routines without heavy external dependencies. - Rationale:
- All input dates observed from extractors follow ISO 8601 (e.g.
2026-08-20T00:36:33-03:00or2026-08-20T03:36:33Z) or RFC 2822 (e.g.Thu, 20 Aug 2026 00:36:33 -0300). datetime.fromisoformat()in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'.email.utils.parsedate_to_datetime()standard library natively handles RFC 2822 dates.- If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order.
- All input dates observed from extractors follow ISO 8601 (e.g.
- Alternatives Considered:
- dateparser / python-dateutil: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors.
Decision 5: URL Validation & Media Filtering
- Decision:
- Validate all URLs with
urllib.parse.urlparse: Scheme must behttporhttps, andnetloc(hostname) must be non-empty. - In Markdown body: Filter out images with relative URLs, empty URLs, or
data:URIs. - Deduplicate identical image URLs in the body, keeping only the first occurrence.
- Validate all URLs with
- Rationale:
- Prevents broken local references or bloated base64 data URIs in downstream pipelines.
Decision 6: Duplicate Title Heading (H1) Stripping
- Decision:
- Check the first top-level ATX heading (
# ...) in the converted body. - If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace.
- Preserve all subsequent H1/H2/H3 headings in the body.
- Check the first top-level ATX heading (
- Rationale:
- Many news articles embed the title in
<h1>inside the article HTML. Since our Markdown schema places# <Resolved Title>at the top of the document, stripping the redundant body H1 avoids awkward repeated headings.
- Many news articles embed the title in
Decision 7: Atomic File Writing & Error Resilience
- Decision:
- Write Markdown output to a temporary file in the same directory (
.<output_filename>.tmp) and atomically replace the destination usingos.replace(orpathlib.Path.replace). - If any error or validation exception occurs during processing, clean up the temporary file and exit with code
1, leaving any pre-existing target file untouched.
- Write Markdown output to a temporary file in the same directory (
- Rationale:
- Guarantees transactional file operations in unattended automated batch pipelines.
3. Technology Stack & Dependencies
- Runtime: Python
>=3.10(tested on 3.10, 3.11, 3.12) - New Dependency:
markdownify>=0.13.0 - Standard Library Modules:
argparse,json,os,sys,pathlib,re,html,urllib.parse,datetime,email.utils - Testing & Quality:
pytest,ruff,mypy