Files
TextNLPClassifierApp/specs/005-convert-json-markdown/research.md
T

5.8 KiB
Raw Blame History

Research: Convert Article JSON to Markdown

Feature: 005-convert-json-markdown | Date: 2026-08-21

1. Executive Summary & Goals

This research addresses the design decisions for converting a single news article JSON with selected_extractor (trafilatura, newspaper4k, or readability) into a standardized, clean, human-readable Markdown file (.md) with deterministic metadata extraction and fallback rules.


2. Technical Decisions & Research Findings

Decision 1: HTML-to-Markdown Engine Selection

  • Decision: Use markdownify with ATX heading style (heading_style=ATX).
  • Rationale:
    • markdownify is a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown.
    • It natively converts <h1>–<h6> to #–###### (ATX style), parses tables, code blocks, lists, quotes, and inline styles (<b>, <i>, <a>, <img>).
    • Unlike broader conversion tools like Microsoft MarkItDown or Pandoc, markdownify has zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions.
  • Alternatives Considered:
    • Microsoft MarkItDown: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion.
    • Custom BeautifulSoup converter: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting).

Decision 2: Direct Markdown Handling for Trafilatura

  • Decision: When selected_extractor == "trafilatura", directly use trafilatura.markdown (or fallback to trafilatura.text) without running HTML-to-Markdown conversion.
  • Rationale:
    • Trafilatura natively emits high-quality Markdown in its extraction output.
    • Plain text (trafilatura.text) is already valid Markdown without special markup.
  • Alternatives Considered:
    • Converting Trafilatura's raw HTML: Inefficient and degrades Trafilatura's native document structural tree.

Decision 3: Metadata Normalization & Priority Resolution Pipeline

  • Decision: Implement a pure Python deterministic metadata resolver supporting:
    • Strict priority tables matching the PRD specification.
    • Scalar normalization: HTML entity decoding (html.unescape), whitespace trimming/collapsing, placeholder discarding (null, none, n/a, unknown, [no-author], no-author).
    • List normalization: Splitting on ; if string, trimming elements, filtering out URL-like authors (http://, https://, www.), deduplicating case-insensitively while preserving initial case and original order.
    • First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging.
  • Rationale:
    • Guarantees 100% deterministic and reproducible metadata output.
    • Prevents corrupt or placeholder values from leaking into final editorial documents.

Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822)

  • Decision: Use Python's standard library datetime.fromisoformat and email.utils.parsedate_to_datetime / standard datetime parsing routines without heavy external dependencies.
  • Rationale:
    • All input dates observed from extractors follow ISO 8601 (e.g. 2026-08-20T00:36:33-03:00 or 2026-08-20T03:36:33Z) or RFC 2822 (e.g. Thu, 20 Aug 2026 00:36:33 -0300).
    • datetime.fromisoformat() in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'.
    • email.utils.parsedate_to_datetime() standard library natively handles RFC 2822 dates.
    • If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order.
  • Alternatives Considered:
    • dateparser / python-dateutil: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors.

Decision 5: URL Validation & Media Filtering

  • Decision:
    • Validate all URLs with urllib.parse.urlparse: Scheme must be http or https, and netloc (hostname) must be non-empty.
    • In Markdown body: Filter out images with relative URLs, empty URLs, or data: URIs.
    • Deduplicate identical image URLs in the body, keeping only the first occurrence.
  • Rationale:
    • Prevents broken local references or bloated base64 data URIs in downstream pipelines.

Decision 6: Duplicate Title Heading (H1) Stripping

  • Decision:
    • Check the first top-level ATX heading (# ...) in the converted body.
    • If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace.
    • Preserve all subsequent H1/H2/H3 headings in the body.
  • Rationale:
    • Many news articles embed the title in <h1> inside the article HTML. Since our Markdown schema places # <Resolved Title> at the top of the document, stripping the redundant body H1 avoids awkward repeated headings.

Decision 7: Atomic File Writing & Error Resilience

  • Decision:
    • Write Markdown output to a temporary file in the same directory (.<output_filename>.tmp) and atomically replace the destination using os.replace (or pathlib.Path.replace).
    • If any error or validation exception occurs during processing, clean up the temporary file and exit with code 1, leaving any pre-existing target file untouched.
  • Rationale:
    • Guarantees transactional file operations in unattended automated batch pipelines.

3. Technology Stack & Dependencies

  • Runtime: Python >=3.10 (tested on 3.10, 3.11, 3.12)
  • New Dependency: markdownify>=0.13.0
  • Standard Library Modules: argparse, json, os, sys, pathlib, re, html, urllib.parse, datetime, email.utils
  • Testing & Quality: pytest, ruff, mypy