Files
TextNLPClassifierApp/specs/005-convert-json-markdown/research.md
T

101 lines
5.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Research: Convert Article JSON to Markdown
**Feature**: `005-convert-json-markdown` | **Date**: 2026-08-21
## 1. Executive Summary & Goals
This research addresses the design decisions for converting a single news article JSON with `selected_extractor` (`trafilatura`, `newspaper4k`, or `readability`) into a standardized, clean, human-readable Markdown file (`.md`) with deterministic metadata extraction and fallback rules.
---
## 2. Technical Decisions & Research Findings
### Decision 1: HTML-to-Markdown Engine Selection
- **Decision**: Use [`markdownify`](https://github.com/matthewwithanm/python-markdownify) with ATX heading style (`heading_style=ATX`).
- **Rationale**:
- `markdownify` is a lightweight, battle-tested Python library focused exclusively on converting HTML trees to clean Markdown.
- It natively converts `<h1>`–`<h6>` to `#`–`######` (ATX style), parses tables, code blocks, lists, quotes, and inline styles (`<b>`, `<i>`, `<a>`, `<img>`).
- Unlike broader conversion tools like Microsoft MarkItDown or Pandoc, `markdownify` has zero external non-Python dependencies, low overhead, and avoids unnecessary multi-format abstractions.
- **Alternatives Considered**:
- *Microsoft MarkItDown*: Evaluated in PRD; rejected because it pulls broader dependencies (PDF, DOCX, audio, Azure AI) that exceed the scope of pure HTML-to-Markdown conversion.
- *Custom BeautifulSoup converter*: Unnecessary wheel reinvention; maintenance burden for complex HTML elements (nested lists, tables, inline formatting).
---
### Decision 2: Direct Markdown Handling for Trafilatura
- **Decision**: When `selected_extractor == "trafilatura"`, directly use `trafilatura.markdown` (or fallback to `trafilatura.text`) without running HTML-to-Markdown conversion.
- **Rationale**:
- Trafilatura natively emits high-quality Markdown in its extraction output.
- Plain text (`trafilatura.text`) is already valid Markdown without special markup.
- **Alternatives Considered**:
- *Converting Trafilatura's raw HTML*: Inefficient and degrades Trafilatura's native document structural tree.
---
### Decision 3: Metadata Normalization & Priority Resolution Pipeline
- **Decision**: Implement a pure Python deterministic metadata resolver supporting:
- Strict priority tables matching the PRD specification.
- Scalar normalization: HTML entity decoding (`html.unescape`), whitespace trimming/collapsing, placeholder discarding (`null`, `none`, `n/a`, `unknown`, `[no-author]`, `no-author`).
- List normalization: Splitting on `;` if string, trimming elements, filtering out URL-like authors (`http://`, `https://`, `www.`), deduplicating case-insensitively while preserving initial case and original order.
- First-valid-source selection: Pick the first source in priority order that yields a valid, non-empty candidate list without cross-source merging.
- **Rationale**:
- Guarantees 100% deterministic and reproducible metadata output.
- Prevents corrupt or placeholder values from leaking into final editorial documents.
---
### Decision 4: Date Parsing Strategy (ISO 8601 & RFC 2822)
- **Decision**: Use Python's standard library `datetime.fromisoformat` and `email.utils.parsedate_to_datetime` / standard datetime parsing routines without heavy external dependencies.
- **Rationale**:
- All input dates observed from extractors follow ISO 8601 (e.g. `2026-08-20T00:36:33-03:00` or `2026-08-20T03:36:33Z`) or RFC 2822 (e.g. `Thu, 20 Aug 2026 00:36:33 -0300`).
- `datetime.fromisoformat()` in Python 3.11+ handles full ISO 8601 with timezone offsets and 'Z'.
- `email.utils.parsedate_to_datetime()` standard library natively handles RFC 2822 dates.
- If a date cannot be parsed, the candidate is discarded and resolution advances to the next source in priority order.
- **Alternatives Considered**:
- *dateparser / python-dateutil*: Adds extra heavy dependency; unnecessary given the standardized datetime formats emitted by upstream extractors.
---
### Decision 5: URL Validation & Media Filtering
- **Decision**:
- Validate all URLs with `urllib.parse.urlparse`: Scheme must be `http` or `https`, and `netloc` (hostname) must be non-empty.
- In Markdown body: Filter out images with relative URLs, empty URLs, or `data:` URIs.
- Deduplicate identical image URLs in the body, keeping only the first occurrence.
- **Rationale**:
- Prevents broken local references or bloated base64 data URIs in downstream pipelines.
---
### Decision 6: Duplicate Title Heading (H1) Stripping
- **Decision**:
- Check the first top-level ATX heading (`# ...`) in the converted body.
- If its text matches the resolved article title (after HTML entity decoding, whitespace collapsing, and case-insensitive comparison), remove that H1 line and preceding/following whitespace.
- Preserve all subsequent H1/H2/H3 headings in the body.
- **Rationale**:
- Many news articles embed the title in `<h1>` inside the article HTML. Since our Markdown schema places `# <Resolved Title>` at the top of the document, stripping the redundant body H1 avoids awkward repeated headings.
---
### Decision 7: Atomic File Writing & Error Resilience
- **Decision**:
- Write Markdown output to a temporary file in the same directory (`.<output_filename>.tmp`) and atomically replace the destination using `os.replace` (or `pathlib.Path.replace`).
- If any error or validation exception occurs during processing, clean up the temporary file and exit with code `1`, leaving any pre-existing target file untouched.
- **Rationale**:
- Guarantees transactional file operations in unattended automated batch pipelines.
---
## 3. Technology Stack & Dependencies
- **Runtime**: Python `>=3.10` (tested on 3.10, 3.11, 3.12)
- **New Dependency**: `markdownify>=0.13.0`
- **Standard Library Modules**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils`
- **Testing & Quality**: `pytest`, `ruff`, `mypy`