Files
TextNLPClassifierApp/specs/005-convert-json-markdown/plan.md
T

6.1 KiB

Implementation Plan: Convert Article JSON to Markdown

Branch: 005-convert-json-markdown | Date: 2026-08-21 | Spec: spec.md

Input: Feature specification from specs/005-convert-json-markdown/spec.md


Summary

Implement a standalone Python script and modular library component (scripts/convert_article_to_markdown.py and supporting functions) that reads a single news article JSON with a selected_extractor attribute, performs deterministic metadata resolution across extractor candidates according to strict priority hierarchies, converts HTML bodies to clean Markdown using markdownify, strips duplicate H1 title headings, sanitizes body image links, and writes the assembled Markdown document atomically.


Technical Context

Language/Version: Python >=3.10 (tested on 3.10, 3.11, 3.12)
Primary Dependencies: markdownify>=0.13.0
Standard Library: argparse, json, os, sys, pathlib, re, html, urllib.parse, datetime, email.utils
Storage: Local filesystem (JSON input, Markdown .md output)
Testing: pytest>=7.0.0 (Unit tests, CLI integration tests, exact byte comparison fixtures)
Quality Gates: ruff (linting/formatting), mypy (type checking), pytest
Target Platform: Cross-platform (Windows, Linux, macOS)
Project Type: CLI Script / Modular Data Pipeline Stage
Performance Goals: <200ms per article on standard hardware; 100% byte-for-byte deterministic output
Constraints: Fully offline / in-memory execution; no network calls; no LLMs/probabilistic algorithms; transactional atomic file writing


Constitution Check

GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.

Principle Requirement Compliance Analysis Status
I. Library-First Self-contained, independently testable modular design Core parsing, normalization, and assembly logic is modularized into testable pure functions. ✅ PASS
II. CLI Interface Clean CLI, text/file I/O, error reporting to stderr, standard exit codes (0, 1, 2) CLI exposes -i/--input and -o/--output, prints diagnostics to stderr, and handles errors cleanly. ✅ PASS
III. Test-First TDD mandatory: test fixtures, unit tests, and CLI tests before implementation Comprehensive test suite planned covering all extractors, fallbacks, and edge cases. ✅ PASS
IV. Integration Testing CLI end-to-end integration and fixture contract tests Golden fixtures for trafilatura, newspaper4k, and readability with exact Markdown match verification. ✅ PASS
V. Simplicity & Observability YAGNI, standard library where possible, markdownify for HTML conversion Lightweight dependencies, pure Python standard library for date/URL/normalization routines. ✅ PASS

Project Structure

Documentation (this feature)

specs/005-convert-json-markdown/
├── spec.md              # Feature specification
├── plan.md              # This file (/speckit-plan command output)
├── research.md          # Technical research & decisions
├── data-model.md        # Entities, normalization rules & priority matrix
├── quickstart.md        # Quickstart & verification guide
├── checklists/
│   └── requirements.md  # Quality checklist
└── contracts/
    ├── cli-contract.md      # CLI interface definition
    └── markdown-schema.md   # Output Markdown schema contract

Source Code & Test Layout

TextNLPClassifierApp/
├── scripts/
│   └── convert_article_to_markdown.py   # CLI entry point and conversion logic
├── tests/
│   ├── fixtures/
│   │   ├── markdown_conversion/         # Test fixtures (JSON inputs & expected MD outputs)
│   │   │   ├── valid_trafilatura.json
│   │   │   ├── valid_trafilatura.md
│   │   │   ├── valid_newspaper4k.json
│   │   │   ├── valid_newspaper4k.md
│   │   │   ├── valid_readability.json
│   │   │   ├── valid_readability.md
│   │   │   ├── batch_articles_invalid.json
│   │   │   └── missing_body_invalid.json
│   └── test_convert_article_to_markdown.py  # Unit and integration test suite
├── requirements.txt                     # Updated with markdownify>=0.13.0
└── README.md                            # Documenting conversion script usage

Implementation Phases

Phase 0: Outline & Research (Completed)

  • Resolved technical decisions in research.md.
  • Confirmed markdownify as the HTML-to-Markdown engine and pure Python stdlib for dates/URLs.

Phase 1: Design & Contracts (Completed)

Phase 2: Tasks & Implementation Breakdown (Next: /speckit-tasks)

  1. Update requirements.txt to include markdownify>=0.13.0.
  2. Build unit test fixtures for each extractor and error condition under tests/fixtures/markdown_conversion/.
  3. Implement metadata extraction, normalization, date parsing, and priority resolution functions.
  4. Implement HTML-to-Markdown conversion, duplicate H1 heading removal, and image URL sanitation.
  5. Implement Markdown document assembly and atomic file writing.
  6. Implement CLI argument parsing and error handling in scripts/convert_article_to_markdown.py.
  7. Write comprehensive test suite in tests/test_convert_article_to_markdown.py.
  8. Run linter (ruff), type checker (mypy), and test suite (pytest).
  9. Update README.md with CLI documentation and pipeline examples.

Complexity Tracking

Violation Why Needed Simpler Alternative Rejected Because
None All principles satisfied without architectural violations. N/A