6.1 KiB
Implementation Plan: Convert Article JSON to Markdown
Branch: 005-convert-json-markdown | Date: 2026-08-21 | Spec: spec.md
Input: Feature specification from specs/005-convert-json-markdown/spec.md
Summary
Implement a standalone Python script and modular library component (scripts/convert_article_to_markdown.py and supporting functions) that reads a single news article JSON with a selected_extractor attribute, performs deterministic metadata resolution across extractor candidates according to strict priority hierarchies, converts HTML bodies to clean Markdown using markdownify, strips duplicate H1 title headings, sanitizes body image links, and writes the assembled Markdown document atomically.
Technical Context
Language/Version: Python >=3.10 (tested on 3.10, 3.11, 3.12)
Primary Dependencies: markdownify>=0.13.0
Standard Library: argparse, json, os, sys, pathlib, re, html, urllib.parse, datetime, email.utils
Storage: Local filesystem (JSON input, Markdown .md output)
Testing: pytest>=7.0.0 (Unit tests, CLI integration tests, exact byte comparison fixtures)
Quality Gates: ruff (linting/formatting), mypy (type checking), pytest
Target Platform: Cross-platform (Windows, Linux, macOS)
Project Type: CLI Script / Modular Data Pipeline Stage
Performance Goals: <200ms per article on standard hardware; 100% byte-for-byte deterministic output
Constraints: Fully offline / in-memory execution; no network calls; no LLMs/probabilistic algorithms; transactional atomic file writing
Constitution Check
GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.
| Principle | Requirement | Compliance Analysis | Status |
|---|---|---|---|
| I. Library-First | Self-contained, independently testable modular design | Core parsing, normalization, and assembly logic is modularized into testable pure functions. | ✅ PASS |
| II. CLI Interface | Clean CLI, text/file I/O, error reporting to stderr, standard exit codes (0, 1, 2) |
CLI exposes -i/--input and -o/--output, prints diagnostics to stderr, and handles errors cleanly. |
✅ PASS |
| III. Test-First | TDD mandatory: test fixtures, unit tests, and CLI tests before implementation | Comprehensive test suite planned covering all extractors, fallbacks, and edge cases. | ✅ PASS |
| IV. Integration Testing | CLI end-to-end integration and fixture contract tests | Golden fixtures for trafilatura, newspaper4k, and readability with exact Markdown match verification. |
✅ PASS |
| V. Simplicity & Observability | YAGNI, standard library where possible, markdownify for HTML conversion |
Lightweight dependencies, pure Python standard library for date/URL/normalization routines. | ✅ PASS |
Project Structure
Documentation (this feature)
specs/005-convert-json-markdown/
├── spec.md # Feature specification
├── plan.md # This file (/speckit-plan command output)
├── research.md # Technical research & decisions
├── data-model.md # Entities, normalization rules & priority matrix
├── quickstart.md # Quickstart & verification guide
├── checklists/
│ └── requirements.md # Quality checklist
└── contracts/
├── cli-contract.md # CLI interface definition
└── markdown-schema.md # Output Markdown schema contract
Source Code & Test Layout
TextNLPClassifierApp/
├── scripts/
│ └── convert_article_to_markdown.py # CLI entry point and conversion logic
├── tests/
│ ├── fixtures/
│ │ ├── markdown_conversion/ # Test fixtures (JSON inputs & expected MD outputs)
│ │ │ ├── valid_trafilatura.json
│ │ │ ├── valid_trafilatura.md
│ │ │ ├── valid_newspaper4k.json
│ │ │ ├── valid_newspaper4k.md
│ │ │ ├── valid_readability.json
│ │ │ ├── valid_readability.md
│ │ │ ├── batch_articles_invalid.json
│ │ │ └── missing_body_invalid.json
│ └── test_convert_article_to_markdown.py # Unit and integration test suite
├── requirements.txt # Updated with markdownify>=0.13.0
└── README.md # Documenting conversion script usage
Implementation Phases
Phase 0: Outline & Research (Completed)
- Resolved technical decisions in research.md.
- Confirmed
markdownifyas the HTML-to-Markdown engine and pure Python stdlib for dates/URLs.
Phase 1: Design & Contracts (Completed)
- Defined domain entities and priority resolution matrix in data-model.md.
- Defined CLI interface contract in contracts/cli-contract.md.
- Defined Markdown output document contract in contracts/markdown-schema.md.
- Created quickstart.md validation instructions.
Phase 2: Tasks & Implementation Breakdown (Next: /speckit-tasks)
- Update
requirements.txtto includemarkdownify>=0.13.0. - Build unit test fixtures for each extractor and error condition under
tests/fixtures/markdown_conversion/. - Implement metadata extraction, normalization, date parsing, and priority resolution functions.
- Implement HTML-to-Markdown conversion, duplicate H1 heading removal, and image URL sanitation.
- Implement Markdown document assembly and atomic file writing.
- Implement CLI argument parsing and error handling in
scripts/convert_article_to_markdown.py. - Write comprehensive test suite in
tests/test_convert_article_to_markdown.py. - Run linter (
ruff), type checker (mypy), and test suite (pytest). - Update
README.mdwith CLI documentation and pipeline examples.
Complexity Tracking
| Violation | Why Needed | Simpler Alternative Rejected Because |
|---|---|---|
| None | All principles satisfied without architectural violations. | N/A |