# Implementation Plan: Convert Article JSON to Markdown **Branch**: `005-convert-json-markdown` | **Date**: 2026-08-21 | **Spec**: [spec.md](spec.md) **Input**: Feature specification from `specs/005-convert-json-markdown/spec.md` --- ## Summary Implement a standalone Python script and modular library component (`scripts/convert_article_to_markdown.py` and supporting functions) that reads a single news article JSON with a `selected_extractor` attribute, performs deterministic metadata resolution across extractor candidates according to strict priority hierarchies, converts HTML bodies to clean Markdown using `markdownify`, strips duplicate H1 title headings, sanitizes body image links, and writes the assembled Markdown document atomically. --- ## Technical Context **Language/Version**: Python `>=3.10` (tested on 3.10, 3.11, 3.12) **Primary Dependencies**: `markdownify>=0.13.0` **Standard Library**: `argparse`, `json`, `os`, `sys`, `pathlib`, `re`, `html`, `urllib.parse`, `datetime`, `email.utils` **Storage**: Local filesystem (JSON input, Markdown `.md` output) **Testing**: `pytest>=7.0.0` (Unit tests, CLI integration tests, exact byte comparison fixtures) **Quality Gates**: `ruff` (linting/formatting), `mypy` (type checking), `pytest` **Target Platform**: Cross-platform (Windows, Linux, macOS) **Project Type**: CLI Script / Modular Data Pipeline Stage **Performance Goals**: `<200ms` per article on standard hardware; 100% byte-for-byte deterministic output **Constraints**: Fully offline / in-memory execution; no network calls; no LLMs/probabilistic algorithms; transactional atomic file writing --- ## Constitution Check *GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.* | Principle | Requirement | Compliance Analysis | Status | |---|---|---|:---:| | **I. Library-First** | Self-contained, independently testable modular design | Core parsing, normalization, and assembly logic is modularized into testable pure functions. | ✅ PASS | | **II. CLI Interface** | Clean CLI, text/file I/O, error reporting to `stderr`, standard exit codes (0, 1, 2) | CLI exposes `-i/--input` and `-o/--output`, prints diagnostics to `stderr`, and handles errors cleanly. | ✅ PASS | | **III. Test-First** | TDD mandatory: test fixtures, unit tests, and CLI tests before implementation | Comprehensive test suite planned covering all extractors, fallbacks, and edge cases. | ✅ PASS | | **IV. Integration Testing** | CLI end-to-end integration and fixture contract tests | Golden fixtures for `trafilatura`, `newspaper4k`, and `readability` with exact Markdown match verification. | ✅ PASS | | **V. Simplicity & Observability** | YAGNI, standard library where possible, `markdownify` for HTML conversion | Lightweight dependencies, pure Python standard library for date/URL/normalization routines. | ✅ PASS | --- ## Project Structure ### Documentation (this feature) ```text specs/005-convert-json-markdown/ ├── spec.md # Feature specification ├── plan.md # This file (/speckit-plan command output) ├── research.md # Technical research & decisions ├── data-model.md # Entities, normalization rules & priority matrix ├── quickstart.md # Quickstart & verification guide ├── checklists/ │ └── requirements.md # Quality checklist └── contracts/ ├── cli-contract.md # CLI interface definition └── markdown-schema.md # Output Markdown schema contract ``` ### Source Code & Test Layout ```text TextNLPClassifierApp/ ├── scripts/ │ └── convert_article_to_markdown.py # CLI entry point and conversion logic ├── tests/ │ ├── fixtures/ │ │ ├── markdown_conversion/ # Test fixtures (JSON inputs & expected MD outputs) │ │ │ ├── valid_trafilatura.json │ │ │ ├── valid_trafilatura.md │ │ │ ├── valid_newspaper4k.json │ │ │ ├── valid_newspaper4k.md │ │ │ ├── valid_readability.json │ │ │ ├── valid_readability.md │ │ │ ├── batch_articles_invalid.json │ │ │ └── missing_body_invalid.json │ └── test_convert_article_to_markdown.py # Unit and integration test suite ├── requirements.txt # Updated with markdownify>=0.13.0 └── README.md # Documenting conversion script usage ``` --- ## Implementation Phases ### Phase 0: Outline & Research *(Completed)* - Resolved technical decisions in [research.md](research.md). - Confirmed `markdownify` as the HTML-to-Markdown engine and pure Python stdlib for dates/URLs. ### Phase 1: Design & Contracts *(Completed)* - Defined domain entities and priority resolution matrix in [data-model.md](data-model.md). - Defined CLI interface contract in [contracts/cli-contract.md](contracts/cli-contract.md). - Defined Markdown output document contract in [contracts/markdown-schema.md](contracts/markdown-schema.md). - Created [quickstart.md](quickstart.md) validation instructions. ### Phase 2: Tasks & Implementation Breakdown *(Next: `/speckit-tasks`)* 1. Update `requirements.txt` to include `markdownify>=0.13.0`. 2. Build unit test fixtures for each extractor and error condition under `tests/fixtures/markdown_conversion/`. 3. Implement metadata extraction, normalization, date parsing, and priority resolution functions. 4. Implement HTML-to-Markdown conversion, duplicate H1 heading removal, and image URL sanitation. 5. Implement Markdown document assembly and atomic file writing. 6. Implement CLI argument parsing and error handling in `scripts/convert_article_to_markdown.py`. 7. Write comprehensive test suite in `tests/test_convert_article_to_markdown.py`. 8. Run linter (`ruff`), type checker (`mypy`), and test suite (`pytest`). 9. Update `README.md` with CLI documentation and pipeline examples. --- ## Complexity Tracking | Violation | Why Needed | Simpler Alternative Rejected Because | |---|---|---| | *None* | All principles satisfied without architectural violations. | N/A |