8.5 KiB
Tasks: Convert Article JSON to Markdown
Branch: 005-convert-json-markdown | Feature: Convert Article JSON to Markdown
Spec: spec.md | Plan: plan.md | Data Model: data-model.md
Phase 1: Setup (Shared Infrastructure)
Purpose: Project initialization, dependency management, and test fixture directory setup
- T001 Add
markdownify>=0.13.0torequirements.txt - T002 [P] Create fixtures directory structure in
tests/fixtures/markdown_conversion/
Phase 2: Foundational (Blocking Prerequisites)
Purpose: Golden fixtures, negative test fixtures, and baseline test infrastructure that MUST be complete before user stories begin
- T003 [P] Create Golden Test Fixtures for Trafilatura in
tests/fixtures/markdown_conversion/valid_trafilatura.jsonandtests/fixtures/markdown_conversion/valid_trafilatura.md - T004 [P] Create Golden Test Fixtures for Newspaper4k in
tests/fixtures/markdown_conversion/valid_newspaper4k.jsonandtests/fixtures/markdown_conversion/valid_newspaper4k.md - T005 [P] Create Golden Test Fixtures for Readability in
tests/fixtures/markdown_conversion/valid_readability.jsonandtests/fixtures/markdown_conversion/valid_readability.md - T006 [P] Create Negative Test Fixtures (
batch_articles_invalid.json,corrupt_json_invalid.json,missing_extractor_invalid.json,missing_body_invalid.json,missing_title_invalid.json,invalid_url_invalid.json) intests/fixtures/markdown_conversion/
Checkpoint: Foundational test fixtures ready. User story implementation can begin in strict TDD order.
Phase 3: User Story 1 - Single Article JSON to Clean Markdown Conversion (Priority: P1) 🎯 MVP
Goal: Convert a single article JSON into a clean Markdown document using exclusively the selected extractor (trafilatura, newspaper4k, readability), supporting direct Markdown reuse and HTML-to-Markdown conversion with intra-extractor fallback.
Independent Test: Provide single-article JSON fixtures for each extractor and verify that the resulting Markdown body text matches the selected extractor's content without cross-extractor leakage.
Tests for User Story 1 (TDD) ⚠️
NOTE: Write these tests FIRST, ensure they FAIL before implementing
- T007 [P] [US1] Write unit and integration tests for extractor body resolution, strict extractor isolation, and HTML-to-Markdown conversion in
tests/test_convert_article_to_markdown.py
Implementation for User Story 1
- T008 [US1] Implement body resolution and strict extractor isolation logic (prohibiting cross-extractor fallback) in
scripts/convert_article_to_markdown.py - T009 [US1] Implement HTML-to-Markdown conversion using
markdownify(ATX headings) and intra-extractor fallback (markdown/html→text) inscripts/convert_article_to_markdown.py
Checkpoint: At this point, User Story 1 is fully functional and testable independently (MVP ready).
Phase 4: User Story 2 - Deterministic Metadata Resolution and Fallback (Priority: P2)
Goal: Resolve mandatory and optional metadata across all extractor candidates and input metadata according to strict priority hierarchies, normalizing strings, lists, dates, and URLs, stripping duplicate H1 headers, and formatting the final Markdown structure.
Independent Test: Pass articles with missing metadata in the selected extractor but present in secondary sources; verify that resolved metadata strictly follows priority chains, discards placeholders, sanitizes dates/URLs/images, and formats output accurately.
Tests for User Story 2 (TDD) ⚠️
NOTE: Write these tests FIRST, ensure they FAIL before implementing
- T010 [P] [US2] Write unit tests for metadata priority chains, scalar normalization, list deduplication, date parsing (ISO 8601 / RFC 2822), URL validation, duplicate H1 heading removal, and body image sanitation in
tests/test_convert_article_to_markdown.py
Implementation for User Story 2
- T011 [US2] Implement scalar and list normalizers (HTML unescape, whitespace collapsing, placeholder filtering, author URL filtering, case-insensitive deduplication) in
scripts/convert_article_to_markdown.py - T012 [US2] Implement ISO 8601 and RFC 2822 date parser (with timezone preservation and
YYYY-MM-DDdate-only output) and URL validator inscripts/convert_article_to_markdown.py - T013 [US2] Implement deterministic metadata priority resolution matrix in
scripts/convert_article_to_markdown.py - T014 [US2] Implement duplicate H1 title heading stripper and body image link sanitizer (removing relative/
data:/empty URLs and deduplicating repeats) inscripts/convert_article_to_markdown.py - T015 [US2] Implement final Markdown document layout assembler (Header, Subtitle, Metadata key-values, Top Image, Separator
---, Body) with LF line endings inscripts/convert_article_to_markdown.py
Checkpoint: At this point, User Stories 1 and 2 work together and pass all unit/integration tests.
Phase 5: User Story 3 - CLI Usability, Validation, and Atomic Output (Priority: P3)
Goal: Provide a production-ready CLI interface with -i/--input and -o/--output flags, standard exit codes (0, 1, 2), error logging to stderr, and atomic file replacement with transactional error cleanup.
Independent Test: Execute the CLI against valid JSON, invalid JSON, batch articles array JSON, and missing files; verify exit codes, stderr diagnostics, and target file integrity.
Tests for User Story 3 (TDD) ⚠️
NOTE: Write these tests FIRST, ensure they FAIL before implementing
- T016 [P] [US3] Write CLI integration tests covering
-i/-oflags, exit codes (0,1,2), stderr logging, rejection of batcharticlesJSON, and atomic replacement rollback on failure intests/test_convert_article_to_markdown.py
Implementation for User Story 3
- T017 [US3] Implement CLI argument parsing, input JSON validation (object root, rejection of
articleskey,selected_extractorcheck), and diagnostic logging tostderrinscripts/convert_article_to_markdown.py - T018 [US3] Implement transactional atomic file writing (
.<output>.tmp+os.replace) with complete cleanup on failure inscripts/convert_article_to_markdown.py
Checkpoint: All user stories are fully implemented, resilient, and verified.
Phase 6: Polish & Cross-Cutting Concerns
Purpose: Quality gate verification, Golden Fixture byte-for-byte validation, and documentation
- T019 [P] Run full test suite (
pytest tests/test_convert_article_to_markdown.py -v) and verify 100% exact Golden Fixtures matching - T020 [P] Run code quality and type checks (
ruff check scripts/ src/ tests/andmypy scripts/convert_article_to_markdown.py) - T021 Update
README.mdwith conversion CLI documentation, pipeline usage examples, and argument references - T022 Run quickstart validation scenarios from
specs/005-convert-json-markdown/quickstart.md
Dependencies & Execution Order
Phase Dependencies
- Setup (Phase 1): No dependencies — start immediately.
- Foundational (Phase 2): Depends on Setup (Phase 1) — BLOCKS all user stories.
- User Story 1 (Phase 3): Depends on Foundational (Phase 2) — Core MVP.
- User Story 2 (Phase 4): Depends on User Story 1 (Phase 3).
- User Story 3 (Phase 5): Depends on User Story 2 (Phase 4).
- Polish (Phase 6): Depends on all user stories (Phases 3–5) being complete.
Parallel Opportunities
- Phase 1:
T002[P] can run in parallel withT001. - Phase 2:
T003[P],T004[P],T005[P], andT006[P] can all be created in parallel. - Phase 3:
T007[P] (tests) can be created in parallel with fixture setup. - Phase 4:
T010[P] (tests) can be written before implementation. - Phase 5:
T016[P] (CLI tests) can be written before CLI wiring. - Phase 6:
T019[P] (pytest) andT020[P] (ruff/mypy) can run in parallel.
Implementation Strategy
MVP First (User Story 1 Only)
- Complete Phase 1 (Setup) and Phase 2 (Foundational Fixtures).
- Complete Phase 3 (User Story 1: TDD tests
T007→ ImplementationT008,T009). - Validate independent execution of User Story 1.
Incremental Delivery
- Foundation Ready → MVP (US1: Body Conversion).
- Deliver US2 (Deterministic Metadata & Sanitization).
- Deliver US3 (CLI, Error Handling & Atomic Output).
- Run Polish & Quality Gates (US1 + US2 + US3 verified with 100% tests passing).