8.0 KiB
Tasks: Deterministic Article Content Selection
Branch: 004-deterministic-content-selection | Spec: spec.md | Plan: plan.md
Phase 1: Setup (Shared Infrastructure)
Purpose: Project initialization and test harness setup
- T001 Initialize script entrypoint and test suite structure in
scripts/select_article_extractor.pyandtests/test_select_article_extractor.py
Phase 2: Foundational (Data Structures & Normalization Engine)
Purpose: Core data models, text normalization, and shingle generation that all user stories depend upon
⚠️ CRITICAL: Must be completed before user story implementation begins
- T002 [P] Implement dataclasses and enumerations (
ExtractorName,CandidateStatus,ExtractorCandidate,ArticleSelectionResult,BatchProcessingResult) inscripts/select_article_extractor.py - T003 [P] Implement text normalization pipeline (
normalize_text, HTML entities unescape, HTML/Markdown tag strip, NFKC, lowercase, Unicode tokenization) inscripts/select_article_extractor.py - T004 Implement 5-token sliding window and short text shingle generator (
generate_shingles) inscripts/select_article_extractor.py - T005 Implement unit tests for normalization, tokenization, and shingle generation in
tests/test_select_article_extractor.py
Checkpoint: Foundation ready — text normalization and shingle generator fully operational and tested.
Phase 3: User Story 1 - Deterministic Selection with Text Consensus (Priority: P1) 🎯 MVP
Goal: Calculate consensus shingles (\ge 2 active extractors), compute Coverage, Support, and F_1 score, apply technical tie margin (\le 0.03), and select the best extractor based on agreement and conciseness.
Independent Test: Execute tests with synthetic and real articles where 2 or 3 extractors agree (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) and verify the winner matches expected score / tie-breaker.
Tests for User Story 1 🧪
NOTE: Write these tests FIRST, ensure they FAIL before implementation
- T006 [P] [US1] Write unit tests for consensus scoring, coverage/support F1 calculation, and technical tie-breaking (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) in
tests/test_select_article_extractor.py
Implementation for User Story 1
- T007 [US1] Implement consensus shingle builder and metric calculator (
calculate_consensus_metrics) inscripts/select_article_extractor.py - T008 [US1] Implement consensus decision selector with technical tie pool (
\le 0.03), smallest shingle count preference, and final priority fallback (select_with_consensus) inscripts/select_article_extractor.py
Checkpoint: User Story 1 (MVP) is fully functional and independently testable for all consensus scenarios.
Phase 4: User Story 2 - Resilient Decision Under Disagreement or Degradation (Priority: P2)
Goal: Ensure zero unassigned or ambiguous selections by handling zero-consensus articles (median of 3, max of 2, single), candidate degradation, all-unavailable extractors, and strict tie-breaking priority (newspaper4k > readability > trafilatura).
Independent Test: Execute tests for zero consensus, degraded errors, single usable candidate, and total extraction failure (CT-004, CT-005, CT-006, CT-007, CT-008).
Tests for User Story 2 🧪
NOTE: Write these tests FIRST, ensure they FAIL before implementation
- T009 [P] [US2] Write unit tests for zero-consensus, degraded candidate handling, single candidate, and total unavailability fallback (CT-004, CT-005, CT-006, CT-007, CT-008) in
tests/test_select_article_extractor.py
Implementation for User Story 2
- T010 [US2] Implement candidate status classifier (
classify_candidate_status) and active candidate set builder (form_active_set) inscripts/select_article_extractor.py - T011 [US2] Implement zero-consensus decision logic (median of 3, max of 2, single candidate, priority hierarchy) (
select_without_consensus) inscripts/select_article_extractor.py - T012 [US2] Implement single article selector orchestrator (
select_article_extractor) integrating Usable, Degraded, Consensus, and Non-Consensus decision branches inscripts/select_article_extractor.py
Checkpoint: User Stories 1 AND 2 are fully functional and handle 100% of single-article decision paths.
Phase 5: User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
Goal: Read JSON batch files containing articles, preserve all original fields and article order, recalculate existing selected_extractor values, validate schema, and write output atomically to <name>_selected.json.
Independent Test: Execute CLI and batch tests (CT-012, CT-013, CT-014), verify non-destructive field preservation, and run end-to-end processing on out/river_plate_extracted.json.
Tests for User Story 3 🧪
NOTE: Write these tests FIRST, ensure they FAIL before implementation
- T013 [P] [US3] Write integration tests for JSON schema validation, error handling, key preservation, recalculation of existing key, and atomic I/O (CT-012, CT-013, CT-014) in
tests/test_select_article_extractor.py
Implementation for User Story 3
- T014 [US3] Implement batch processor (
process_batch) and atomic file saver (atomic_save_json) inscripts/select_article_extractor.py - T015 [US3] Implement CLI interface (
main) withargparse, options (-o,--indent,--verbose), exit codes (0, 1, 2),stderrlogs, andstdoutJSON summary inscripts/select_article_extractor.py
Checkpoint: Full end-to-end batch processing operational and tested against contracts and real datasets.
Phase 6: Polish & Cross-Cutting Concerns
Purpose: Validation, performance checks, and documentation verification
- T016 [P] Execute quickstart validation scenarios on
out/river_plate_extracted.jsonperspecs/004-deterministic-content-selection/quickstart.md - T017 Run full test suite with coverage via
pytest tests/test_select_article_extractor.py -vensuring all 14 mandatory test cases (CT-001 to CT-014) pass
Dependencies & Execution Order
graph TD
T001[T001: Setup Harness] --> T002[T002: Data Models]
T001 --> T003[T003: Text Normalization]
T002 --> T004[T004: Shingle Generator]
T003 --> T004
T004 --> T005[T005: Foundation Tests]
T005 --> T006[T006: US1 Tests]
T006 --> T007[T007: US1 Consensus Metrics]
T007 --> T008[T008: US1 Consensus Selection]
T008 --> T009[T009: US2 Tests]
T009 --> T010[T010: US2 Candidate Classifier]
T010 --> T011[T011: US2 Zero-Consensus Logic]
T011 --> T012[T012: US2 Orchestrator]
T012 --> T013[T013: US3 Integration Tests]
T013 --> T014[T014: US3 Batch & Atomic I/O]
T014 --> T015[T015: US3 CLI Interface]
T015 --> T016[T016: Quickstart Validation]
T016 --> T017[T017: Full Test Suite CT-001..CT-014]
Parallel Opportunities
- Phase 2 (Foundations):
T002(Models) andT003(Normalization) can be developed in parallel. - Phase 3 (User Story 1):
T006(Unit tests) can be written in parallel with data model test setups. - Phase 4 (User Story 2):
T009(Unit tests) can be written in parallel with candidate state transition logic. - Phase 5 (User Story 3):
T013(Integration tests) can be developed alongside CLI option parser definitions.
Implementation Strategy
MVP First (Phases 1, 2 & 3)
- Complete Setup and Foundations (
T001-T005). - Implement User Story 1 (
T006-T008): text consensus and scoring engine. - Validate MVP: test consensus-based decisions on multi-extractor articles.
Incremental Delivery (Phases 4, 5 & 6)
- Add User Story 2 (
T009-T012): resilient fallbacks, zero consensus, degraded candidate evaluation. - Add User Story 3 (
T013-T015): non-destructive batch JSON processing, atomic save, and CLI tool. - Polish & Final Verification (
T016-T017): run quickstart validation and full test suite covering CT-001 to CT-014.