Files
TextNLPClassifierApp/specs/004-deterministic-content-selection/tasks.md
T

8.0 KiB

Tasks: Deterministic Article Content Selection

Branch: 004-deterministic-content-selection | Spec: spec.md | Plan: plan.md


Phase 1: Setup (Shared Infrastructure)

Purpose: Project initialization and test harness setup

  • T001 Initialize script entrypoint and test suite structure in scripts/select_article_extractor.py and tests/test_select_article_extractor.py

Phase 2: Foundational (Data Structures & Normalization Engine)

Purpose: Core data models, text normalization, and shingle generation that all user stories depend upon

⚠️ CRITICAL: Must be completed before user story implementation begins

  • T002 [P] Implement dataclasses and enumerations (ExtractorName, CandidateStatus, ExtractorCandidate, ArticleSelectionResult, BatchProcessingResult) in scripts/select_article_extractor.py
  • T003 [P] Implement text normalization pipeline (normalize_text, HTML entities unescape, HTML/Markdown tag strip, NFKC, lowercase, Unicode tokenization) in scripts/select_article_extractor.py
  • T004 Implement 5-token sliding window and short text shingle generator (generate_shingles) in scripts/select_article_extractor.py
  • T005 Implement unit tests for normalization, tokenization, and shingle generation in tests/test_select_article_extractor.py

Checkpoint: Foundation ready — text normalization and shingle generator fully operational and tested.


Phase 3: User Story 1 - Deterministic Selection with Text Consensus (Priority: P1) 🎯 MVP

Goal: Calculate consensus shingles (\ge 2 active extractors), compute Coverage, Support, and F_1 score, apply technical tie margin (\le 0.03), and select the best extractor based on agreement and conciseness.

Independent Test: Execute tests with synthetic and real articles where 2 or 3 extractors agree (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) and verify the winner matches expected score / tie-breaker.

Tests for User Story 1 🧪

NOTE: Write these tests FIRST, ensure they FAIL before implementation

  • T006 [P] [US1] Write unit tests for consensus scoring, coverage/support F1 calculation, and technical tie-breaking (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) in tests/test_select_article_extractor.py

Implementation for User Story 1

  • T007 [US1] Implement consensus shingle builder and metric calculator (calculate_consensus_metrics) in scripts/select_article_extractor.py
  • T008 [US1] Implement consensus decision selector with technical tie pool (\le 0.03), smallest shingle count preference, and final priority fallback (select_with_consensus) in scripts/select_article_extractor.py

Checkpoint: User Story 1 (MVP) is fully functional and independently testable for all consensus scenarios.


Phase 4: User Story 2 - Resilient Decision Under Disagreement or Degradation (Priority: P2)

Goal: Ensure zero unassigned or ambiguous selections by handling zero-consensus articles (median of 3, max of 2, single), candidate degradation, all-unavailable extractors, and strict tie-breaking priority (newspaper4k > readability > trafilatura).

Independent Test: Execute tests for zero consensus, degraded errors, single usable candidate, and total extraction failure (CT-004, CT-005, CT-006, CT-007, CT-008).

Tests for User Story 2 🧪

NOTE: Write these tests FIRST, ensure they FAIL before implementation

  • T009 [P] [US2] Write unit tests for zero-consensus, degraded candidate handling, single candidate, and total unavailability fallback (CT-004, CT-005, CT-006, CT-007, CT-008) in tests/test_select_article_extractor.py

Implementation for User Story 2

  • T010 [US2] Implement candidate status classifier (classify_candidate_status) and active candidate set builder (form_active_set) in scripts/select_article_extractor.py
  • T011 [US2] Implement zero-consensus decision logic (median of 3, max of 2, single candidate, priority hierarchy) (select_without_consensus) in scripts/select_article_extractor.py
  • T012 [US2] Implement single article selector orchestrator (select_article_extractor) integrating Usable, Degraded, Consensus, and Non-Consensus decision branches in scripts/select_article_extractor.py

Checkpoint: User Stories 1 AND 2 are fully functional and handle 100% of single-article decision paths.


Phase 5: User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)

Goal: Read JSON batch files containing articles, preserve all original fields and article order, recalculate existing selected_extractor values, validate schema, and write output atomically to <name>_selected.json.

Independent Test: Execute CLI and batch tests (CT-012, CT-013, CT-014), verify non-destructive field preservation, and run end-to-end processing on out/river_plate_extracted.json.

Tests for User Story 3 🧪

NOTE: Write these tests FIRST, ensure they FAIL before implementation

  • T013 [P] [US3] Write integration tests for JSON schema validation, error handling, key preservation, recalculation of existing key, and atomic I/O (CT-012, CT-013, CT-014) in tests/test_select_article_extractor.py

Implementation for User Story 3

  • T014 [US3] Implement batch processor (process_batch) and atomic file saver (atomic_save_json) in scripts/select_article_extractor.py
  • T015 [US3] Implement CLI interface (main) with argparse, options (-o, --indent, --verbose), exit codes (0, 1, 2), stderr logs, and stdout JSON summary in scripts/select_article_extractor.py

Checkpoint: Full end-to-end batch processing operational and tested against contracts and real datasets.


Phase 6: Polish & Cross-Cutting Concerns

Purpose: Validation, performance checks, and documentation verification

  • T016 [P] Execute quickstart validation scenarios on out/river_plate_extracted.json per specs/004-deterministic-content-selection/quickstart.md
  • T017 Run full test suite with coverage via pytest tests/test_select_article_extractor.py -v ensuring all 14 mandatory test cases (CT-001 to CT-014) pass

Dependencies & Execution Order

graph TD
    T001[T001: Setup Harness] --> T002[T002: Data Models]
    T001 --> T003[T003: Text Normalization]
    T002 --> T004[T004: Shingle Generator]
    T003 --> T004
    T004 --> T005[T005: Foundation Tests]
    
    T005 --> T006[T006: US1 Tests]
    T006 --> T007[T007: US1 Consensus Metrics]
    T007 --> T008[T008: US1 Consensus Selection]
    
    T008 --> T009[T009: US2 Tests]
    T009 --> T010[T010: US2 Candidate Classifier]
    T010 --> T011[T011: US2 Zero-Consensus Logic]
    T011 --> T012[T012: US2 Orchestrator]
    
    T012 --> T013[T013: US3 Integration Tests]
    T013 --> T014[T014: US3 Batch & Atomic I/O]
    T014 --> T015[T015: US3 CLI Interface]
    
    T015 --> T016[T016: Quickstart Validation]
    T016 --> T017[T017: Full Test Suite CT-001..CT-014]

Parallel Opportunities

  • Phase 2 (Foundations): T002 (Models) and T003 (Normalization) can be developed in parallel.
  • Phase 3 (User Story 1): T006 (Unit tests) can be written in parallel with data model test setups.
  • Phase 4 (User Story 2): T009 (Unit tests) can be written in parallel with candidate state transition logic.
  • Phase 5 (User Story 3): T013 (Integration tests) can be developed alongside CLI option parser definitions.

Implementation Strategy

MVP First (Phases 1, 2 & 3)

  1. Complete Setup and Foundations (T001 - T005).
  2. Implement User Story 1 (T006 - T008): text consensus and scoring engine.
  3. Validate MVP: test consensus-based decisions on multi-extractor articles.

Incremental Delivery (Phases 4, 5 & 6)

  1. Add User Story 2 (T009 - T012): resilient fallbacks, zero consensus, degraded candidate evaluation.
  2. Add User Story 3 (T013 - T015): non-destructive batch JSON processing, atomic save, and CLI tool.
  3. Polish & Final Verification (T016 - T017): run quickstart validation and full test suite covering CT-001 to CT-014.