feat: add deterministic content extractor selector engine with F1 consensus

This commit is contained in:
2026-08-20 22:09:43 -03:00
parent 6a45368cb0
commit ff7a50e0eb
46 changed files with 18503 additions and 2813 deletions
@@ -0,0 +1,148 @@
# Tasks: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Spec**: [spec.md](spec.md) | **Plan**: [plan.md](plan.md)
---
## Phase 1: Setup (Shared Infrastructure)
**Purpose**: Project initialization and test harness setup
- [X] T001 Initialize script entrypoint and test suite structure in `scripts/select_article_extractor.py` and `tests/test_select_article_extractor.py`
---
## Phase 2: Foundational (Data Structures & Normalization Engine)
**Purpose**: Core data models, text normalization, and shingle generation that all user stories depend upon
**⚠️ CRITICAL**: Must be completed before user story implementation begins
- [X] T002 [P] Implement dataclasses and enumerations (`ExtractorName`, `CandidateStatus`, `ExtractorCandidate`, `ArticleSelectionResult`, `BatchProcessingResult`) in `scripts/select_article_extractor.py`
- [X] T003 [P] Implement text normalization pipeline (`normalize_text`, HTML entities unescape, HTML/Markdown tag strip, NFKC, lowercase, Unicode tokenization) in `scripts/select_article_extractor.py`
- [X] T004 Implement 5-token sliding window and short text shingle generator (`generate_shingles`) in `scripts/select_article_extractor.py`
- [X] T005 Implement unit tests for normalization, tokenization, and shingle generation in `tests/test_select_article_extractor.py`
**Checkpoint**: Foundation ready — text normalization and shingle generator fully operational and tested.
---
## Phase 3: User Story 1 - Deterministic Selection with Text Consensus (Priority: P1) 🎯 MVP
**Goal**: Calculate consensus shingles ($\ge 2$ active extractors), compute Coverage, Support, and $F_1$ score, apply technical tie margin ($\le 0.03$), and select the best extractor based on agreement and conciseness.
**Independent Test**: Execute tests with synthetic and real articles where 2 or 3 extractors agree (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) and verify the winner matches expected score / tie-breaker.
### Tests for User Story 1 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T006 [P] [US1] Write unit tests for consensus scoring, coverage/support F1 calculation, and technical tie-breaking (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) in `tests/test_select_article_extractor.py`
### Implementation for User Story 1
- [X] T007 [US1] Implement consensus shingle builder and metric calculator (`calculate_consensus_metrics`) in `scripts/select_article_extractor.py`
- [X] T008 [US1] Implement consensus decision selector with technical tie pool ($\le 0.03$), smallest shingle count preference, and final priority fallback (`select_with_consensus`) in `scripts/select_article_extractor.py`
**Checkpoint**: User Story 1 (MVP) is fully functional and independently testable for all consensus scenarios.
---
## Phase 4: User Story 2 - Resilient Decision Under Disagreement or Degradation (Priority: P2)
**Goal**: Ensure zero unassigned or ambiguous selections by handling zero-consensus articles (median of 3, max of 2, single), candidate degradation, all-unavailable extractors, and strict tie-breaking priority (`newspaper4k` > `readability` > `trafilatura`).
**Independent Test**: Execute tests for zero consensus, degraded errors, single usable candidate, and total extraction failure (CT-004, CT-005, CT-006, CT-007, CT-008).
### Tests for User Story 2 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T009 [P] [US2] Write unit tests for zero-consensus, degraded candidate handling, single candidate, and total unavailability fallback (CT-004, CT-005, CT-006, CT-007, CT-008) in `tests/test_select_article_extractor.py`
### Implementation for User Story 2
- [X] T010 [US2] Implement candidate status classifier (`classify_candidate_status`) and active candidate set builder (`form_active_set`) in `scripts/select_article_extractor.py`
- [X] T011 [US2] Implement zero-consensus decision logic (median of 3, max of 2, single candidate, priority hierarchy) (`select_without_consensus`) in `scripts/select_article_extractor.py`
- [X] T012 [US2] Implement single article selector orchestrator (`select_article_extractor`) integrating Usable, Degraded, Consensus, and Non-Consensus decision branches in `scripts/select_article_extractor.py`
**Checkpoint**: User Stories 1 AND 2 are fully functional and handle 100% of single-article decision paths.
---
## Phase 5: User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
**Goal**: Read JSON batch files containing `articles`, preserve all original fields and article order, recalculate existing `selected_extractor` values, validate schema, and write output atomically to `<name>_selected.json`.
**Independent Test**: Execute CLI and batch tests (CT-012, CT-013, CT-014), verify non-destructive field preservation, and run end-to-end processing on `out/river_plate_extracted.json`.
### Tests for User Story 3 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T013 [P] [US3] Write integration tests for JSON schema validation, error handling, key preservation, recalculation of existing key, and atomic I/O (CT-012, CT-013, CT-014) in `tests/test_select_article_extractor.py`
### Implementation for User Story 3
- [X] T014 [US3] Implement batch processor (`process_batch`) and atomic file saver (`atomic_save_json`) in `scripts/select_article_extractor.py`
- [X] T015 [US3] Implement CLI interface (`main`) with `argparse`, options (`-o`, `--indent`, `--verbose`), exit codes (0, 1, 2), `stderr` logs, and `stdout` JSON summary in `scripts/select_article_extractor.py`
**Checkpoint**: Full end-to-end batch processing operational and tested against contracts and real datasets.
---
## Phase 6: Polish & Cross-Cutting Concerns
**Purpose**: Validation, performance checks, and documentation verification
- [X] T016 [P] Execute quickstart validation scenarios on `out/river_plate_extracted.json` per `specs/004-deterministic-content-selection/quickstart.md`
- [X] T017 Run full test suite with coverage via `pytest tests/test_select_article_extractor.py -v` ensuring all 14 mandatory test cases (CT-001 to CT-014) pass
---
## Dependencies & Execution Order
```mermaid
graph TD
T001[T001: Setup Harness] --> T002[T002: Data Models]
T001 --> T003[T003: Text Normalization]
T002 --> T004[T004: Shingle Generator]
T003 --> T004
T004 --> T005[T005: Foundation Tests]
T005 --> T006[T006: US1 Tests]
T006 --> T007[T007: US1 Consensus Metrics]
T007 --> T008[T008: US1 Consensus Selection]
T008 --> T009[T009: US2 Tests]
T009 --> T010[T010: US2 Candidate Classifier]
T010 --> T011[T011: US2 Zero-Consensus Logic]
T011 --> T012[T012: US2 Orchestrator]
T012 --> T013[T013: US3 Integration Tests]
T013 --> T014[T014: US3 Batch & Atomic I/O]
T014 --> T015[T015: US3 CLI Interface]
T015 --> T016[T016: Quickstart Validation]
T016 --> T017[T017: Full Test Suite CT-001..CT-014]
```
---
## Parallel Opportunities
- **Phase 2 (Foundations)**: `T002` (Models) and `T003` (Normalization) can be developed in parallel.
- **Phase 3 (User Story 1)**: `T006` (Unit tests) can be written in parallel with data model test setups.
- **Phase 4 (User Story 2)**: `T009` (Unit tests) can be written in parallel with candidate state transition logic.
- **Phase 5 (User Story 3)**: `T013` (Integration tests) can be developed alongside CLI option parser definitions.
---
## Implementation Strategy
### MVP First (Phases 1, 2 & 3)
1. Complete Setup and Foundations (`T001` - `T005`).
2. Implement User Story 1 (`T006` - `T008`): text consensus and scoring engine.
3. Validate MVP: test consensus-based decisions on multi-extractor articles.
### Incremental Delivery (Phases 4, 5 & 6)
4. Add User Story 2 (`T009` - `T012`): resilient fallbacks, zero consensus, degraded candidate evaluation.
5. Add User Story 3 (`T013` - `T015`): non-destructive batch JSON processing, atomic save, and CLI tool.
6. Polish & Final Verification (`T016` - `T017`): run quickstart validation and full test suite covering CT-001 to CT-014.