feat: add deterministic content extractor selector engine with F1 consensus

This commit is contained in:
2026-08-20 22:09:43 -03:00
parent 6a45368cb0
commit ff7a50e0eb
46 changed files with 18503 additions and 2813 deletions
@@ -0,0 +1,54 @@
# Deterministic Content Selection Checklist: End-to-End Requirements Quality
**Purpose**: Validate the completeness, clarity, consistency, and measurability of requirements for the deterministic extractor selection pipeline
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
---
## Text Normalization & Tokenization Quality
- [x] CHK001 Are Unicode NFKC normalization rules explicitly specified for multilingual text content? [Completeness, Spec §FR-007]
- [x] CHK002 Is HTML and Markdown tag stripping behavior defined to prevent accidental concatenation of neighboring words? [Clarity, Spec §FR-007]
- [x] CHK003 Are anchor text extraction rules for Markdown and HTML links documented unambiguously? [Clarity, Spec §FR-007]
- [x] CHK004 Is the tokenization behavior (Unicode alphanumeric tokens, punctuation exclusion, lowercase) completely specified? [Completeness, Spec §FR-007]
## Shingles & Consensus Metric Formulation
- [x] CHK005 Is the sliding window shingle size (5-tokens) and the fallback rule for short texts (< 5 tokens) explicitly defined? [Clarity, Spec §FR-008]
- [x] CHK006 Are the mathematical formulas for Coverage, Support, and F1 Score defined with explicit zero-division handling? [Measurability, Spec §FR-010]
- [x] CHK007 Is the threshold for a shingle to enter the Consensus set (presence in $\ge 2$ active candidates) unambiguously stated? [Clarity, Spec §FR-009]
## Decision & Tie-Breaking Hierarchy
- [x] CHK008 Is the technical tie threshold ($\le 0.03$) quantified with exact comparison semantics? [Clarity, Spec §FR-011]
- [x] CHK009 Is the tie-breaker preference for the smaller candidate (fewest shingles) explicitly constrained to candidates within the technical tie pool? [Consistency, Spec §FR-011]
- [x] CHK010 Is the zero-consensus fallback hierarchy (median of 3, maximum of 2, single candidate) completely specified without ambiguous gaps? [Coverage, Spec §FR-012]
- [x] CHK011 Is the final mandatory priority order (`newspaper4k` > `readability` > `trafilatura`) consistent across all tie scenarios? [Consistency, Spec §FR-011, §FR-012]
## Candidate State Transitions & Resilience
- [x] CHK012 Are the criteria distinguishing Usable, Degraded, and Unavailable candidates defined unambiguously? [Completeness, Spec §FR-005]
- [x] CHK013 Does the spec define the exact behavior and fallback when all 3 extractors are Unavailable? [Edge Case, Spec §FR-013]
- [x] CHK014 Does the spec define what occurs when degraded candidates exist but no usable candidates are present? [Coverage, Spec §FR-006]
## JSON Schema Integrity & Atomic I/O
- [x] CHK015 Are requirements explicit that 100% of pre-existing fields, structures, and article order must be preserved unchanged? [Completeness, Spec §FR-003, §FR-010]
- [x] CHK016 Is the output filename pattern `<original_name_without_extension>_selected.json` specified for default CLI execution? [Clarity, Spec §FR-016]
- [x] CHK017 Are atomic write requirements (temporary file + atomic replacement) defined to prevent partial or corrupted files on disk? [Non-Functional, Spec §FR-002, §FR-016]
- [x] CHK018 Is the behavior for recalculating an already present `selected_extractor` key explicitly specified? [Clarity, Spec §FR-014]
---
## Notes
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied.
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation.
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers.
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`.
- Items are numbered sequentially (CHK001 - CHK018) for easy reference.
@@ -0,0 +1,36 @@
# Specification Quality Checklist: Deterministic Content Selection
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- All requirements are derived directly from PRD `docs/prd_deterministic_content_selection.md`.
- No ambiguity remains; all edge cases and tie-breaking hierarchies are fully specified.
- Ready for `/speckit-plan`.
@@ -0,0 +1,54 @@
# CLI Interface Contract: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](../spec.md)
---
## 1. Command Syntax
```bash
python scripts/select_article_extractor.py <input_file> [-o OUTPUT] [--indent INDENT] [--verbose]
```
---
## 2. Arguments and Flags
| Argumento / Flag | Tipo | Obrigatório | Padrão | Descrição |
|---|---|:---:|---|---|
| `input_file` | `Path` (Posicional) | Sim | - | Caminho para o arquivo JSON contendo a coleção `articles` extraída. |
| `-o`, `--output` | `Path` | Não | `<input_file_without_ext>_selected.json` | Caminho do arquivo JSON de destino. Se omitido, grava no mesmo diretório com sufixo `_selected.json`. |
| `--indent` | `int` | Não | `2` | Número de espaços para indentação do JSON de saída. Use `0` para JSON compacto em linha única. |
| `-v`, `--verbose` | `flag` | Não | `False` | Exibe no `stderr` detalhes da pontuação e justificativa de escolha por artigo. |
---
## 3. Standard Streams (I/O)
- **`stdout`**:
- Emite o sumário operacional em JSON ou texto resumido ao término da execução:
```json
{
"status": "success",
"input_file": "out/river_plate_extracted.json",
"output_file": "out/river_plate_extracted_selected.json",
"total_articles": 20,
"distribution": {
"newspaper4k": 9,
"readability": 9,
"trafilatura": 2
}
}
```
- **`stderr`**:
- Mensagens de log, progresso da barra/processamento de artigos e erros de validação ou exceções.
---
## 4. Exit Codes
| Código | Significado | Comportamento |
|:---:|---|---|
| `0` | **Sucesso** | Todos os artigos foram processados e o arquivo final foi gravado atomicamente com sucesso. |
| `1` | **Erro de I/O ou JSON Inválido** | Arquivo não encontrado, JSON malformado ou permissão negada. Nenhum arquivo de saída é gerado. |
| `2` | **Erro de Validação de Estrutura** | Raiz não é objeto ou chave `articles` não é uma lista. Nenhum arquivo de saída é gerado. |
@@ -0,0 +1,80 @@
# JSON Schema Contract: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](../spec.md)
---
## 1. Input JSON Schema
O arquivo de entrada deve conter uma lista de artigos sob a chave `articles`.
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["articles"],
"properties": {
"articles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"trafilatura": {
"type": "object",
"properties": {
"text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
},
"newspaper4k": {
"type": "object",
"properties": {
"text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
},
"readability": {
"type": "object",
"properties": {
"cleaned_text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
}
}
}
}
}
}
```
---
## 2. Output JSON Schema
O arquivo de saída mantém todos os campos, metadados e ordem originais, adicionando obrigatoriamente `selected_extractor`.
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["articles"],
"properties": {
"articles": {
"type": "array",
"items": {
"type": "object",
"required": ["selected_extractor"],
"properties": {
"selected_extractor": {
"type": "string",
"enum": ["trafilatura", "newspaper4k", "readability"]
},
"trafilatura": { "type": "object" },
"newspaper4k": { "type": "object" },
"readability": { "type": "object" }
}
}
}
}
}
```
@@ -0,0 +1,137 @@
# Data Model: Deterministic Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Domain Entities & Value Types
```mermaid
classDiagram
class ExtractorName {
<<enumeration>>
TRAFILATURA = "trafilatura"
NEWSPAPER4K = "newspaper4k"
READABILITY = "readability"
}
class CandidateStatus {
<<enumeration>>
USABLE
DEGRADED
UNAVAILABLE
}
class ExtractorCandidate {
+ExtractorName name
+str raw_text
+str error
+CandidateStatus status
+List~str~ tokens
+Set~Tuple~ shingles
+int shingle_count
+float coverage
+float support
+float score
}
class ArticleSelectionResult {
+int article_index
+ExtractorName selected_extractor
+str selection_reason
+int active_candidates_count
+int consensus_shingles_count
+Dict~ExtractorName, ExtractorCandidate~ candidates
}
class BatchProcessingResult {
+int total_articles
+int processed_count
+Dict~str, int~ selection_distribution
+str input_file
+str output_file
}
ExtractorCandidate --> ExtractorName
ExtractorCandidate --> CandidateStatus
ArticleSelectionResult --> ExtractorName
ArticleSelectionResult --> ExtractorCandidate
BatchProcessingResult --> ArticleSelectionResult
```
---
## 2. Entity Descriptions & Fields
### `ExtractorName` (Enum / Literal)
Enumeração estrita com os três motores de extração suportados:
- `"trafilatura"`
- `"newspaper4k"`
- `"readability"`
### `CandidateStatus` (Enum)
Classificação do estado de cada extrator em um dado artigo:
- `USABLE`: Campo de texto contém string não-vazia após normalização e campo `error` é nulo/vazio.
- `DEGRADED`: Campo de texto contém string não-vazia após normalização, porém campo `error` não é nulo.
- `UNAVAILABLE`: Campo de texto é ausente, nulo, tipo diferente de string ou vazio após normalização.
### `ExtractorCandidate` (Dataclass)
Representação estruturada de um candidato durante o cálculo:
| Campo | Tipo | Descrição |
|---|---|---|
| `name` | `ExtractorName` | Identificador do motor de extração (`trafilatura`, `newspaper4k`, `readability`). |
| `raw_text` | `str \| None` | Texto bruto obtido do campo correspondente no JSON (`trafilatura.text`, `newspaper4k.text`, `readability.cleaned_text`). |
| `error` | `str \| None` | Mensagem de erro do motor, se houver (`trafilatura.error`, etc.). |
| `status` | `CandidateStatus` | Estado de viabilidade do candidato (`USABLE`, `DEGRADED`, `UNAVAILABLE`). |
| `tokens` | `list[str]` | Sequência ordenada de tokens alfanuméricos minúsculos após normalização NFKC. |
| `shingles` | `set[tuple[str, ...]]` | Conjunto de n-grams consecutivos de 5 tokens (ou 1 n-gram se $1 \le \text{tokens} \le 4$). |
| `shingle_count` | `int` | Quantidade total de shingles gerados (`len(shingles)`). |
| `coverage` | `float` | Proporção de shingles do consenso presentes no candidato ($[0.0, 1.0]$). |
| `support` | `float` | Proporção de shingles do candidato que pertencem ao consenso ($[0.0, 1.0]$). |
| `score` | `float` | Pontuação $F_1$ baseada em cobertura e suporte ($[0.0, 1.0]$). |
---
### `ArticleSelectionResult` (Dataclass)
Resultado detalhado da avaliação para um único artigo:
| Campo | Tipo | Descrição |
|---|---|---|
| `article_index` | `int` | Posição ordinal do artigo no array `articles` original (0-indexed). |
| `selected_extractor` | `ExtractorName` | Vencedor da seleção determinística (`trafilatura`, `newspaper4k`, `readability`). |
| `selection_reason` | `str` | Justificativa rastreável da escolha (ex: `"highest_score"`, `"technical_tie_smallest_shingles"`, `"no_consensus_median_shingles"`, `"single_usable_candidate"`, `"fallback_all_unavailable"`). |
| `active_candidates_count` | `int` | Número de candidatos que formaram o conjunto ativo avaliado. |
| `consensus_shingles_count` | `int` | Quantidade de shingles no conjunto de consenso. |
| `candidates` | `dict[ExtractorName, ExtractorCandidate]` | Dicionário com o detalhamento de cada um dos 3 motores. |
---
### `BatchProcessingResult` (Dataclass)
Sumário da execução do lote:
| Campo | Tipo | Descrição |
|---|---|---|
| `total_articles` | `int` | Total de artigos encontrados no arquivo de entrada. |
| `processed_count` | `int` | Total de artigos processados e enriquecidos com sucesso. |
| `selection_distribution` | `dict[str, int]` | Contagem de seleções por motor (`{"trafilatura": X, "newspaper4k": Y, "readability": Z}`). |
| `input_file` | `str` | Caminho do arquivo lido. |
| `output_file` | `str` | Caminho do arquivo gerado de forma atômica. |
---
## 3. JSON Schema Mapping
### Entrada
- Raiz: Objeto contendo chave `articles: list[dict]`.
- Cada item em `articles`:
- `trafilatura` (objeto opcional): `{ "text": str | null, "error": str | null, ... }`
- `newspaper4k` (objeto opcional): `{ "text": str | null, "error": str | null, ... }`
- `readability` (objeto opcional): `{ "cleaned_text": str | null, "error": str | null, ... }`
### Saída
- Mesma estrutura exata da entrada, preservando 100% dos dados anteriores e ordem da lista `articles`.
- Em cada item de `articles`, adição/atualização da chave:
```json
"selected_extractor": "newspaper4k" | "readability" | "trafilatura"
```
@@ -0,0 +1,93 @@
# Implementation Plan: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
**Input**: Feature specification from `specs/004-deterministic-content-selection/spec.md`
---
## Summary
Implementação do motor determinístico de seleção de extratores (`scripts/select_article_extractor.py`), capaz de consumir arquivos JSON consolidados com saídas do **Trafilatura**, **Newspaper4k** e **Readability**, aplicar normalização de texto, geração de shingles (5-tokens), pontuação $F_1$ baseada em consenso e regras de desempate técnico / hierárquico estritas, gerando um novo arquivo JSON enriquecido exclusivamente com a chave `selected_extractor` em cada artigo de forma não-destrutiva e atômica.
---
## Technical Context
**Language/Version**: Python 3.10+
**Primary Dependencies**: Standard Library (`json`, `re`, `unicodedata`, `html`, `argparse`, `dataclasses`, `pathlib`, `tempfile`, `os`)
**Storage**: Arquivos JSON locais no diretório `out/`
**Testing**: `pytest` com testes unitários e de integração cobrindo 100% dos casos de teste obrigatórios (CT-001 a CT-014)
**Target Platform**: Windows / Linux / macOS (Terminal CLI & Módulo Python)
**Project Type**: CLI tool & modular selection engine
**Performance Goals**: Processamento em lote de centenas de artigos em menos de 1 segundo (complexidade linear $O(N)$ em memória)
**Constraints**: 100% determinístico, 0 chamadas de rede, sem uso de LLMs ou embeddings, escrita atômica em disco
**Scale/Scope**: Lotes de 1 a 10.000+ artigos
---
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Princípio | Avaliação | Status |
|---|---|---|
| **I. Library / Modular Design** | Módulo estruturado com funções puras e dataclasses desacopladas (`normalize_text`, `generate_shingles`, `calculate_consensus_metrics`, `select_best_candidate`, `process_batch`). | ✅ Aprovado |
| **II. CLI Interface** | CLI via `scripts/select_article_extractor.py` com flags descritivas, streams padronizados (`stdout` para resumo e `stderr` para logs/erros) e códigos de saída específicos. | ✅ Aprovado |
| **III. Test-First (NON-NEGOTIABLE)** | TDD com suíte automatizada em `tests/test_select_article_extractor.py` cobrindo todos os cenários (CT-001 a CT-014) e validação end-to-end com o arquivo real `out/river_plate_extracted.json`. | ✅ Aprovado |
| **IV. Integration Testing** | Testes de integração validando leitura, enriquecimento de `selected_extractor`, não-destrutividade de campos e escrita atômica. | ✅ Aprovado |
| **V. Simplicity & YAGNI** | Uso exclusivo da biblioteca padrão do Python, sem dependências adicionais pesadas. | ✅ Aprovado |
---
## Project Structure
### Documentation (this feature)
```text
specs/004-deterministic-content-selection/
├── spec.md # Especificação de requisitos funcionais e critérios
├── plan.md # Este plano de implementação (/speckit-plan)
├── research.md # Decisões técnicas e algoritmos (Phase 0)
├── data-model.md # Entidades e modelos de dados (Phase 1)
├── quickstart.md # Guia de validação e execução (Phase 1)
├── contracts/
│ ├── cli-contract.md # Contrato de linha de comando
│ └── json-schema.md # Esquemas JSON de entrada e saída
└── checklists/
└── requirements.md # Checklist de validação da especificação
```
### Source Code Layout
```text
scripts/
├── extract_google_news.py # Extrator RSS do Google News
├── extract_article_contents.py # Extrator multimotor de artigos
└── select_article_extractor.py # [NEW] Seletor determinístico de extrator por artigo
tests/
├── test_extract_google_news.py # Testes do extrator Google News
├── test_extract_article_contents.py # Testes do extrator multimotor
└── test_select_article_extractor.py # [NEW] Testes unitários e de integração do seletor
```
**Structure Decision**: Criação de `scripts/select_article_extractor.py` como ferramenta CLI e biblioteca modular autônoma, e `tests/test_select_article_extractor.py` contendo a suíte de testes de alta fidelidade aos requisitos do PRD.
---
## Implementation Phases
### Phase 0: Outline & Research *(Completed)*
- Normalização de texto via biblioteca padrão (`html.unescape`, `unicodedata.normalize('NFKC')`, regex Unicode).
- Estratégia de geração de shingles de 5 tokens e cálculo de $F_1$ sobre consenso compartilhado por $\ge 2$ motores.
- Regras de desempate técnico (`<= 0.03`), desempate sem consenso (mediana/máximo) e fallback prioritário (`newspaper4k` > `readability` > `trafilatura`).
- Documentado em [research.md](research.md).
### Phase 1: Design & Contracts *(Completed)*
- Modelos de dados e dataclasses estruturados em [data-model.md](data-model.md).
- Contratos de linha de comando e JSON schema definidos em [contracts/](contracts/).
- Guia prático de execução e validação estruturado em [quickstart.md](quickstart.md).
### Phase 2: Tasks & Execution *(Next Step via `/speckit-tasks`)*
- Criação das tarefas de implementação e testes orientados a TDD em `tasks.md`.
@@ -0,0 +1,66 @@
# Quickstart: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Pré-requisitos
- Python 3.10+
- Ambiente virtual configurado com dependências do projeto instaladas (`pip install -r requirements.txt`).
---
## 2. Execução Rápida via CLI
### Cenário 1: Selecionar o melhor extrator para uma extração existente
```bash
python scripts/select_article_extractor.py out/river_plate_extracted.json
```
**Resultado esperado**:
- Arquivo `out/river_plate_extracted_selected.json` gerado contendo todos os 20 artigos com a chave `selected_extractor` devidamente preenchida (`trafilatura`, `newspaper4k` ou `readability`).
- O arquivo original `out/river_plate_extracted.json` permanece inalterado.
### Cenário 2: Especificar caminho de saída customizado e modo verboso
```bash
python scripts/select_article_extractor.py out/river_plate_extracted.json -o out/meu_resultado.json --verbose
```
**Resultado esperado**:
- Logs detalhados no `stderr` mostrando as pontuações e a regra acionada (ex: `highest_score`, `technical_tie`, etc.).
---
## 3. Execução dos Testes Automatizados
Para rodar a suíte completa de testes unitários e de integração (cobrindo os casos CT-001 a CT-014):
```bash
pytest tests/test_select_article_extractor.py -v
```
---
## 4. Validação Programática / Uso como Módulo Python
```python
from scripts.select_article_extractor import select_article_extractor
article_data = {
"trafilatura": {"text": "El Club Atlético River Plate venció 2-0 anoche.", "error": None},
"newspaper4k": {
"text": "El Club Atlético River Plate venció 2-0 anoche en el Monumental.",
"error": None,
},
"readability": {
"cleaned_text": "El Club Atlético River Plate venció 2-0 anoche.",
"error": None,
},
}
result = select_article_extractor(article_data)
print("Extrator selecionado:", result.selected_extractor.value)
print("Motivo da escolha:", result.selection_reason)
# Output esperado:
# Extrator selecionado: newspaper4k (ou readability dependendo do desempate de shingles)
# Motivo da escolha: technical_tie_smallest_shingles (ou highest_score)
```
@@ -0,0 +1,107 @@
# Research & Architectural Decisions: Deterministic Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Text Normalization Pipeline
### Context
Cada extrator (Trafilatura, Newspaper4k e Readability) gera textos com diferentes resíduos de formatação (entidades HTML como `&amp;` ou `&nbsp;`, links no formato markdown `[texto](url)` ou tags `<a href="...">texto</a>`, variações de quebras de linha e pontuações). A comparação textual para consenso exige uma normalização uniforme, determinística e de alta performance.
### Decisions
1. **Decodificação de entidades HTML**: Utilizar `html.unescape()` da biblioteca padrão do Python.
2. **Remoção de imagens Markdown**: Expressão regular `re.compile(r'!\s*\[[^\]]*\]\([^)]*\)')` substituindo imagens Markdown por espaços para descartar marcação de mídia não-textual e evitar falsos consensos com legendas.
3. **Preservação de texto de links Markdown**: Expressão regular `re.compile(r'\[([^\]]+)\]\([^)]+\)')` substituindo links Markdown pelo texto âncora `\1`.
4. **Remoção de tags HTML**: Expressão regular `re.compile(r'<[^>]+>')` substituindo tags por espaços para evitar fusão acidental de palavras vizinhas.
5. **Normalização Unicode**: `unicodedata.normalize('NFKC', text)` para uniformizar caracteres compostos, ligaduras e variantes tipográficas.
6. **Conversão para minúsculas**: `.lower()` após NFKC.
7. **Colapso de espaços em branco**: `re.sub(r'\s+', ' ', text).strip()`.
8. **Tokenização**: Extração de sequências alfanuméricas com `re.findall(r'[\w]+', text, flags=re.UNICODE)`. Pontuações são descartadas naturalmente sem remoção semântica de palavras.
### Rationale
- 100% implementável com módulos padrão do Python (`re`, `unicodedata`, `html`), garantindo portabilidade em qualquer ambiente sem novas dependências externas.
- Complexidade linear $O(N)$ no tamanho do texto, com execução em frações de milissegundo por artigo.
### Alternatives Considered
- `BeautifulSoup` para strip de tags: Rejeitado por ser mais lento e desnecessário para textos já extraídos.
- `nltk` ou `spacy`: Rejeitados por adicionarem dependências pesadas, download de modelos e lentidão desnecessária para uma tarefa de tokenização alfanumérica pura.
---
## 2. 5-Token Shingles & Consensus Metrics
### Context
O algoritmo compara a sobreposição textual entre os candidatos ativos através de janelas deslizantes consecutivas de 5 tokens (shingles).
### Decisions
1. **Geração de Shingles**:
- Para um candidato com $T$ tokens ordenados $[t_0, t_1, \dots, t_{T-1}]$:
- Se $T \ge 5$: conjunto de tuplas de 5 tokens $\{ (t_i, t_{i+1}, t_{i+2}, t_{i+3}, t_{i+4}) \mid 0 \le i \le T-5 \}$.
- Se $1 \le T \le 4$: conjunto contendo uma única tupla com todos os tokens $\{ (t_0, \dots, t_{T-1}) \}$.
- Se $T = 0$: conjunto vazio $\emptyset$.
2. **Construção do Consenso**:
- Para cada shingle único observado nos candidatos ativos, conta-se em quantos candidatos distintos ele aparece.
- $\text{Consenso} = \{ s \mid \text{contagem}(s) \ge 2 \}$.
3. **Métricas por Candidato Ativo $C$**:
- $\text{cobertura}(C) = \frac{|C_{\text{shingles}} \cap \text{Consenso}|}{|\text{Consenso}|}$
- $\text{suporte}(C) = \frac{|C_{\text{shingles}} \cap \text{Consenso}|}{|C_{\text{shingles}}|}$
- $\text{score}(C) = \frac{2 \times \text{cobertura}(C) \times \text{suporte}(C)}{\text{cobertura}(C) + \text{suporte}(C)}$ (se denominador for zero, $\text{score} = 0.0$).
### Rationale
- A métrica de pontuação $F_1$ penaliza tanto extratores que perderam conteúdo essencial (baixa cobertura) quanto extratores que trouxeram excesso de lixo/boilerplate do site (baixo suporte).
- A representação por `set` de tuplas em Python permite operações de intersecção (`&`) com complexidade ótima de tempo $O(|C|)$.
---
## 3. Regras de Decisão, Empate Técnico e Desempate Hierárquico
### Context
O sistema precisa garantir uma escolha única e determinística em todas as variações possíveis de entrada.
### Decisions
1. **Formação do Conjunto Ativo**:
- Classificação:
- `Usável`: `text` é string não vazia após normalização e `error` é `None`/vazio.
- `Degradado`: `text` é string não vazia após normalização, mas `error` não é `None`.
- `Indisponível`: `text` é nulo, ausente, não-string ou vazio.
- Se houver $\ge 1$ Usável $\to$ Ativos = Usáveis.
- Senão, se houver $\ge 1$ Degradado $\to$ Ativos = Degradados.
- Senão $\to$ Seleciona `newspaper4k` diretamente (Fallback Final).
- Se $|\text{Ativos}| = 1 \to$ Seleciona o único candidato ativo imediatamente.
2. **Seleção Com Consenso ($|\text{Consenso}| > 0$)**:
- Maior score $S_{\max} = \max_{C \in \text{Ativos}} \text{score}(C)$.
- Grupo de empate técnico: $\{ C \in \text{Ativos} \mid S_{\max} - \text{score}(C) \le 0.03 + 10^{-9} \}$.
- Se grupo tiver 1 candidato $\to$ Seleciona ele.
- Se grupo tiver $\ge 2$ candidatos $\to$ Seleciona o candidato com menor $|C_{\text{shingles}}|$ (menor conteúdo excedente).
- Se ainda houver empate no número de shingles $\to$ Desempate por prioridade fixa: `newspaper4k` > `readability` > `trafilatura`.
3. **Seleção Sem Consenso ($|\text{Consenso}| = 0$)**:
- Se $|\text{Ativos}| = 3 \to$ Seleciona candidato com quantidade **mediana** de shingles.
- Se $|\text{Ativos}| = 2 \to$ Seleciona candidato com **maior** quantidade de shingles.
- Se $|\text{Ativos}| = 1 \to$ Seleciona o único candidato.
- Empates na quantidade de shingles $\to$ Prioridade fixa: `newspaper4k` > `readability` > `trafilatura`.
### Rationale
- Total aderência às seções 7.1 a 7.6 do PRD. A tolerância de $10^{-9}$ evita imprecisões de ponto flutuante em comparações `<= 0.03`.
---
## 4. Estratégia de I/O Não Destrutiva e Escrita Atômica
### Context
O processamento em lote deve preservar a ordem dos artigos e todos os campos originais do JSON, gravando o resultado sem risco de corrupção de arquivos em caso de interrupção.
### Decisions
1. **Entrada e Saída**:
- Nome padrão de saída: `<nome_original_sem_extensão>_selected.json`.
- Suporte a argumento opcional de saída `--output / -o`.
2. **Gravação Atômica**:
- Gravar os dados em um arquivo temporário no mesmo diretório (`<saida>.tmp.<pid>`).
- Executar substituição atômica via `os.replace(temp_path, target_path)`.
3. **Preservação de Conteúdo**:
- Carregar o JSON original em estruturas nativas de dicionário/lista.
- Inserir a chave `selected_extractor` diretamente em cada dicionário de artigo.
- Se `selected_extractor` já existir na entrada, sobrescrever com o novo valor recalculado.
### Rationale
- Garante integridade absoluta dos dados contra falhas de disco ou encerramentos abruptos.
@@ -0,0 +1,135 @@
# Feature Specification: Deterministic Content Selection
**Feature Branch**: `004-deterministic-content-selection`
**Created**: 2026-08-20
**Status**: Draft
**Input**: User description: "usando o PRD: docs/prd_deterministic_content_selection.md"
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Deterministic Selection with Text Consensus (Priority: P1)
As a data pipeline consumer or analyst, I want the system to automatically analyze the extracted text from Trafilatura, Newspaper4k, and Readability for each article and pick the single best extractor using consensus and coverage scoring, so that our dataset has high-quality, standardized content without human review.
**Why this priority**: Core value of the feature. Resolves the primary dilemma of choosing between 3 extractor outputs per article based on mutual agreement (consensus shingles) and concise content.
**Independent Test**: Can be tested independently by running the selection algorithm on articles where extractors have high agreement or partial variations, verifying that the extractor with highest F1 score (or closest score with fewest excess shingles) is selected.
**Acceptance Scenarios**:
1. **Given** an article with usable extracts from all 3 libraries where 2 or 3 libraries agree closely, **When** selection is evaluated, **Then** the library with the highest consensus F1-score (or the more concise candidate within a 0.03 technical tie margin) is set in `selected_extractor`.
2. **Given** an extractor with excess boilerplate/noise and two extractors with clean common content, **When** selection is evaluated, **Then** the noisy extractor suffers lower support score and the clean agreeing extractor is selected.
3. **Given** an extractor with only a small snippet and two extractors with complete text, **When** selection is evaluated, **Then** the short snippet loses due to low consensus coverage.
---
### User Story 2 - Resilient Decision Under Total Disagreement or Degradation (Priority: P2)
As a pipeline maintainer, I want the selection algorithm to make a deterministic and sensible fallback choice even when extractors completely disagree, produce errors, or return empty/degraded content, so that the pipeline never halts or leaves an article without a chosen extractor.
**Why this priority**: Essential for pipeline stability. The system must guarantee that every article gets an unambiguous winner without throwing runtime exceptions or generating `null`/`ambiguous` states.
**Independent Test**: Can be tested with synthetic articles representing edge cases: all extractors returning non-overlapping text, extractors reporting errors, or all extractors failing.
**Acceptance Scenarios**:
1. **Given** 3 active candidates with 0 consensus shingles, **When** selection runs, **Then** the candidate with the median shingle length is selected.
2. **Given** 2 active candidates with 0 consensus shingles, **When** selection runs, **Then** the candidate with the larger shingle count is selected.
3. **Given** an article where all usable candidates are absent but degraded candidates exist, **When** selection runs, **Then** the algorithm evaluates only the degraded candidates.
4. **Given** an article where all 3 extractors failed or returned empty content, **When** selection runs, **Then** `newspaper4k` is selected via the mandatory final fallback rule.
---
### User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
As a system operator, I want to pass a JSON file with an `articles` array, execute the deterministic selector, and receive a new file `<original_name>_selected.json` with all original data and order intact plus the `selected_extractor` field, leaving the original file completely untouched.
**Why this priority**: Guarantees data preservation, idempotency, and clean pipeline integration.
**Independent Test**: Can be tested by running the process on a full batch JSON file (such as `out/river_plate_extracted.json`) and comparing input vs output keys, element counts, article order, and field contents.
**Acceptance Scenarios**:
1. **Given** a valid JSON file with $N$ articles, **When** the batch selection is executed, **Then** a new file `<original_name>_selected.json` is generated containing exactly $N$ articles in identical order, each with all original fields plus `selected_extractor`.
2. **Given** an input JSON file where `selected_extractor` already exists, **When** the batch selection is executed, **Then** `selected_extractor` is recalculated and updated.
3. **Given** an invalid JSON file or a file where `articles` is not a list, **When** execution runs, **Then** the process terminates with an error and does not produce a partial or corrupted output file.
---
### Edge Cases
- **Empty `articles` list (`[]`)**: Produces a valid output JSON containing an empty `articles: []` list without errors.
- **Exact score & shingle count tie**: Resolved deterministically by the strict fallback hierarchy: `newspaper4k` > `readability` > `trafilatura`.
- **Single active candidate**: When only 1 library produces usable output, it is selected immediately without computing consensus.
- **Short texts (< 5 tokens)**: When candidate text has between 1 and 4 tokens, the entire token sequence forms a single shingle.
- **Malformed fields / Type mismatch**: If a content field is not a string or missing, the candidate is classified as unavailable.
- **Missing library block**: If an article does not contain a `trafilatura`, `newspaper4k`, or `readability` block, that candidate is treated as unavailable.
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST accept a valid JSON file path containing a root object with an `articles` array.
- **FR-002**: System MUST validate input structure (root is object, `articles` is list) and terminate immediately without creating an output file if validation fails.
- **FR-003**: System MUST process all articles in `articles`, preserving their exact sequence and all existing fields and values without modification.
- **FR-004**: System MUST evaluate extractor candidates using exclusively:
- `trafilatura.text` for Trafilatura
- `newspaper4k.text` for Newspaper4k
- `readability.cleaned_text` for Readability
- **FR-005**: System MUST categorize each candidate into one of three states:
- *Usable*: Content is non-empty string after normalization and extractor `error` is null/empty.
- *Degraded*: Content is non-empty string after normalization but extractor `error` is non-null.
- *Unavailable*: Content is missing, not a string, or empty after normalization.
- **FR-006**: System MUST form the active candidate set per article: Usable candidates if any exist; otherwise Degraded candidates if any exist; otherwise trigger final fallback.
- **FR-007**: System MUST perform deterministic in-memory normalization for candidate comparisons:
1. Decode HTML entities.
2. Strip Markdown images (`![alt](url)`), removing non-textual media embeds.
3. In Markdown links (`[text](url)`), preserve anchor text and strip URL targets.
4. Strip HTML tags, maintaining spacing between adjacent words.
5. Apply Unicode NFKC normalization.
6. Convert to lowercase.
7. Collapse multiple whitespace/newlines/tabs into a single space.
8. Tokenize retaining Unicode letters and digits.
9. Ignore punctuation symbols.
- **FR-008**: System MUST generate 5-token sliding window shingles from the ordered token sequence of each candidate (or single $N$-token shingle if $1 \le N \le 4$).
- **FR-009**: System MUST construct the consensus shingle set (shingles appearing in at least 2 active candidates).
- **FR-010**: System MUST compute `coverage`, `support`, and `score` ($F_1 = 2 \times \text{coverage} \times \text{support} / (\text{coverage} + \text{support})$) for each active candidate against the consensus shingles (or 0 if denominator is 0).
- **FR-011**: When consensus shingles exist, the system MUST:
1. Sort candidates descending by score.
2. Identify all candidates within a `0.03` difference from the top score (technical tie pool).
3. If technical tie pool has 1 candidate, select it.
4. If multiple candidates are in technical tie, select the one with the smallest total shingle count (least surplus).
5. If shingle count is also tied, apply priority hierarchy: `newspaper4k` > `readability` > `trafilatura`.
- **FR-012**: When 0 consensus shingles exist, the system MUST:
- With 3 active candidates: select candidate with median shingle count.
- With 2 active candidates: select candidate with maximum shingle count.
- With 1 active candidate: select that single candidate.
- In shingle count ties: apply priority hierarchy (`newspaper4k` > `readability` > `trafilatura`).
- **FR-013**: When 0 active candidates exist (all unavailable), system MUST assign `newspaper4k`.
- **FR-014**: System MUST inject or replace `selected_extractor` in each article item with exactly one value from `{"trafilatura", "newspaper4k", "readability"}`.
- **FR-015**: System MUST never output `null`, empty string, `ambiguous`, or leave an article without a selection.
- **FR-016**: System MUST write the result atomically to `<original_name_without_extension>_selected.json` in the same directory or specified target, leaving the input file unchanged.
- **FR-017**: System MUST produce 100% deterministic and identical outputs across repeated runs with identical inputs.
### Key Entities *(include if feature involves data)*
- **Article Input Batch**: Root JSON container with metadata and an ordered list of `articles`.
- **Article Record**: Object representing an article, containing source metadata, extraction results from the 3 extractors (`trafilatura`, `newspaper4k`, `readability`), and the resulting `selected_extractor` tag.
- **Extractor Candidate**: Evaluation model for an individual extractor containing raw text, error state, candidate usability state (`Usable`, `Degraded`, `Unavailable`), normalized token stream, 5-token shingles, and computed metrics (`coverage`, `support`, `score`, `shingle_count`).
- **Consensus Shingle Set**: Set of unique 5-token shingles shared by 2 or more active extractor candidates.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: **100% Selection Completeness**: 100% of articles in the input collection receive a valid `selected_extractor` value from the closed set `['trafilatura', 'newspaper4k', 'readability']`.
- **SC-002**: **0% Ambiguity**: Exactly 0 articles result in `null`, missing, empty, or ambiguous selection states.
- **SC-003**: **100% Deterministic Reproducibility**: 100% identical `selected_extractor` values when executing across multiple runs on identical input datasets.
- **SC-004**: **100% Non-Destructive Integrity**: 100% of pre-existing keys, nested objects, article counts, and article ordering are preserved identically in the output JSON.
- **SC-005**: **100% Test Case Coverage**: Passes 100% of defined mandatory test cases (CT-001 through CT-014).
- **SC-006**: **Atomic Operation**: 0 partial or corrupted output files generated on process failure or invalid JSON inputs.
## Assumptions
- The input JSON is generated by the extraction pipeline and contains `articles` where each item may have `trafilatura`, `newspaper4k`, and `readability` sub-objects.
- All three extraction libraries operated on the exact same HTML source document.
- No external dependencies (LLM APIs, embedding services, or network calls) are permitted during the selection process.
- Unicode NFKC normalization and standard tokenization cover multilingual article content (e.g. Portuguese, Spanish, English).
- Default output file path naming convention `<name>_selected.json` is sufficient, with CLI support for optional custom destination.
@@ -0,0 +1,148 @@
# Tasks: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Spec**: [spec.md](spec.md) | **Plan**: [plan.md](plan.md)
---
## Phase 1: Setup (Shared Infrastructure)
**Purpose**: Project initialization and test harness setup
- [X] T001 Initialize script entrypoint and test suite structure in `scripts/select_article_extractor.py` and `tests/test_select_article_extractor.py`
---
## Phase 2: Foundational (Data Structures & Normalization Engine)
**Purpose**: Core data models, text normalization, and shingle generation that all user stories depend upon
**⚠️ CRITICAL**: Must be completed before user story implementation begins
- [X] T002 [P] Implement dataclasses and enumerations (`ExtractorName`, `CandidateStatus`, `ExtractorCandidate`, `ArticleSelectionResult`, `BatchProcessingResult`) in `scripts/select_article_extractor.py`
- [X] T003 [P] Implement text normalization pipeline (`normalize_text`, HTML entities unescape, HTML/Markdown tag strip, NFKC, lowercase, Unicode tokenization) in `scripts/select_article_extractor.py`
- [X] T004 Implement 5-token sliding window and short text shingle generator (`generate_shingles`) in `scripts/select_article_extractor.py`
- [X] T005 Implement unit tests for normalization, tokenization, and shingle generation in `tests/test_select_article_extractor.py`
**Checkpoint**: Foundation ready — text normalization and shingle generator fully operational and tested.
---
## Phase 3: User Story 1 - Deterministic Selection with Text Consensus (Priority: P1) 🎯 MVP
**Goal**: Calculate consensus shingles ($\ge 2$ active extractors), compute Coverage, Support, and $F_1$ score, apply technical tie margin ($\le 0.03$), and select the best extractor based on agreement and conciseness.
**Independent Test**: Execute tests with synthetic and real articles where 2 or 3 extractors agree (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) and verify the winner matches expected score / tie-breaker.
### Tests for User Story 1 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T006 [P] [US1] Write unit tests for consensus scoring, coverage/support F1 calculation, and technical tie-breaking (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) in `tests/test_select_article_extractor.py`
### Implementation for User Story 1
- [X] T007 [US1] Implement consensus shingle builder and metric calculator (`calculate_consensus_metrics`) in `scripts/select_article_extractor.py`
- [X] T008 [US1] Implement consensus decision selector with technical tie pool ($\le 0.03$), smallest shingle count preference, and final priority fallback (`select_with_consensus`) in `scripts/select_article_extractor.py`
**Checkpoint**: User Story 1 (MVP) is fully functional and independently testable for all consensus scenarios.
---
## Phase 4: User Story 2 - Resilient Decision Under Disagreement or Degradation (Priority: P2)
**Goal**: Ensure zero unassigned or ambiguous selections by handling zero-consensus articles (median of 3, max of 2, single), candidate degradation, all-unavailable extractors, and strict tie-breaking priority (`newspaper4k` > `readability` > `trafilatura`).
**Independent Test**: Execute tests for zero consensus, degraded errors, single usable candidate, and total extraction failure (CT-004, CT-005, CT-006, CT-007, CT-008).
### Tests for User Story 2 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T009 [P] [US2] Write unit tests for zero-consensus, degraded candidate handling, single candidate, and total unavailability fallback (CT-004, CT-005, CT-006, CT-007, CT-008) in `tests/test_select_article_extractor.py`
### Implementation for User Story 2
- [X] T010 [US2] Implement candidate status classifier (`classify_candidate_status`) and active candidate set builder (`form_active_set`) in `scripts/select_article_extractor.py`
- [X] T011 [US2] Implement zero-consensus decision logic (median of 3, max of 2, single candidate, priority hierarchy) (`select_without_consensus`) in `scripts/select_article_extractor.py`
- [X] T012 [US2] Implement single article selector orchestrator (`select_article_extractor`) integrating Usable, Degraded, Consensus, and Non-Consensus decision branches in `scripts/select_article_extractor.py`
**Checkpoint**: User Stories 1 AND 2 are fully functional and handle 100% of single-article decision paths.
---
## Phase 5: User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
**Goal**: Read JSON batch files containing `articles`, preserve all original fields and article order, recalculate existing `selected_extractor` values, validate schema, and write output atomically to `<name>_selected.json`.
**Independent Test**: Execute CLI and batch tests (CT-012, CT-013, CT-014), verify non-destructive field preservation, and run end-to-end processing on `out/river_plate_extracted.json`.
### Tests for User Story 3 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T013 [P] [US3] Write integration tests for JSON schema validation, error handling, key preservation, recalculation of existing key, and atomic I/O (CT-012, CT-013, CT-014) in `tests/test_select_article_extractor.py`
### Implementation for User Story 3
- [X] T014 [US3] Implement batch processor (`process_batch`) and atomic file saver (`atomic_save_json`) in `scripts/select_article_extractor.py`
- [X] T015 [US3] Implement CLI interface (`main`) with `argparse`, options (`-o`, `--indent`, `--verbose`), exit codes (0, 1, 2), `stderr` logs, and `stdout` JSON summary in `scripts/select_article_extractor.py`
**Checkpoint**: Full end-to-end batch processing operational and tested against contracts and real datasets.
---
## Phase 6: Polish & Cross-Cutting Concerns
**Purpose**: Validation, performance checks, and documentation verification
- [X] T016 [P] Execute quickstart validation scenarios on `out/river_plate_extracted.json` per `specs/004-deterministic-content-selection/quickstart.md`
- [X] T017 Run full test suite with coverage via `pytest tests/test_select_article_extractor.py -v` ensuring all 14 mandatory test cases (CT-001 to CT-014) pass
---
## Dependencies & Execution Order
```mermaid
graph TD
T001[T001: Setup Harness] --> T002[T002: Data Models]
T001 --> T003[T003: Text Normalization]
T002 --> T004[T004: Shingle Generator]
T003 --> T004
T004 --> T005[T005: Foundation Tests]
T005 --> T006[T006: US1 Tests]
T006 --> T007[T007: US1 Consensus Metrics]
T007 --> T008[T008: US1 Consensus Selection]
T008 --> T009[T009: US2 Tests]
T009 --> T010[T010: US2 Candidate Classifier]
T010 --> T011[T011: US2 Zero-Consensus Logic]
T011 --> T012[T012: US2 Orchestrator]
T012 --> T013[T013: US3 Integration Tests]
T013 --> T014[T014: US3 Batch & Atomic I/O]
T014 --> T015[T015: US3 CLI Interface]
T015 --> T016[T016: Quickstart Validation]
T016 --> T017[T017: Full Test Suite CT-001..CT-014]
```
---
## Parallel Opportunities
- **Phase 2 (Foundations)**: `T002` (Models) and `T003` (Normalization) can be developed in parallel.
- **Phase 3 (User Story 1)**: `T006` (Unit tests) can be written in parallel with data model test setups.
- **Phase 4 (User Story 2)**: `T009` (Unit tests) can be written in parallel with candidate state transition logic.
- **Phase 5 (User Story 3)**: `T013` (Integration tests) can be developed alongside CLI option parser definitions.
---
## Implementation Strategy
### MVP First (Phases 1, 2 & 3)
1. Complete Setup and Foundations (`T001` - `T005`).
2. Implement User Story 1 (`T006` - `T008`): text consensus and scoring engine.
3. Validate MVP: test consensus-based decisions on multi-extractor articles.
### Incremental Delivery (Phases 4, 5 & 6)
4. Add User Story 2 (`T009` - `T012`): resilient fallbacks, zero consensus, degraded candidate evaluation.
5. Add User Story 3 (`T013` - `T015`): non-destructive batch JSON processing, atomic save, and CLI tool.
6. Polish & Final Verification (`T016` - `T017`): run quickstart validation and full test suite covering CT-001 to CT-014.