feat: add deterministic content extractor selector engine with F1 consensus

This commit is contained in:
2026-08-20 22:09:43 -03:00
parent 6a45368cb0
commit ff7a50e0eb
46 changed files with 18503 additions and 2813 deletions
+148 -42
View File
@@ -6,7 +6,7 @@
[![Type Checked](https://img.shields.io/badge/Type%20Check-Mypy-blue.svg)](https://mypy-lang.org/) [![Type Checked](https://img.shields.io/badge/Type%20Check-Mypy-blue.svg)](https://mypy-lang.org/)
[![Tests](https://img.shields.io/badge/Tests-Pytest%20(100%25%20Passing)-brightgreen.svg)](tests/) [![Tests](https://img.shields.io/badge/Tests-Pytest%20(100%25%20Passing)-brightgreen.svg)](tests/)
> Plataforma modular em Python para **Classificação Multilíngue de Inerência de Entidades (NLP/LLM/ECP)** e **Extração Inteligente de Manchetes de Notícias com Evasão Anti-Bot (Google News RSS & Foxcape)**. > Plataforma modular em Python para **Classificação Multilíngue de Inerência de Entidades (NLP/LLM/ECP)**, **Extração Inteligente de Manchetes (Google News RSS & Foxcape)**, **Extração Multimotor de Artigos (Trafilatura, Newspaper4k, Readability)** e **Seleção Determinística de Conteúdo por Consenso Textual ($F_1$ Shingles)**.
--- ---
@@ -29,6 +29,12 @@
- [Visão Geral e Tríplice Extração](#visão-geral-e-tríplice-extração) - [Visão Geral e Tríplice Extração](#visão-geral-e-tríplice-extração)
- [Argumentos e Flags CLI](#argumentos-e-flags-cli) - [Argumentos e Flags CLI](#argumentos-e-flags-cli)
- [Exemplos de Uso](#exemplos-de-uso) - [Exemplos de Uso](#exemplos-de-uso)
- [4. Seletor Determinístico de Conteúdo de Artigos](#4--seletor-determinístico-de-conteúdo-de-artigos)
- [Visão Geral e Algoritmo de Consenso ($F_1$)](#visão-geral-e-algoritmo-de-consenso-f_1)
- [Pipeline de Normalização e Shingles](#pipeline-de-normalização-e-shingles)
- [Critérios de Desempate Técnico e Resiliência](#critérios-de-desempate-técnico-e-resiliência)
- [Argumentos e Flags CLI](#argumentos-e-flags-cli-1)
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso-1)
- [Estrutura do Projeto](#-estrutura-do-projeto) - [Estrutura do Projeto](#-estrutura-do-projeto)
- [Testes e Qualidade de Código](#-testes-e-qualidade-de-código) - [Testes e Qualidade de Código](#-testes-e-qualidade-de-código)
- [Licença](#-licença) - [Licença](#-licença)
@@ -37,11 +43,12 @@
## 🌟 Visão Geral ## 🌟 Visão Geral
O **TextNLPClassifierApp** reúne ferramentas de engenharia de dados e processamento de linguagem natural: O **TextNLPClassifierApp** reúne um ecossistema completo de ferramentas de engenharia de dados e processamento de linguagem natural:
1. **`classify.py`**: Motor de classificação semântica e contextual que determina o grau de aderência e inerência de um documento Markdown em relação a uma entidade alvo definida em um **ECP Snapshot (Entity Context Profile)**. 1. **`classify.py`**: Motor de classificação semântica e contextual que determina o grau de aderência e inerência de um documento Markdown em relação a uma entidade alvo descrita em um **ECP Snapshot (Entity Context Profile)**.
2. **`scripts/extract_google_news.py`**: Extrator de notícias por palavra-chave, idioma e região geográfica que utiliza o motor stealth **Foxcape** (em modo headless), decodificação paralela de URLs para os links reais dos portais de notícias e feedback em tempo real. 2. **`scripts/extract_google_news.py`**: Extrator de notícias por palavra-chave, idioma e região geográfica utilizando navegação stealth **Foxcape** (headless), decodificação paralela de URLs para links reais e feedback em tempo real.
3. **`scripts/extract_article_contents.py`**: Extrator e parser de artigos multimotor com navegação stealth Foxcape headless e extração combinada via **Trafilatura**, **Newspaper4k** (NLP) e **Readability**, consolidando texto higienizado, autores, datas, imagens e resumos em JSON estruturado. 3. **`scripts/extract_article_contents.py`**: Extrator e parser de artigos multimotor com navegação stealth Foxcape headless e extração simultânea via **Trafilatura**, **Newspaper4k** (NLP) e **Readability**, consolidando texto higienizado, autores, datas, imagens e resumos.
4. **`scripts/select_article_extractor.py`**: Motor determinístico de seleção de extratores que avalia as saídas dos três motores, aplica normalização em memória, calcula métricas de consenso de shingles (5-tokens) com pontuação $F_1$, desempata tecnicamente ($\le 0.03$) favorecendo menor concisão/ruído e enriquece os dados de forma não-destrutiva e atômica.
--- ---
@@ -157,7 +164,7 @@ python classify.py --ecp ecp.json --content artigo.md --enable-embeddings --enab
### O que é e Como Funciona ### O que é e Como Funciona
O script [`scripts/extract_google_news.py`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/scripts/extract_google_news.py) é um extrator CLI autônomo projetado para consultar o feed RSS do Google News com máxima velocidade, resiliência e integridade de dados. O script [`scripts/extract_google_news.py`](scripts/extract_google_news.py) é um extrator CLI autônomo projetado para consultar o feed RSS do Google News com máxima velocidade, resiliência e integridade de dados.
### Diferenciais Técnicos ### Diferenciais Técnicos
@@ -182,42 +189,24 @@ O script [`scripts/extract_google_news.py`](file:///c:/Users/aferr/Projects/AFTe
### Exemplos Práticos de Uso ### Exemplos Práticos de Uso
#### 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo)
```bash ```bash
# River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo)
python scripts/extract_google_news.py -q "River Plate" -l es --locale AR -p 2 -o out/river_plate.json python scripts/extract_google_news.py -q "River Plate" -l es --locale AR -p 2 -o out/river_plate.json
```
* **Saída no Terminal**:
```text
[INFO] 🔍 Consultando Google News: 'River Plate' (idioma: es, locale: AR, max_pages: 2)...
[INFO] 📥 Feed RSS recebido (162117 bytes).
[INFO] 📰 20 artigos extraídos do feed XML.
[INFO] 🔗 Decodificando 20 URLs do Google News para os portais reais...
[INFO] ✅ 20/20 URLs resolvidas com sucesso para os domínios de origem.
[INFO] 💾 Arquivo salvo com sucesso: 'out/river_plate.json' (20 notícias).
```
#### 2. Cruzeiro (Brasil / Português / Formatado no Terminal) # Cruzeiro (Brasil / Português / Formatado no Terminal)
```bash
python scripts/extract_google_news.py --query "Cruzeiro" --lang pt --locale BR --pretty python scripts/extract_google_news.py --query "Cruzeiro" --lang pt --locale BR --pretty
```
#### 3. Fórmula 1 (Inglaterra / Inglês) # Filtragem com jq em modo silencioso
```bash
python scripts/extract_google_news.py --query "Formula 1" --lang en --locale GB --pretty
```
#### 4. Filtragem com `jq` em Modo Silencioso
```bash
python scripts/extract_google_news.py -q "inteligência artificial" -s | jq '.items[].url' python scripts/extract_google_news.py -q "inteligência artificial" -s | jq '.items[].url'
``` ```
--- ---
## 3. 📰 Extrator e Parser Multimotor de Artigos ## 3. 📄 Extrator e Parser Multimotor de Artigos
### Visão Geral e Tríplice Extração ### Visão Geral e Tríplice Extração
O script `scripts/extract_article_contents.py` lê os arquivos JSON gerados pelo extrator do Google News (ou qualquer lista contendo `items` com `url`), acessa cada página via **Foxcape** em modo stealth headless (reutilizando uma única sessão de navegador ativa com espera do evento `domcontentloaded`), e executa simultaneamente 3 motores especializados de extração: O script [`scripts/extract_article_contents.py`](scripts/extract_article_contents.py) lê os arquivos JSON gerados pelo extrator do Google News (ou qualquer lista contendo `items` com `url`), acessa cada página via **Foxcape** em modo stealth headless (reutilizando uma única sessão de navegador ativa com espera do evento `domcontentloaded`), e executa simultaneamente 3 motores especializados de extração:
1. **Trafilatura**: Texto principal higienizado, autores, data de publicação, categorias, tags, URL canônica e payload estruturado nativo. 1. **Trafilatura**: Texto principal higienizado, autores, data de publicação, categorias, tags, URL canônica e payload estruturado nativo.
2. **Newspaper4k**: Artigo completo, autores, imagens (`top_image` e galeria), resumo automático e palavras-chave (*keywords*) extraídas por NLP nativo. 2. **Newspaper4k**: Artigo completo, autores, imagens (`top_image` e galeria), resumo automático e palavras-chave (*keywords*) extraídas por NLP nativo.
@@ -238,20 +227,128 @@ O JSON final consolidado é salvo em `out/` com descarte de strings HTML brutas
### Exemplos de Uso ### Exemplos de Uso
#### 1. Extração Completa Automática
```bash ```bash
# Extração Completa Automática (gera out/river_plate_extracted.json)
python scripts/extract_article_contents.py -i out/river_plate.json python scripts/extract_article_contents.py -i out/river_plate.json
# Gera automaticamente out/river_plate_extracted.json
```
#### 2. Amostragem Rápida (Limit 2 Notícias) # Amostragem Rápida (Limit 2 Notícias)
```bash
python scripts/extract_article_contents.py -i out/river_plate.json --limit 2 python scripts/extract_article_contents.py -i out/river_plate.json --limit 2
# Destino Customizado e Timeout Ajustado
python scripts/extract_article_contents.py -i out/petrobras_result.json -o out/petrobras_full.json --timeout 45
``` ```
#### 3. Destino Customizado e Timeout Ajustado ---
## 4. 🎯 Seletor Determinístico de Conteúdo de Artigos
### Visão Geral e Algoritmo de Consenso ($F_1$)
O script [`scripts/select_article_extractor.py`](scripts/select_article_extractor.py) é uma ferramenta autônoma, leve e 100% determinística (baseada exclusivamente na biblioteca padrão do Python) que resolve o problema de divergência entre múltiplos extratores de texto.
O motor analisa as saídas dos candidatos ativos (`trafilatura`, `newspaper4k` e `readability`), compara a sobreposição de conteúdo através de **shingles de 5 tokens** e calcula o índice de concordância mútua via pontuação $F_1$:
$$\text{Coverage} = \frac{|\text{Shingles do Candidato} \cap \text{Consenso}|}{|\text{Consenso}|}$$
$$\text{Support} = \frac{|\text{Shingles do Candidato} \cap \text{Consenso}|}{|\text{Shingles do Candidato}|}$$
$$F_1 = \frac{2 \times \text{Coverage} \times \text{Support}}{\text{Coverage} + \text{Support}}$$
```mermaid
flowchart TD
Art[Artigo com Trafilatura, Newspaper4k e Readability] --> Class[Classificação de Viabilidade: Usable, Degraded, Unavailable]
Class --> Set[Formação do Conjunto Ativo]
Set --> Norm[Normalização NFKC & Shingles de 5 Tokens]
Norm --> Metric[Cálculo de Consenso & Métricas F1]
Metric --> Check{Existe Consenso >= 2?}
Check -- Sim --> TopScore[Avaliação do Top Score]
TopScore --> TiePool{Empate Técnico <= 0.03?}
TiePool -- Sim --> SmallestShingle[Menor Quantidade de Shingles / Menos Ruído]
SmallestShingle --> Winner[Extrator Selecionado]
TiePool -- Não --> HighestScore[Maior Pontuação F1]
HighestScore --> Winner
Check -- Não --> ZeroCons[Desempate Sem Consenso: Mediana / Max / Prioridade]
ZeroCons --> Winner
```
### Pipeline de Normalização e Shingles
A normalização ocorre em memória exclusivamente para fins comparativos:
1. **Decodificação de entidades HTML** (`html.unescape`).
2. **Descarte de imagens Markdown** (`![alt](url)`) para evitar que descrições de imagens criem falsos consensos com legendas.
3. **Preservação de links Markdown** (`[texto](url)` $\to$ `texto`).
4. **Remoção de tags HTML** preservando espaçamento entre palavras adjacentes.
5. **Normalização Unicode NFKC** e conversão para minúsculas.
6. **Colapso de espaços em branco**.
7. **Tokenização Unicode alfanumérica** com descarte de pontuações.
8. **Geração de Shingles**: Janela deslizante de 5 tokens (ou tupla única para textos curtos de 1 a 4 tokens).
### Critérios de Desempate Técnico e Resiliência
* **Empate Técnico ($\le 0.03$)**: Quando dois ou mais extratores atingem pontuações com diferença $\le 0.03$, o algoritmo seleciona aquele com **menor quantidade de shingles** (penalizando *boilerplate*, cabeçalhos ou menus excedentes).
* **Desempate Hierárquico Estrito**: Em caso de empate absoluto de pontuação e tamanho, aplica-se a hierarquia fixa:
$$\text{newspaper4k} > \text{readability} > \text{trafilatura}$$
* **Cenários de 0 Consenso**:
* 3 candidatos ativos $\to$ seleciona o de **tamanho mediano de shingles**.
* 2 candidatos ativos $\to$ seleciona o de **maior tamanho de shingles**.
* 1 candidato ativo $\to$ seleciona o único utilizável.
* Todos indisponíveis $\to$ *fallback* obrigatório em `newspaper4k`.
* **Gravação Atômica e Não-Destrutiva**: Criação de arquivo temporário com substituição atômica (`os.replace`), preservando 100% dos dados pré-existentes, ordem de artigos e propriedades originais.
### Argumentos e Flags CLI
| Parâmetro | Tipo | Padrão | Descrição |
|---|---|---|---|
| `input_file` | Posicional (obrigatório) | — | Caminho para o arquivo JSON contendo a coleção `articles`. |
| `-o, --output` | Caminho (opcional) | `<input_stem>_selected.json` | Caminho do arquivo JSON de destino. |
| `--indent` | Inteiro (opcional) | `2` | Espaços de indentação do JSON (`0` para compacto). |
| `-v, --verbose` | Flag booleana | `False` | Emite no `stderr` os detalhes de pontuação, ativos e regra de escolha por artigo. |
### Exemplos Práticos de Uso
#### 1. Execução Padrão Automática
```bash ```bash
python scripts/extract_article_contents.py -i out/petrobras_result.json -o out/petrobras_full.json --timeout 45 python scripts/select_article_extractor.py out/river_plate_extracted.json
# Gera automaticamente out/river_plate_extracted_selected.json
```
#### 2. Execução com Modo Verboso
```bash
python scripts/select_article_extractor.py out/river_plate_extracted.json --verbose
```
* **Saída no Terminal**:
```text
[Artigo #001] Extrator: newspaper4k | Motivo: technical_tie_smallest_shingles | Ativos: 3 | Consenso: 1048
[Artigo #002] Extrator: newspaper4k | Motivo: highest_score | Ativos: 3 | Consenso: 333
[Artigo #003] Extrator: readability | Motivo: highest_score | Ativos: 3 | Consenso: 778
...
{
"status": "success",
"input_file": "out/river_plate_extracted.json",
"output_file": "out/river_plate_extracted_selected.json",
"total_articles": 20,
"processed_count": 20,
"distribution": {
"newspaper4k": 9,
"readability": 9,
"trafilatura": 2
}
}
```
#### 3. Uso Programático como Módulo Python
```python
from scripts.select_article_extractor import select_article_extractor
article_data = {
"trafilatura": {"text": "River Plate venceu ontem por 3-0.", "error": None},
"newspaper4k": {"text": "River Plate venceu ontem por 3-0 no Monumental.", "error": None},
"readability": {"cleaned_text": "River Plate venceu ontem por 3-0.", "error": None},
}
result = select_article_extractor(article_data)
print("Extrator Selecionado:", result.selected_extractor.value)
print("Motivo:", result.selection_reason)
``` ```
--- ---
@@ -264,7 +361,8 @@ TextNLPClassifierApp/
├── scripts/ ├── scripts/
│ ├── __init__.py # Pacote utilitário de scripts │ ├── __init__.py # Pacote utilitário de scripts
│ ├── extract_google_news.py # CLI de Extração de Manchetes do Google News │ ├── extract_google_news.py # CLI de Extração de Manchetes do Google News
│ └── extract_article_contents.py # CLI de Extração e Parsing Multimotor de Artigos │ ├── extract_article_contents.py # CLI de Extração e Parsing Multimotor de Artigos
│ └── select_article_extractor.py # CLI de Seleção Determinística de Extrator
├── src/ # Módulos centrais do classificador ├── src/ # Módulos centrais do classificador
│ ├── classifier.py # Orquestrador de classificação (Tier 1, 2, 3) │ ├── classifier.py # Orquestrador de classificação (Tier 1, 2, 3)
│ ├── models.py # Modelos de dados e esquemas (ECPSnapshot, Decision) │ ├── models.py # Modelos de dados e esquemas (ECPSnapshot, Decision)
@@ -273,11 +371,13 @@ TextNLPClassifierApp/
├── specs/ # Especificações e planos arquiteturais (Speckit) ├── specs/ # Especificações e planos arquiteturais (Speckit)
│ ├── 001-multilingual-entity-classifier/ │ ├── 001-multilingual-entity-classifier/
│ ├── 002-google-news-extractor/ │ ├── 002-google-news-extractor/
│ └── 003-article-content-extractor/ # Specs da feature de extração multimotor │ ├── 003-article-content-extractor/
│ └── 004-deterministic-content-selection/ # Specs da seleção determinística
├── tests/ # Suíte de testes automatizados ├── tests/ # Suíte de testes automatizados
│ ├── test_classifier.py │ ├── test_classifier.py
│ ├── test_extract_google_news.py │ ├── test_extract_google_news.py
│ └── test_extract_article_contents.py # Testes do extrator de conteúdo │ ├── test_extract_article_contents.py
│ └── test_select_article_extractor.py # Testes do seletor determinístico
├── requirements.txt # Dependências do projeto ├── requirements.txt # Dependências do projeto
├── pyproject.toml # Configurações de ferramentas (pytest, ruff, mypy) ├── pyproject.toml # Configurações de ferramentas (pytest, ruff, mypy)
└── README.md # Documentação principal └── README.md # Documentação principal
@@ -287,13 +387,19 @@ TextNLPClassifierApp/
## 🧪 Testes e Qualidade de Código ## 🧪 Testes e Qualidade de Código
O repositório possui cobertura com testes unitários, testes de integração e testes End-to-End (E2E) com requisição de rede ao vivo: O repositório possui **120 testes automatizados** com 100% de aprovação cobrindo testes unitários, de regressão, de integração e testes End-to-End (E2E) via CLI subprocess:
```bash ```bash
# Executar todos os testes do projeto # Executar toda a suíte de testes do projeto (120 testes)
pytest -v pytest -v
# Executar especificamente os testes do Extrator de Notícias # Executar especificamente os testes do Seletor Determinístico
pytest tests/test_select_article_extractor.py -v
# Executar os testes do Extrator de Conteúdo Multimotor
pytest tests/test_extract_article_contents.py -v
# Executar os testes do Extrator do Google News
pytest tests/test_extract_google_news.py -v pytest tests/test_extract_google_news.py -v
# Validação e correção automática de formatação com Ruff # Validação e correção automática de formatação com Ruff
+279
View File
@@ -0,0 +1,279 @@
# PRD — Seleção determinística da biblioteca de extração de conteúdo
**Versão:** 1.0
**Data:** 20/08/2026
**Status:** Pronto para implementação
## 1. Contexto
Cada artigo é processado pelas bibliotecas Trafilatura, Newspaper4k e Readability a partir do mesmo HTML. O resultado é consolidado em um arquivo JSON que contém, para cada artigo, as saídas das três bibliotecas.
É necessário escolher deterministicamente uma única biblioteca por artigo. A escolha deve ocorrer mesmo quando todos os resultados forem ruins. O processamento não pode retornar estado ambíguo nem deixar um artigo sem seleção.
## 2. Objetivo
Receber um arquivo JSON no formato do arquivo de referência, selecionar a melhor saída de extração disponível para cada artigo e gerar um novo arquivo JSON com os mesmos dados, acrescentando somente a chave `selected_extractor` em cada item de `articles`.
## 3. Escopo
### 3.1 Incluído
- Ler um arquivo JSON com uma coleção `articles`.
- Comparar as saídas de Trafilatura, Newspaper4k e Readability de cada artigo.
- Escolher obrigatoriamente uma das três bibliotecas.
- Adicionar `selected_extractor` em cada artigo.
- Gerar um novo arquivo JSON.
- Preservar os dados e a ordem dos artigos recebidos.
### 3.2 Fora do escopo
- Baixar ou renderizar páginas.
- Verificar se as bibliotecas receberam o mesmo HTML.
- Executar novamente as bibliotecas de extração.
- Limpar, recortar, combinar ou reescrever o conteúdo extraído.
- Gerar Markdown.
- Usar LLM, embeddings ou regras específicas por domínio.
- Alterar qualquer campo existente no JSON.
- Adicionar métricas, justificativas ou outras chaves ao arquivo de saída.
## 4. Premissas
- As três bibliotecas processaram exatamente o mesmo HTML.
- A entrada segue a estrutura do JSON de referência.
- A seleção é executada individualmente para cada artigo.
- O algoritmo deve sempre produzir uma escolha, inclusive em situações sem concordância entre as bibliotecas.
## 5. Entrada
### 5.1 Arquivo
- Formato: JSON válido.
- A raiz deve conter `articles` como uma lista.
- Cada item de `articles` representa um artigo.
### 5.2 Conteúdos comparados
| Biblioteca | Campo usado na comparação |
| ----------- | -------------------------- |
| Trafilatura | `trafilatura.text` |
| Newspaper4k | `newspaper4k.text` |
| Readability | `readability.cleaned_text` |
Os demais campos, incluindo título, descrição, resumo, palavras-chave, HTML estruturado e metadados, não participam da seleção.
### 5.3 Estado de um candidato
Para cada biblioteca, o candidato é classificado em um dos seguintes estados:
| Estado | Condição |
| ------------ | ---------------------------------------------------------------------------------- |
| Utilizável | Campo de conteúdo é uma string não vazia após normalização e `error` é nulo |
| Degradado | Campo de conteúdo é uma string não vazia após normalização, mas `error` não é nulo |
| Indisponível | Campo ausente, nulo, de tipo diferente de string ou vazio após normalização |
O campo de erro considerado é `trafilatura.error`, `newspaper4k.error` ou `readability.error`, conforme a biblioteca.
## 6. Saída
### 6.1 Arquivo
O arquivo de entrada nunca deve ser alterado. Deve ser criado um novo arquivo com o nome:
`<nome_original_sem_extensão>_selected.json`
Exemplo: `river_plate_extracted(2).json` gera `river_plate_extracted(2)_selected.json`.
### 6.2 Alteração permitida
Cada item de `articles` deve receber exatamente uma nova chave no mesmo nível de `trafilatura`, `newspaper4k` e `readability`:
`selected_extractor`
Valores permitidos:
- `trafilatura`
- `newspaper4k`
- `readability`
Não são permitidos `null`, string vazia, `ambiguous` ou qualquer outro valor.
### 6.3 Preservação da entrada
- Todas as chaves e valores existentes devem permanecer semanticamente idênticos.
- A ordem dos itens de `articles` deve ser preservada.
- Nenhum artigo pode ser adicionado ou removido.
- Espaçamento, indentação e ordem textual das chaves do JSON não fazem parte do contrato, pois o arquivo pode ser serializado novamente.
- Caso `selected_extractor` já exista, seu valor deve ser recalculado e substituído.
## 7. Algoritmo determinístico de seleção
### 7.1 Formar o conjunto ativo
Para cada artigo:
1. Identificar os candidatos utilizáveis.
2. Se existir pelo menos um utilizável, considerar somente os utilizáveis.
3. Se não existir utilizável, considerar os candidatos degradados.
4. Se não existir candidato utilizável nem degradado, selecionar `newspaper4k` pelo desempate final obrigatório.
5. Se o conjunto ativo possuir somente um candidato, selecioná-lo imediatamente.
### 7.2 Normalizar os conteúdos
A normalização serve apenas para comparação e não modifica o JSON de saída.
Para cada candidato ativo:
1. Decodificar entidades HTML.
2. Remover marcação HTML e Markdown, preservando o texto visível.
3. Em links, preservar o texto e remover o endereço.
4. Aplicar normalização Unicode NFKC.
5. Converter o texto para minúsculas.
6. Substituir toda sequência de espaços, tabulações ou quebras de linha por um único espaço.
7. Tokenizar mantendo letras e números Unicode.
8. Desconsiderar pontuação.
Nenhuma palavra ou trecho pode ser removido por interpretação semântica.
### 7.3 Gerar shingles
- Gerar a sequência ordenada de tokens de cada candidato.
- Formar o conjunto de todas as janelas consecutivas de cinco tokens.
- Quando o candidato possuir entre um e quatro tokens, usar a sequência completa como um único shingle.
- Candidato sem token é indisponível e não participa do conjunto ativo.
### 7.4 Construir o consenso
O consenso é o conjunto de shingles presentes em pelo menos dois candidatos ativos.
Para cada candidato ativo, calcular:
**Cobertura:**
`coverage = quantidade de shingles do consenso presentes no candidato / quantidade total de shingles do consenso`
**Suporte:**
`support = quantidade de shingles do candidato presentes no consenso / quantidade total de shingles do candidato`
**Pontuação:**
`score = 2 × coverage × support / (coverage + support)`
Quando `coverage + support` for zero, a pontuação será zero.
### 7.5 Selecionar quando existe consenso
1. Ordenar os candidatos por `score`, do maior para o menor.
2. Identificar o maior `score`.
3. Considerar empate técnico todo candidato cuja diferença para o maior `score` seja menor ou igual a `0,03`.
4. Se houver apenas um candidato no empate técnico, selecioná-lo.
5. Se houver empate técnico, selecionar o candidato com a menor quantidade de shingles.
6. Se a quantidade de shingles também empatar, aplicar a prioridade final:
1. `newspaper4k`
2. `readability`
3. `trafilatura`
A preferência pelo menor candidato ocorre somente no empate técnico. Nesse cenário, os candidatos possuem qualidade de concordância equivalente, e a decisão favorece menor conteúdo excedente.
### 7.6 Selecionar quando não existe consenso
Quando nenhum shingle aparece em pelo menos dois candidatos ativos:
- Com três candidatos ativos: selecionar o candidato com a quantidade mediana de shingles.
- Com dois candidatos ativos: selecionar o candidato com a maior quantidade de shingles.
- Com um candidato ativo: selecionar o único candidato.
- Em empate de quantidade: aplicar a prioridade `newspaper4k`, `readability`, `trafilatura`.
- Sem candidato ativo: selecionar `newspaper4k`.
Essas regras garantem uma escolha mesmo quando não existe concordância textual.
### 7.7 Gravar a escolha
Adicionar ou substituir `selected_extractor` no artigo com o identificador da biblioteca vencedora. Repetir o processo até que todos os itens de `articles` tenham sido processados.
## 8. Requisitos funcionais
| ID | Requisito |
| ------ | ------------------------------------------------------------------------------------------------------- |
| FR-001 | O sistema deve aceitar um arquivo JSON como entrada. |
| FR-002 | O sistema deve validar que a raiz é um objeto e que `articles` é uma lista. |
| FR-003 | O sistema deve processar todos os artigos, preservando sua ordem. |
| FR-004 | O sistema deve usar exclusivamente os três campos de conteúdo definidos no PRD para calcular a escolha. |
| FR-005 | O sistema deve executar a normalização e a comparação conforme o algoritmo deste PRD. |
| FR-006 | O sistema deve selecionar exatamente uma biblioteca por artigo. |
| FR-007 | O sistema nunca deve produzir resultado ambíguo. |
| FR-008 | O sistema deve adicionar somente `selected_extractor` em cada artigo. |
| FR-009 | O valor de `selected_extractor` deve pertencer ao catálogo fechado de valores permitidos. |
| FR-010 | O sistema deve preservar todas as chaves e valores recebidos. |
| FR-011 | O sistema deve gerar um novo arquivo e manter o arquivo original inalterado. |
| FR-012 | O sistema deve produzir a mesma seleção sempre que receber exatamente a mesma entrada. |
| FR-013 | O sistema deve recalcular `selected_extractor` quando a chave já existir. |
## 9. Tratamento de erros
| Situação | Comportamento obrigatório |
| ---------------------------------------- | ----------------------------------------------------- |
| JSON inválido | Encerrar o processamento e não gerar arquivo de saída |
| Raiz diferente de objeto | Encerrar o processamento e não gerar arquivo de saída |
| `articles` ausente ou diferente de lista | Encerrar o processamento e não gerar arquivo de saída |
| `articles` vazio | Gerar arquivo com lista vazia e sem outras alterações |
| Estrutura de uma biblioteca ausente | Tratar seu candidato como indisponível |
| Campo de conteúdo com tipo inválido | Tratar seu candidato como indisponível |
| Todas as bibliotecas indisponíveis | Selecionar `newspaper4k` |
| Falha ao gravar o arquivo | Não deixar arquivo de saída parcialmente gravado |
Um erro em um artigo não pode impedir a seleção dos demais artigos, desde que o JSON e a lista `articles` sejam válidos.
## 10. Requisitos não funcionais
| ID | Requisito |
| ------- | ----------------------------------------------------------------------------------------- |
| NFR-001 | O processamento deve ser totalmente determinístico. |
| NFR-002 | O processamento não deve realizar chamadas de rede. |
| NFR-003 | O processamento não deve depender de LLM, embeddings ou serviços externos. |
| NFR-004 | O processamento deve operar somente sobre os dados do arquivo recebido. |
| NFR-005 | A gravação do arquivo deve ser atômica: sucesso completo ou ausência do arquivo de saída. |
## 11. Critérios de aceite
1. Dado o JSON de referência com 20 artigos, o arquivo de saída contém os mesmos 20 artigos na mesma ordem.
2. Cada artigo contém exatamente um `selected_extractor` válido.
3. Nenhum artigo contém `selected_extractor` nulo, vazio ou ambíguo.
4. Todas as chaves e valores anteriores permanecem semanticamente idênticos.
5. Nenhuma chave adicional, além de `selected_extractor`, é criada.
6. O arquivo original permanece inalterado.
7. Duas execuções sobre o mesmo arquivo produzem os mesmos valores de `selected_extractor`.
8. A seleção usa somente `trafilatura.text`, `newspaper4k.text` e `readability.cleaned_text`.
9. Quando todas as saídas estiverem vazias ou indisponíveis, `selected_extractor` recebe `newspaper4k`.
10. Quando não houver consenso, o desempate segue exatamente as regras da seção 7.6.
11. Quando houver empate técnico, o desempate segue exatamente as regras da seção 7.5.
12. Um JSON inválido ou sem `articles` válido não produz arquivo parcial.
## 12. Casos obrigatórios de teste
| Caso | Condição | Resultado esperado |
| ------ | ---------------------------------------------------------------------------- | ------------------------------------------------- |
| CT-001 | Três candidatos com consenso e um vencedor claro | Selecionar o maior `score` |
| CT-002 | Dois ou mais candidatos dentro de `0,03` do maior `score` | Selecionar o de menor quantidade de shingles |
| CT-003 | Empate técnico e mesma quantidade de shingles | Aplicar prioridade final |
| CT-004 | Três candidatos sem consenso | Selecionar a quantidade mediana de shingles |
| CT-005 | Dois candidatos sem consenso | Selecionar a maior quantidade de shingles |
| CT-006 | Somente um candidato utilizável | Selecionar esse candidato |
| CT-007 | Nenhum utilizável, mas existe candidato degradado | Executar o algoritmo somente com os degradados |
| CT-008 | Todos os candidatos indisponíveis | Selecionar `newspaper4k` |
| CT-009 | Readability retorna apenas um fragmento pequeno enquanto os outros concordam | O fragmento perde por baixa cobertura do consenso |
| CT-010 | Um candidato contém o conteúdo comum e muito conteúdo excedente | O candidato perde suporte e reduz sua pontuação |
| CT-011 | Um candidato contém somente parte do conteúdo comum | O candidato perde cobertura e reduz sua pontuação |
| CT-012 | A entrada já contém `selected_extractor` | Recalcular e substituir somente essa chave |
| CT-013 | `articles` está vazio | Gerar saída válida com `articles` vazio |
| CT-014 | JSON inválido | Não gerar saída |
## 13. Definition of Done
- Todos os requisitos funcionais foram implementados.
- Todos os casos obrigatórios de teste foram automatizados e aprovados.
- O JSON de referência com 20 artigos é processado integralmente.
- A saída contém somente a inclusão de `selected_extractor` em cada artigo.
- O arquivo original permanece inalterado.
- Execuções repetidas sobre a mesma entrada produzem as mesmas escolhas.
- Não existe caminho de execução que produza `ambiguous`, `null` ou artigo sem seleção.
+22 -6
View File
@@ -44,10 +44,10 @@
"42": "1. Input Schemas", "42": "1. Input Schemas",
"43": "2. Basic CLI Usage Examples", "43": "2. Basic CLI Usage Examples",
"44": "2. Standard Streams & Exit Codes", "44": "2. Standard Streams & Exit Codes",
"45": "classifier.py", "45": "ClassificationResult",
"46": "InherenceClassifier", "46": "InherenceClassifier",
"47": "detect_language", "47": "classifier.py",
"48": "ClassificationError", "48": "main",
"49": "content_northvolt_de.md", "49": "content_northvolt_de.md",
"50": "content_presal_pt.md", "50": "content_presal_pt.md",
"51": "content_tangential_es.md", "51": "content_tangential_es.md",
@@ -99,12 +99,12 @@
"97": "🧠 TextNLPClassifierApp", "97": "🧠 TextNLPClassifierApp",
"98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor", "98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor",
"99": "sample_rss_xml", "99": "sample_rss_xml",
"100": "main", "100": "models.py",
"101": "ECPSnapshot", "101": "ECPSnapshot",
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)", "102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
"103": "4. Requisitos Funcionais (FR)", "103": "4. Requisitos Funcionais (FR)",
"104": "Tasks: Article Content Multi-Engine Extractor", "104": "Tasks: Article Content Multi-Engine Extractor",
"105": "models.py", "105": "LocalEmbeddingsAdapter",
"106": "Implementation Plan: Article Content Multi-Engine Extractor", "106": "Implementation Plan: Article Content Multi-Engine Extractor",
"107": "2. Cenários de Validação", "107": "2. Cenários de Validação",
"108": "1. Technical Decisions & Tradeoffs", "108": "1. Technical Decisions & Tradeoffs",
@@ -113,5 +113,21 @@
"111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier", "111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier",
"112": "CLI Contract: Article Content Multi-Engine Extractor", "112": "CLI Contract: Article Content Multi-Engine Extractor",
"113": "001-multilingual-entity-classifier/spec.md", "113": "001-multilingual-entity-classifier/spec.md",
"114": "JSON Schema Contract: Article Content Multi-Engine Extractor" "114": "JSON Schema Contract: Article Content Multi-Engine Extractor",
"115": "PRD — Seleção determinística da biblioteca de extração de conteúdo",
"116": "select_article_extractor.py",
"117": "1. Text Normalization Pipeline",
"118": "Tasks: Deterministic Article Content Selection",
"119": "select_article_extractor",
"120": "process_batch",
"122": "test_select_article_extractor.py",
"123": "Feature Specification: Deterministic Content Selection",
"124": "2. Entity Descriptions & Fields",
"125": "Implementation Plan: Deterministic Article Content Selection",
"126": "Deterministic Content Selection Checklist: End-to-End Requirements Quality",
"127": "Quickstart: Deterministic Article Content Selection",
"128": "Specification Quality Checklist: Deterministic Content Selection",
"129": "CLI Interface Contract: Deterministic Article Content Selection",
"130": "004-deterministic-content-selection/spec.md",
"131": "JSON Schema Contract: Deterministic Article Content Selection"
} }
+1 -1
View File
@@ -1 +1 @@
{"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "5b8462a3f82d188c", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "66b69189c0acc3ff", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "70b30e0a139bbd59", "46": "58f3596265f5f902", "47": "f89e819a4cc2d299", "48": "b45478cb1443df68", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6", "80": "6ff8a97e63c9a2f3", "81": "a38f84ae3d895236", "82": "08e48bd11f9714df", "83": "5095122914e83cf5", "84": "1aef305bd7d7d63f", "85": "f8bfd0cfe9e8b478", "86": "410d15a346bd5894", "87": "6b41d288cfd834ab", "88": "5aa6db96312a8811", "89": "80225792bb62ba04", "90": "fd291228c3311f40", "91": "d4579c5b7aa2742a", "92": "7b9ba7c3bff11361", "93": "71cd9c1fa4a857f0", "94": "34cd980be3c32d21", "95": "970093453f3b7d90", "96": "9e96780a2b7c4bd6", "97": "a4e57776e229c794", "98": "089ea6a55861c693", "99": "8968e9e7d55afcbe", "100": "daa6156e559cda08", "101": "14e5b6323d34c337", "102": "6aa00d5a83295f11", "103": "f58668f5b10ccdeb", "104": "4ec787414cc6f50b", "105": "de672873b6dec6e2", "106": "edcd5d9bb3c4b00f", "107": "37f2f47110fe3eaa", "108": "b7ad5abb1da8cf8d", "109": "cb48a9c4f54efa38", "110": "f6dd36fd7f3edbe5", "111": "2925b620f0b1fd17", "112": "d8b3099917c3b711", "113": "3bb61caa0302c804", "114": "0d4f1d08dd056bb9"} {"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "5b8462a3f82d188c", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "66b69189c0acc3ff", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "1880db165e83768c", "46": "58f3596265f5f902", "47": "4e9a638bf1f93e1b", "48": "cc757e9c9987201b", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6", "80": "6ff8a97e63c9a2f3", "81": "a38f84ae3d895236", "82": "08e48bd11f9714df", "83": "5095122914e83cf5", "84": "1aef305bd7d7d63f", "85": "f8bfd0cfe9e8b478", "86": "410d15a346bd5894", "87": "6b41d288cfd834ab", "88": "5aa6db96312a8811", "89": "80225792bb62ba04", "90": "fd291228c3311f40", "91": "d4579c5b7aa2742a", "92": "7b9ba7c3bff11361", "93": "71cd9c1fa4a857f0", "94": "34cd980be3c32d21", "95": "970093453f3b7d90", "96": "9e96780a2b7c4bd6", "97": "e091504212d41fa4", "98": "089ea6a55861c693", "99": "8968e9e7d55afcbe", "100": "a7b5a49d77797f1b", "101": "21a45c852f9aeb96", "102": "6aa00d5a83295f11", "103": "f58668f5b10ccdeb", "104": "4ec787414cc6f50b", "105": "b42f25c7ec3542ab", "106": "edcd5d9bb3c4b00f", "107": "37f2f47110fe3eaa", "108": "b7ad5abb1da8cf8d", "109": "cb48a9c4f54efa38", "110": "f6dd36fd7f3edbe5", "111": "2925b620f0b1fd17", "112": "d8b3099917c3b711", "113": "3bb61caa0302c804", "114": "0d4f1d08dd056bb9", "115": "4ac2dcddeec2ff11", "116": "0bd2e5150cdcc835", "117": "196f63e0c4536d30", "118": "ade84262e3cfac12", "119": "6adb71671a1f363f", "120": "3d4cabafb8ebe1b7", "122": "80cc7c4359559d19", "123": "96618c9a362af46c", "124": "83f104cbb62fd03e", "125": "6db738fb27190349", "126": "6a087a22cbcef972", "127": "85fd71a0cad8d3a5", "128": "22dd4feed96c4229", "129": "c4d2f60f532e6f16", "130": "f6b0aa8a1568926b", "131": "56747bad6345d66b"}
+21 -5
View File
@@ -44,10 +44,10 @@
"42": "1. Input Schemas", "42": "1. Input Schemas",
"43": "2. Basic CLI Usage Examples", "43": "2. Basic CLI Usage Examples",
"44": "2. Standard Streams & Exit Codes", "44": "2. Standard Streams & Exit Codes",
"45": "LocalEmbeddingsAdapter", "45": "ClassificationResult",
"46": "InherenceClassifier", "46": "InherenceClassifier",
"47": "detect_language", "47": "detect_language",
"48": "classifier.py", "48": "models.py",
"49": "content_northvolt_de.md", "49": "content_northvolt_de.md",
"50": "content_presal_pt.md", "50": "content_presal_pt.md",
"51": "content_tangential_es.md", "51": "content_tangential_es.md",
@@ -99,12 +99,12 @@
"97": "🧠 TextNLPClassifierApp", "97": "🧠 TextNLPClassifierApp",
"98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor", "98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor",
"99": "sample_rss_xml", "99": "sample_rss_xml",
"100": "models.py", "100": "ClassificationError",
"101": "ECPSnapshot", "101": "ECPSnapshot",
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)", "102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
"103": "4. Requisitos Funcionais (FR)", "103": "4. Requisitos Funcionais (FR)",
"104": "Tasks: Article Content Multi-Engine Extractor", "104": "Tasks: Article Content Multi-Engine Extractor",
"105": "ClassificationResult", "105": "classifier.py",
"106": "Implementation Plan: Article Content Multi-Engine Extractor", "106": "Implementation Plan: Article Content Multi-Engine Extractor",
"107": "2. Cenários de Validação", "107": "2. Cenários de Validação",
"108": "1. Technical Decisions & Tradeoffs", "108": "1. Technical Decisions & Tradeoffs",
@@ -113,5 +113,21 @@
"111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier", "111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier",
"112": "CLI Contract: Article Content Multi-Engine Extractor", "112": "CLI Contract: Article Content Multi-Engine Extractor",
"113": "001-multilingual-entity-classifier/spec.md", "113": "001-multilingual-entity-classifier/spec.md",
"114": "JSON Schema Contract: Article Content Multi-Engine Extractor" "114": "JSON Schema Contract: Article Content Multi-Engine Extractor",
"115": "PRD — Seleção determinística da biblioteca de extração de conteúdo",
"116": "select_article_extractor.py",
"117": "1. Text Normalization Pipeline",
"118": "Tasks: Deterministic Article Content Selection",
"119": "select_article_extractor",
"120": "process_batch",
"122": "test_select_article_extractor.py",
"123": "Feature Specification: Deterministic Content Selection",
"124": "2. Entity Descriptions & Fields",
"125": "Implementation Plan: Deterministic Article Content Selection",
"126": "Deterministic Content Selection Checklist: End-to-End Requirements Quality",
"127": "Quickstart: Deterministic Article Content Selection",
"128": "Specification Quality Checklist: Deterministic Content Selection",
"129": "CLI Interface Contract: Deterministic Article Content Selection",
"130": "004-deterministic-content-selection/spec.md",
"131": "JSON Schema Contract: Deterministic Article Content Selection"
} }
+119 -44
View File
@@ -1,16 +1,16 @@
# Graph Report - TextNLPClassifierApp (2026-08-20) # Graph Report - TextNLPClassifierApp (2026-08-20)
## Corpus Check ## Corpus Check
- 161 files · ~78,476 words - 174 files · ~91,254 words
- Verdict: corpus is large enough that graph structure adds value. - Verdict: corpus is large enough that graph structure adds value.
## Summary ## Summary
- 1000 nodes · 1226 edges · 115 communities (77 shown, 38 thin omitted) - 1233 nodes · 1546 edges · 131 communities (93 shown, 38 thin omitted)
- Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 39 edges (avg confidence: 0.95) - Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 51 edges (avg confidence: 0.95)
- Token cost: 0 input · 0 output - Token cost: 0 input · 0 output
## Graph Freshness ## Graph Freshness
- Built from commit: `6e3d5761` - Built from commit: `6a45368c`
- Run `git rev-parse HEAD` and compare to check if the graph is stale. - Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost). - Run `graphify update .` after code changes (no API cost).
@@ -56,10 +56,10 @@
- 1. Input Schemas - 1. Input Schemas
- 2. Basic CLI Usage Examples - 2. Basic CLI Usage Examples
- 2. Standard Streams & Exit Codes - 2. Standard Streams & Exit Codes
- LocalEmbeddingsAdapter - ClassificationResult
- InherenceClassifier - InherenceClassifier
- detect_language - detect_language
- classifier.py - models.py
- content_northvolt_de.md - content_northvolt_de.md
- content_presal_pt.md - content_presal_pt.md
- content_tangential_es.md - content_tangential_es.md
@@ -110,12 +110,12 @@
- 🧠 TextNLPClassifierApp - 🧠 TextNLPClassifierApp
- Extraction Pipeline Checklist: Article Content Multi-Engine Extractor - Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
- sample_rss_xml - sample_rss_xml
- models.py - ClassificationError
- ECPSnapshot - ECPSnapshot
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC) - Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- 4. Requisitos Funcionais (FR) - 4. Requisitos Funcionais (FR)
- Tasks: Article Content Multi-Engine Extractor - Tasks: Article Content Multi-Engine Extractor
- ClassificationResult - classifier.py
- Implementation Plan: Article Content Multi-Engine Extractor - Implementation Plan: Article Content Multi-Engine Extractor
- 2. Cenários de Validação - 2. Cenários de Validação
- 1. Technical Decisions & Tradeoffs - 1. Technical Decisions & Tradeoffs
@@ -124,18 +124,33 @@
- Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier - Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
- CLI Contract: Article Content Multi-Engine Extractor - CLI Contract: Article Content Multi-Engine Extractor
- JSON Schema Contract: Article Content Multi-Engine Extractor - JSON Schema Contract: Article Content Multi-Engine Extractor
- PRD — Seleção determinística da biblioteca de extração de conteúdo
- select_article_extractor.py
- 1. Text Normalization Pipeline
- Tasks: Deterministic Article Content Selection
- select_article_extractor
- process_batch
- test_select_article_extractor.py
- Feature Specification: Deterministic Content Selection
- 2. Entity Descriptions & Fields
- Implementation Plan: Deterministic Article Content Selection
- Deterministic Content Selection Checklist: End-to-End Requirements Quality
- Quickstart: Deterministic Article Content Selection
- Specification Quality Checklist: Deterministic Content Selection
- CLI Interface Contract: Deterministic Article Content Selection
- JSON Schema Contract: Deterministic Article Content Selection
## God Nodes (most connected - your core abstractions) ## God Nodes (most connected - your core abstractions)
1. `ECPSnapshot` - 31 edges 1. `ECPSnapshot` - 31 edges
2. `InherenceClassifier` - 25 edges 2. `InherenceClassifier` - 25 edges
3. `DecisionCategory` - 18 edges 3. `select_article_extractor()` - 23 edges
4. `ClassificationResult` - 17 edges 4. `ExtractorName` - 21 edges
5. `process_batch()` - 15 edges 5. `DecisionCategory` - 17 edges
6. `LocalEmbeddingsAdapter` - 14 edges 6. `ClassificationResult` - 17 edges
7. `LLMFallbackAdapter` - 14 edges 7. `process_batch()` - 15 edges
8. `detect_language()` - 14 edges 8. `process_batch()` - 14 edges
9. `main()` - 13 edges 9. `LocalEmbeddingsAdapter` - 14 edges
10. `ArticleCrawler` - 13 edges 10. `LLMFallbackAdapter` - 14 edges
## Surprising Connections (you probably didn't know these) ## Surprising Connections (you probably didn't know these)
- `main()` --uses--> `ECPSnapshot` [INFERRED] - `main()` --uses--> `ECPSnapshot` [INFERRED]
@@ -146,13 +161,13 @@
tests/test_benchmark_24.py → src/classifier.py tests/test_benchmark_24.py → src/classifier.py
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED] - `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
tests/test_models.py → src/models.py tests/test_models.py → src/models.py
- `petrobras_ecp()` --uses--> `ECPSnapshot` [INFERRED] - `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
tests/test_classifier.py → src/models.py tests/test_models.py → src/models.py
## Import Cycles ## Import Cycles
- None detected. - None detected.
## Communities (115 total, 38 thin omitted) ## Communities (131 total, 38 thin omitted)
### Community 0 - "Task Planning" ### Community 0 - "Task Planning"
Cohesion: 0.07 Cohesion: 0.07
@@ -294,21 +309,21 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
Cohesion: 0.29 Cohesion: 0.29
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC) Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
### Community 45 - "LocalEmbeddingsAdapter" ### Community 45 - "ClassificationResult"
Cohesion: 0.14 Cohesion: 0.17
Nodes (8): LocalEmbeddingsAdapter, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface(), test_llm_adapter_interface() Nodes (10): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional LLM fallback adapter (Tier 3). Disabled by default. Provides fallback… (+2 more)
### Community 46 - "InherenceClassifier" ### Community 46 - "InherenceClassifier"
Cohesion: 0.12 Cohesion: 0.12
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more) Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess without --output and verify stdout is pure parseable…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC. (+19 more)
### Community 47 - "detect_language" ### Community 47 - "detect_language"
Cohesion: 0.14 Cohesion: 0.09
Nodes (21): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+13 more) Nodes (30): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+22 more)
### Community 48 - "classifier.py" ### Community 48 - "models.py"
Cohesion: 0.24 Cohesion: 0.19
Nodes (11): Core deterministic classification engine (Tier 1 core)., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+3 more) Nodes (15): emit_error(), main(), parse_args(), Namespace, ErrorCode, MatchedGraphEntity, Enum, str (+7 more)
### Community 80 - "test_extract_article_contents.py" ### Community 80 - "test_extract_article_contents.py"
Cohesion: 0.06 Cohesion: 0.06
@@ -382,13 +397,13 @@ Nodes (34): 1. Requirement Completeness, 2. Requirement Clarity & Non-Ambiguity,
Cohesion: 0.67 Cohesion: 0.67
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml() Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
### Community 100 - "models.py" ### Community 100 - "ClassificationError"
Cohesion: 0.15 Cohesion: 0.29
Nodes (17): emit_error(), main(), parse_args(), Namespace, Enum, ClassificationError, ErrorCode, MatchedGraphEntity (+9 more) Nodes (3): ClassificationError, Any, test_classification_error_serialization()
### Community 101 - "ECPSnapshot" ### Community 101 - "ECPSnapshot"
Cohesion: 0.22 Cohesion: 0.24
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more) Nodes (10): parametrize, ECPSnapshot, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), Unit tests for ECP models, schema validation, and structured error handling., test_ecp_snapshot_defaults() (+2 more)
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)" ### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14 Cohesion: 0.14
@@ -402,13 +417,13 @@ Nodes (20): 1.1 Objetivo do Produto, 1. Visão Geral e Contexto, 2. Personas e C
Cohesion: 0.11 Cohesion: 0.11
Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more) Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more)
### Community 105 - "ClassificationResult" ### Community 105 - "classifier.py"
Cohesion: 0.16 Cohesion: 0.13
Nodes (11): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional local vector embeddings adapter (Tier 2). Disabled by default.… (+3 more) Nodes (10): LocalEmbeddingsAdapter, Optional local vector embeddings adapter (Tier 2). Disabled by default.…, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Core deterministic classification engine (Tier 1 core)., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags() (+2 more)
### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor" ### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor"
Cohesion: 0.17 Cohesion: 0.17
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Article Content Multi-Engine Extractor, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Next: `/speckit-tasks`)*, Project Structure (+3 more) Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Article Content Multi-Engine Extractor, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Completed)*, Project Structure (+3 more)
### Community 107 - "2. Cenários de Validação" ### Community 107 - "2. Cenários de Validação"
Cohesion: 0.22 Cohesion: 0.22
@@ -438,25 +453,85 @@ Nodes (5): 1. Comando de Execução, 2. Argumentos e Flags, 3. Códigos de Saíd
Cohesion: 0.50 Cohesion: 0.50
Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor
### Community 115 - "PRD — Seleção determinística da biblioteca de extração de conteúdo"
Cohesion: 0.07
Nodes (29): 10. Requisitos não funcionais, 11. Critérios de aceite, 12. Casos obrigatórios de teste, 13. Definition of Done, 1. Contexto, 2. Objetivo, 3.1 Incluído, 3.2 Fora do escopo (+21 more)
### Community 116 - "select_article_extractor.py"
Cohesion: 0.16
Nodes (17): BatchProcessingResult, break_priority_tie(), calculate_consensus_metrics(), ExtractorCandidate, form_active_set(), main(), parse_args(), Namespace (+9 more)
### Community 117 - "1. Text Normalization Pipeline"
Cohesion: 0.11
Nodes (18): 1. Text Normalization Pipeline, 2. 5-Token Shingles & Consensus Metrics, 3. Regras de Decisão, Empate Técnico e Desempate Hierárquico, 4. Estratégia de I/O Não Destrutiva e Escrita Atômica, Alternatives Considered, Context, Context, Context (+10 more)
### Community 118 - "Tasks: Deterministic Article Content Selection"
Cohesion: 0.11
Nodes (18): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Incremental Delivery (Phases 4, 5 & 6), MVP First (Phases 1, 2 & 3), Parallel Opportunities (+10 more)
### Community 119 - "select_article_extractor"
Cohesion: 0.08
Nodes (37): ArticleSelectionResult, CandidateStatus, extract_candidate_data(), ExtractorName, Any, Enum, str, Extrai campo de texto, erro e calcula tokens/shingles para um motor. (+29 more)
### Community 120 - "process_batch"
Cohesion: 0.10
Nodes (26): atomic_save_json(), process_batch(), Path, Salva dados em JSON de forma atômica utilizando arquivo temporário e rename., Lê o JSON de entrada, valida a estrutura, processa todos os artigos e grava o…, Path, CT-012: A entrada já contém selected_extractor -> Recalcular e substituir…, CT-013: articles está vazio -> Gerar saída válida com articles vazio. (+18 more)
### Community 122 - "test_select_article_extractor.py"
Cohesion: 0.21
Nodes (14): generate_shingles(), normalize_text(), Executa a normalização determinística para comparação: 1. Decodificar entidades…, Gera conjunto de shingles ordenados de tamanho window_size (padrão 5). - Se…, Suíte de Testes Automatizados para o Seletor Determinístico de Extrator. Cobre…, Garante que marcação de imagem Markdown ![alt](url) seja descartada e link…, test_generate_shingles_empty(), test_generate_shingles_short_text() (+6 more)
### Community 123 - "Feature Specification: Deterministic Content Selection"
Cohesion: 0.17
Nodes (12): Assumptions, Edge Cases, Feature Specification: Deterministic Content Selection, Functional Requirements, Key Entities *(include if feature involves data)*, Measurable Outcomes, Requirements *(mandatory)*, Success Criteria *(mandatory)* (+4 more)
### Community 124 - "2. Entity Descriptions & Fields"
Cohesion: 0.18
Nodes (11): 1. Domain Entities & Value Types, 2. Entity Descriptions & Fields, 3. JSON Schema Mapping, `ArticleSelectionResult` (Dataclass), `BatchProcessingResult` (Dataclass), `CandidateStatus` (Enum), Data Model: Deterministic Content Selection, Entrada (+3 more)
### Community 125 - "Implementation Plan: Deterministic Article Content Selection"
Cohesion: 0.18
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Deterministic Article Content Selection, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Next Step via `/speckit-tasks`)*, Project Structure (+3 more)
### Community 126 - "Deterministic Content Selection Checklist: End-to-End Requirements Quality"
Cohesion: 0.25
Nodes (7): Candidate State Transitions & Resilience, Decision & Tie-Breaking Hierarchy, Deterministic Content Selection Checklist: End-to-End Requirements Quality, JSON Schema Integrity & Atomic I/O, Notes, Shingles & Consensus Metric Formulation, Text Normalization & Tokenization Quality
### Community 127 - "Quickstart: Deterministic Article Content Selection"
Cohesion: 0.29
Nodes (7): 1. Pré-requisitos, 2. Execução Rápida via CLI, 3. Execução dos Testes Automatizados, 4. Validação Programática / Uso como Módulo Python, Cenário 1: Selecionar o melhor extrator para uma extração existente, Cenário 2: Especificar caminho de saída customizado e modo verboso, Quickstart: Deterministic Article Content Selection
### Community 128 - "Specification Quality Checklist: Deterministic Content Selection"
Cohesion: 0.33
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Deterministic Content Selection
### Community 129 - "CLI Interface Contract: Deterministic Article Content Selection"
Cohesion: 0.33
Nodes (5): 1. Command Syntax, 2. Arguments and Flags, 3. Standard Streams (I/O), 4. Exit Codes, CLI Interface Contract: Deterministic Article Content Selection
### Community 131 - "JSON Schema Contract: Deterministic Article Content Selection"
Cohesion: 0.50
Nodes (3): 1. Input JSON Schema, 2. Output JSON Schema, JSON Schema Contract: Deterministic Article Content Selection
## Knowledge Gaps ## Knowledge Gaps
- **465 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+460 more) - **560 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+555 more)
These have ≤1 connection - possible missing edges or undocumented components. These have ≤1 connection - possible missing edges or undocumented components.
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes. - **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions ## Suggested Questions
_Questions this graph is uniquely positioned to answer:_ _Questions this graph is uniquely positioned to answer:_
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `models.py`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `InherenceClassifier`, `detect_language`, `classifier.py`?** - **Why does `ECPSnapshot` connect `ECPSnapshot` to `classifier.py`, `ClassificationResult`, `InherenceClassifier`, `detect_language`, `models.py`?**
_High betweenness centrality (0.005) - this node is a cross-community bridge._
- **Why does `InherenceClassifier` connect `InherenceClassifier` to `models.py`, `ECPSnapshot`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `detect_language`, `classifier.py`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._ _High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?** - **Why does `LLMFallbackAdapter` connect `classifier.py` to `ECPSnapshot`, `ClassificationResult`, `InherenceClassifier`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Why does `Research & Architectural Decisions: Deterministic Content Selection` connect `1. Text Normalization Pipeline` to `004-deterministic-content-selection/spec.md`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._ _High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?** - **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._ _`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?** - **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._ _`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
- **Are the 12 inferred relationships involving `ExtractorName` (e.g. with `test_article_1_regression_technical_tie_markdown_images()` and `test_ct_001_three_candidates_clear_winner()`) actually correct?**
_`ExtractorName` has 12 INFERRED edges - model-reasoned connections that need verification._
- **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?** - **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?**
_`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._ _`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
File diff suppressed because it is too large Load Diff
+150 -72
View File
@@ -294,15 +294,15 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"classify.py": { "classify.py": {
"mtime": 1787197110.3753626, "mtime": 1787264412.2457643,
"seen": 1787197159.9828603, "seen": 1787264519.0539427,
"ast_hash": "d09a35a5e25d42f6ecc537d5b67bef8e", "ast_hash": "e89fc4b64bd7606fc466d5338108705b",
"semantic_hash": "" "semantic_hash": ""
}, },
"pyproject.toml": { "pyproject.toml": {
"mtime": 1787234982.768605, "mtime": 1787264482.8509245,
"seen": 1787235020.6850271, "seen": 1787264519.0539448,
"ast_hash": "26f5f979658067bf447f375f044a8f21", "ast_hash": "409c1fc0d7bbda6830ccb71c6196cb2f",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/__init__.py": { "src/__init__.py": {
@@ -318,45 +318,45 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"src/adapters/base.py": { "src/adapters/base.py": {
"mtime": 1787196396.6199949, "mtime": 1787264412.2437606,
"seen": 1787196463.0953116, "seen": 1787264519.0542011,
"ast_hash": "40dc6e175c8708467748c1f42a0e064a", "ast_hash": "220635d5d74e73f58259069bf5207908",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/adapters/embeddings.py": { "src/adapters/embeddings.py": {
"mtime": 1787196402.7484262, "mtime": 1787264437.92723,
"seen": 1787196463.0953145, "seen": 1787264519.0542023,
"ast_hash": "0444823bb05720d4c4ca1662a00b2c01", "ast_hash": "7beb0ecfb7482c1fd10422da99f69abe",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/adapters/llm.py": { "src/adapters/llm.py": {
"mtime": 1787196408.180451, "mtime": 1787264412.2437606,
"seen": 1787196463.0953166, "seen": 1787264519.0542035,
"ast_hash": "3d95d98625df3fcd343f2fecb78707c5", "ast_hash": "ba2990328f5b25ac6cdfc3b57acc9d39",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/classifier.py": { "src/classifier.py": {
"mtime": 1787197085.0250447, "mtime": 1787264459.052089,
"seen": 1787197159.985785, "seen": 1787264519.0542045,
"ast_hash": "a2e70f968109d213fdab0ae81ac819ff", "ast_hash": "b8fb440374ecd66bb8bf67c70c19ed1f",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/language.py": { "src/language.py": {
"mtime": 1787196384.5545745, "mtime": 1787264437.9312274,
"seen": 1787196463.0953217, "seen": 1787264519.0542057,
"ast_hash": "985619013e58a7f55af2e955817f223c", "ast_hash": "cbaf52272ebb05372e34a05cd201c270",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/models.py": { "src/models.py": {
"mtime": 1787195855.3266282, "mtime": 1787264437.9302285,
"seen": 1787196463.0953236, "seen": 1787264519.054207,
"ast_hash": "ce73bfe85eb54f907e2052c80636ecaf", "ast_hash": "16e604c14f7d66634dc9e2ed7959dd7c",
"semantic_hash": "" "semantic_hash": ""
}, },
"src/parser.py": { "src/parser.py": {
"mtime": 1787195873.9119804, "mtime": 1787264437.9312274,
"seen": 1787196463.0953257, "seen": 1787264519.054208,
"ast_hash": "0a1d64ee7088b266125af071f85f3826", "ast_hash": "d43fca3f063534e3626f86f36a97a2cc",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/__init__.py": { "tests/__init__.py": {
@@ -366,39 +366,39 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_adapters.py": { "tests/test_adapters.py": {
"mtime": 1787196418.6328156, "mtime": 1787264437.9292288,
"seen": 1787196463.09533, "seen": 1787264519.0543096,
"ast_hash": "5caf95327ea030dbcc37a99bd9052ebd", "ast_hash": "7ed9ba6f62fef404bfc28a2e16ada647",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_benchmark_24.py": { "tests/test_benchmark_24.py": {
"mtime": 1787196325.5409796, "mtime": 1787264437.9302285,
"seen": 1787196463.0953324, "seen": 1787264519.0543122,
"ast_hash": "487d30d6de2dee529fa91b86a1d62c0a", "ast_hash": "052a51561002be035de1a79455256b67",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_classifier.py": { "tests/test_classifier.py": {
"mtime": 1787196364.7195728, "mtime": 1787264437.92723,
"seen": 1787196463.0953343, "seen": 1787264519.0543132,
"ast_hash": "4350a5b11a08d12ccb85a3f332d71e87", "ast_hash": "14da0c7a7d4c08762437ca777c086dea",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_cli.py": { "tests/test_cli.py": {
"mtime": 1787195939.567137, "mtime": 1787264437.9292288,
"seen": 1787196463.0953364, "seen": 1787264519.0543146,
"ast_hash": "90b0131d43cfc64a76e499c7a65d95de", "ast_hash": "b6f57994937363f8a1b0f2f66ea53d3a",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_language.py": { "tests/test_language.py": {
"mtime": 1787195891.3329668, "mtime": 1787264437.9292288,
"seen": 1787196463.0953388, "seen": 1787264519.0543184,
"ast_hash": "66191b43fcd7aef47c07cd20d5a8f03c", "ast_hash": "f1a09702410864434274f1b347769742",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_models.py": { "tests/test_models.py": {
"mtime": 1787195886.5009987, "mtime": 1787264437.9302285,
"seen": 1787196463.095341, "seen": 1787264519.0543194,
"ast_hash": "1fd725c8ca0f52e52f56c7223e9574d0", "ast_hash": "7ec72b612bf758be8c9a3617ce18b132",
"semantic_hash": "" "semantic_hash": ""
}, },
"examples/content_northvolt_de.md": { "examples/content_northvolt_de.md": {
@@ -570,27 +570,27 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_adversarial.py": { "tests/test_adversarial.py": {
"mtime": 1787197141.2340336, "mtime": 1787264437.9242287,
"seen": 1787197159.9881344, "seen": 1787264519.0543108,
"ast_hash": "5e8daf517ee32c261617d96264ef0473", "ast_hash": "d353633ce31a624a8dcda790efec4d5e",
"semantic_hash": "" "semantic_hash": ""
}, },
"scripts/extract_google_news.py": { "scripts/extract_google_news.py": {
"mtime": 1787236453.5231855, "mtime": 1787264437.92723,
"seen": 1787236509.1147037, "seen": 1787264519.0540311,
"ast_hash": "4216de490a87378d21dab707a3166679", "ast_hash": "97d8a193901d5e2edf6359fad256b7ef",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_extract_google_news.py": { "tests/test_extract_google_news.py": {
"mtime": 1787236301.513448, "mtime": 1787264437.9312274,
"seen": 1787236353.5582018, "seen": 1787264519.054317,
"ast_hash": "1fc9a108a8abfecf1d37c8642a885e9e", "ast_hash": "7561ad8fedcf1bc6cc8ed1e86a345532",
"semantic_hash": "" "semantic_hash": ""
}, },
"docs/googlenews_extractor_guia_completo.md": { "docs/googlenews_extractor_guia_completo.md": {
"mtime": 1787231605.0830677, "mtime": 1787264437.9282286,
"seen": 1787234772.8112028, "seen": 1787264519.0578,
"ast_hash": "59db8e652000ba6087a00289a0ce0959", "ast_hash": "d7c3c64e2009dded496a08c7c91f63b9",
"semantic_hash": "" "semantic_hash": ""
}, },
"specs/002-google-news-extractor/checklists/readiness.md": { "specs/002-google-news-extractor/checklists/readiness.md": {
@@ -654,21 +654,21 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"README.md": { "README.md": {
"mtime": 1787240502.4633865, "mtime": 1787264334.4498444,
"seen": 1787240516.5990582, "seen": 1787264519.0577974,
"ast_hash": "a075f140e2dd8383bec3991f92b32b26", "ast_hash": "196db110cfde06179fb750fc75afb6fe",
"semantic_hash": "" "semantic_hash": ""
}, },
"scripts/extract_article_contents.py": { "scripts/extract_article_contents.py": {
"mtime": 1787241061.3097162, "mtime": 1787264465.3415215,
"seen": 1787241108.1811028, "seen": 1787264519.0540295,
"ast_hash": "08cd37738da1500b6912ffa1fb2fc0cc", "ast_hash": "2bea6b2a048526cb49d95fb88a001600",
"semantic_hash": "" "semantic_hash": ""
}, },
"tests/test_extract_article_contents.py": { "tests/test_extract_article_contents.py": {
"mtime": 1787240910.5980105, "mtime": 1787264437.9292288,
"seen": 1787241108.1826632, "seen": 1787264519.0543158,
"ast_hash": "98cf3b30fa6ce044d943eff381eaf8a7", "ast_hash": "56aee72480f05c714865791a935d539c",
"semantic_hash": "" "semantic_hash": ""
}, },
"docs/prd_extrator_artigos_nlp.md": { "docs/prd_extrator_artigos_nlp.md": {
@@ -708,9 +708,9 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"specs/003-article-content-extractor/plan.md": { "specs/003-article-content-extractor/plan.md": {
"mtime": 1787239601.395844, "mtime": 1787264323.707266,
"seen": 1787240516.6008782, "seen": 1787264519.0605621,
"ast_hash": "406c3ebad1d32984c1b98f8a9b7ae9f1", "ast_hash": "b2232491829a303a49e749c75f3adca2",
"semantic_hash": "" "semantic_hash": ""
}, },
"specs/003-article-content-extractor/quickstart.md": { "specs/003-article-content-extractor/quickstart.md": {
@@ -726,9 +726,9 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"specs/003-article-content-extractor/spec.md": { "specs/003-article-content-extractor/spec.md": {
"mtime": 1787239197.9064422, "mtime": 1787264319.2883234,
"seen": 1787240516.6008813, "seen": 1787264519.060745,
"ast_hash": "9fd196b7ba4d9ba7adb2cad25d094764", "ast_hash": "187b1307e28a02f0ac44895b5f9cca55",
"semantic_hash": "" "semantic_hash": ""
}, },
"specs/003-article-content-extractor/tasks.md": { "specs/003-article-content-extractor/tasks.md": {
@@ -736,5 +736,83 @@
"seen": 1787240516.6008823, "seen": 1787240516.6008823,
"ast_hash": "d95c4763728496703cd89590288a2bed", "ast_hash": "d95c4763728496703cd89590288a2bed",
"semantic_hash": "" "semantic_hash": ""
},
"scripts/select_article_extractor.py": {
"mtime": 1787273522.903319,
"seen": 1787273548.728257,
"ast_hash": "3e6938964c7e4549051cb1637207698d",
"semantic_hash": ""
},
"tests/test_select_article_extractor.py": {
"mtime": 1787274189.7777488,
"seen": 1787274207.7989414,
"ast_hash": "9f91e09d47f06140b243b6f516d14c60",
"semantic_hash": ""
},
"docs/prd_deterministic_content_selection.md": {
"mtime": 1787264674.2038286,
"seen": 1787272611.6604555,
"ast_hash": "2c421402df8b6469ce60eaa984dce666",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/checklists/deterministic-selection.md": {
"mtime": 1787272346.8692143,
"seen": 1787272611.6713927,
"ast_hash": "6d9bcda99dbd284ffcbce372102fd868",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/checklists/requirements.md": {
"mtime": 1787264728.7202728,
"seen": 1787272611.6713977,
"ast_hash": "074efe6ba551e93f7d005c5462805693",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/contracts/cli-contract.md": {
"mtime": 1787274329.2175446,
"seen": 1787274402.0944166,
"ast_hash": "2cd795eaed6198dc343424bd20673728",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/contracts/json-schema.md": {
"mtime": 1787271837.7328393,
"seen": 1787272611.6714034,
"ast_hash": "155f0bccb7bcffabcedb6ef86b8003c5",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/data-model.md": {
"mtime": 1787271825.5970395,
"seen": 1787272611.671406,
"ast_hash": "23ec4a144f4aeeed9fee259b6165bea3",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/plan.md": {
"mtime": 1787271850.1744206,
"seen": 1787272611.671409,
"ast_hash": "3bebc2a73a9632d42ccca646af592133",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/quickstart.md": {
"mtime": 1787274324.1467915,
"seen": 1787274402.0946271,
"ast_hash": "a7886ca618af43f33971213e2255b3c6",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/research.md": {
"mtime": 1787274316.8544433,
"seen": 1787274402.0946283,
"ast_hash": "73ac1637d31472e25b0a1709cd137e5c",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/spec.md": {
"mtime": 1787274311.1564865,
"seen": 1787274402.0946293,
"ast_hash": "b2c994a14871c2e6e1cf4a11db7dbeb5",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/tasks.md": {
"mtime": 1787272595.8186145,
"seen": 1787272611.6714194,
"ast_hash": "3a2617478f44c6400ed57088d2039f93",
"semantic_hash": ""
} }
} }
+121 -46
View File
@@ -1,16 +1,16 @@
# Graph Report - TextNLPClassifierApp (2026-08-20) # Graph Report - TextNLPClassifierApp (2026-08-20)
## Corpus Check ## Corpus Check
- 161 files · ~78,549 words - 174 files · ~91,961 words
- Verdict: corpus is large enough that graph structure adds value. - Verdict: corpus is large enough that graph structure adds value.
## Summary ## Summary
- 1000 nodes · 1220 edges · 115 communities (77 shown, 38 thin omitted) - 1235 nodes · 1548 edges · 131 communities (93 shown, 38 thin omitted)
- Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 39 edges (avg confidence: 0.95) - Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 51 edges (avg confidence: 0.95)
- Token cost: 0 input · 0 output - Token cost: 0 input · 0 output
## Graph Freshness ## Graph Freshness
- Built from commit: `6e3d5761` - Built from commit: `6a45368c`
- Run `git rev-parse HEAD` and compare to check if the graph is stale. - Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost). - Run `graphify update .` after code changes (no API cost).
@@ -56,10 +56,10 @@
- 1. Input Schemas - 1. Input Schemas
- 2. Basic CLI Usage Examples - 2. Basic CLI Usage Examples
- 2. Standard Streams & Exit Codes - 2. Standard Streams & Exit Codes
- classifier.py - ClassificationResult
- InherenceClassifier - InherenceClassifier
- detect_language - classifier.py
- ClassificationError - main
- content_northvolt_de.md - content_northvolt_de.md
- content_presal_pt.md - content_presal_pt.md
- content_tangential_es.md - content_tangential_es.md
@@ -110,12 +110,12 @@
- 🧠 TextNLPClassifierApp - 🧠 TextNLPClassifierApp
- Extraction Pipeline Checklist: Article Content Multi-Engine Extractor - Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
- sample_rss_xml - sample_rss_xml
- main - models.py
- ECPSnapshot - ECPSnapshot
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC) - Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- 4. Requisitos Funcionais (FR) - 4. Requisitos Funcionais (FR)
- Tasks: Article Content Multi-Engine Extractor - Tasks: Article Content Multi-Engine Extractor
- models.py - LocalEmbeddingsAdapter
- Implementation Plan: Article Content Multi-Engine Extractor - Implementation Plan: Article Content Multi-Engine Extractor
- 2. Cenários de Validação - 2. Cenários de Validação
- 1. Technical Decisions & Tradeoffs - 1. Technical Decisions & Tradeoffs
@@ -124,35 +124,50 @@
- Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier - Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
- CLI Contract: Article Content Multi-Engine Extractor - CLI Contract: Article Content Multi-Engine Extractor
- JSON Schema Contract: Article Content Multi-Engine Extractor - JSON Schema Contract: Article Content Multi-Engine Extractor
- PRD — Seleção determinística da biblioteca de extração de conteúdo
- select_article_extractor.py
- 1. Text Normalization Pipeline
- Tasks: Deterministic Article Content Selection
- select_article_extractor
- process_batch
- test_select_article_extractor.py
- Feature Specification: Deterministic Content Selection
- 2. Entity Descriptions & Fields
- Implementation Plan: Deterministic Article Content Selection
- Deterministic Content Selection Checklist: End-to-End Requirements Quality
- Quickstart: Deterministic Article Content Selection
- Specification Quality Checklist: Deterministic Content Selection
- CLI Interface Contract: Deterministic Article Content Selection
- JSON Schema Contract: Deterministic Article Content Selection
## God Nodes (most connected - your core abstractions) ## God Nodes (most connected - your core abstractions)
1. `ECPSnapshot` - 31 edges 1. `ECPSnapshot` - 31 edges
2. `InherenceClassifier` - 25 edges 2. `InherenceClassifier` - 25 edges
3. `DecisionCategory` - 17 edges 3. `select_article_extractor()` - 23 edges
4. `ClassificationResult` - 17 edges 4. `ExtractorName` - 21 edges
5. `process_batch()` - 15 edges 5. `DecisionCategory` - 17 edges
6. `LocalEmbeddingsAdapter` - 14 edges 6. `ClassificationResult` - 17 edges
7. `LLMFallbackAdapter` - 14 edges 7. `process_batch()` - 15 edges
8. `detect_language()` - 14 edges 8. `process_batch()` - 14 edges
9. `main()` - 13 edges 9. `LocalEmbeddingsAdapter` - 14 edges
10. `ArticleCrawler` - 13 edges 10. `LLMFallbackAdapter` - 14 edges
## Surprising Connections (you probably didn't know these) ## Surprising Connections (you probably didn't know these)
- `main()` --uses--> `ECPSnapshot` [INFERRED] - `main()` --uses--> `ECPSnapshot` [INFERRED]
classify.py → src/models.py classify.py → src/models.py
- `main()` --uses--> `ErrorCode` [INFERRED]
classify.py → src/models.py
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED] - `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
tests/test_extract_google_news.py → scripts/extract_google_news.py tests/test_extract_google_news.py → scripts/extract_google_news.py
- `test_llm_adapter_interface()` --calls--> `LLMFallbackAdapter` [EXTRACTED]
tests/test_adapters.py → src/adapters/llm.py
- `classifier()` --uses--> `InherenceClassifier` [INFERRED] - `classifier()` --uses--> `InherenceClassifier` [INFERRED]
tests/test_benchmark_24.py → src/classifier.py tests/test_benchmark_24.py → src/classifier.py
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
tests/test_models.py → src/models.py
- `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
tests/test_models.py → src/models.py
## Import Cycles ## Import Cycles
- None detected. - None detected.
## Communities (115 total, 38 thin omitted) ## Communities (131 total, 38 thin omitted)
### Community 0 - "Task Planning" ### Community 0 - "Task Planning"
Cohesion: 0.07 Cohesion: 0.07
@@ -294,21 +309,21 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
Cohesion: 0.29 Cohesion: 0.29
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC) Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
### Community 45 - "classifier.py" ### Community 45 - "ClassificationResult"
Cohesion: 0.14 Cohesion: 0.14
Nodes (9): LocalEmbeddingsAdapter, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Core deterministic classification engine (Tier 1 core)., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface() (+1 more) Nodes (12): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., LLMFallbackAdapter (+4 more)
### Community 46 - "InherenceClassifier" ### Community 46 - "InherenceClassifier"
Cohesion: 0.12 Cohesion: 0.12
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess without --output and verify stdout is pure parseable…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC. (+19 more) Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess without --output and verify stdout is pure parseable…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC. (+19 more)
### Community 47 - "detect_language" ### Community 47 - "classifier.py"
Cohesion: 0.09 Cohesion: 0.10
Nodes (30): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+22 more) Nodes (31): count_phrase_occurrences(), match_phrase_in_text(), Core deterministic classification engine (Tier 1 core)., Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words() (+23 more)
### Community 48 - "ClassificationError" ### Community 48 - "main"
Cohesion: 0.29 Cohesion: 0.31
Nodes (3): ClassificationError, Any, test_classification_error_serialization() Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
### Community 80 - "test_extract_article_contents.py" ### Community 80 - "test_extract_article_contents.py"
Cohesion: 0.06 Cohesion: 0.06
@@ -372,7 +387,7 @@ Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Proto
### Community 97 - "🧠 TextNLPClassifierApp" ### Community 97 - "🧠 TextNLPClassifierApp"
Cohesion: 0.06 Cohesion: 0.06
Nodes (33): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. Extração Completa Automática, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Amostragem Rápida (Limit 2 Notícias), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências (+25 more) Nodes (35): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. Execução Padrão Automática, 2. Execução com Modo Verboso, 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências, 3. Baixar Binários do Navegador Stealth (Camoufox), 3. 📄 Extrator e Parser Multimotor de Artigos (+27 more)
### Community 98 - "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor" ### Community 98 - "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor"
Cohesion: 0.05 Cohesion: 0.05
@@ -382,13 +397,13 @@ Nodes (34): 1. Requirement Completeness, 2. Requirement Clarity & Non-Ambiguity,
Cohesion: 0.67 Cohesion: 0.67
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml() Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
### Community 100 - "main" ### Community 100 - "models.py"
Cohesion: 0.23 Cohesion: 0.24
Nodes (13): emit_error(), main(), parse_args(), Namespace, Enum, ErrorCode, str, CLI execution tests covering flags, arguments, stdout, and error handling. (+5 more) Nodes (8): emit_error(), ClassificationError, ErrorCode, MatchedGraphEntity, Enum, str, Data models and validation schemas for Multilingual NLP Entity Inherence…, test_classification_error_serialization()
### Community 101 - "ECPSnapshot" ### Community 101 - "ECPSnapshot"
Cohesion: 0.22 Cohesion: 0.19
Nodes (11): parametrize, ECPSnapshot, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), Unit tests for ECP models, schema validation, and structured error handling., test_classification_result_serialization() (+3 more) Nodes (11): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), Unit tests for ECP models, schema validation, and structured error handling. (+3 more)
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)" ### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14 Cohesion: 0.14
@@ -402,9 +417,9 @@ Nodes (20): 1.1 Objetivo do Produto, 1. Visão Geral e Contexto, 2. Personas e C
Cohesion: 0.11 Cohesion: 0.11
Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more) Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more)
### Community 105 - "models.py" ### Community 105 - "LocalEmbeddingsAdapter"
Cohesion: 0.16 Cohesion: 0.18
Nodes (12): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional local vector embeddings adapter (Tier 2). Disabled by default.… (+4 more) Nodes (7): LocalEmbeddingsAdapter, Optional local vector embeddings adapter (Tier 2). Disabled by default.…, Optional adapter for local multilingual semantic vector embeddings., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface(), test_llm_adapter_interface()
### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor" ### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor"
Cohesion: 0.17 Cohesion: 0.17
@@ -438,25 +453,85 @@ Nodes (5): 1. Comando de Execução, 2. Argumentos e Flags, 3. Códigos de Saíd
Cohesion: 0.50 Cohesion: 0.50
Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor
### Community 115 - "PRD — Seleção determinística da biblioteca de extração de conteúdo"
Cohesion: 0.07
Nodes (29): 10. Requisitos não funcionais, 11. Critérios de aceite, 12. Casos obrigatórios de teste, 13. Definition of Done, 1. Contexto, 2. Objetivo, 3.1 Incluído, 3.2 Fora do escopo (+21 more)
### Community 116 - "select_article_extractor.py"
Cohesion: 0.16
Nodes (17): BatchProcessingResult, break_priority_tie(), calculate_consensus_metrics(), ExtractorCandidate, form_active_set(), main(), parse_args(), Namespace (+9 more)
### Community 117 - "1. Text Normalization Pipeline"
Cohesion: 0.11
Nodes (18): 1. Text Normalization Pipeline, 2. 5-Token Shingles & Consensus Metrics, 3. Regras de Decisão, Empate Técnico e Desempate Hierárquico, 4. Estratégia de I/O Não Destrutiva e Escrita Atômica, Alternatives Considered, Context, Context, Context (+10 more)
### Community 118 - "Tasks: Deterministic Article Content Selection"
Cohesion: 0.11
Nodes (18): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Incremental Delivery (Phases 4, 5 & 6), MVP First (Phases 1, 2 & 3), Parallel Opportunities (+10 more)
### Community 119 - "select_article_extractor"
Cohesion: 0.08
Nodes (37): ArticleSelectionResult, CandidateStatus, extract_candidate_data(), ExtractorName, Any, Enum, str, Extrai campo de texto, erro e calcula tokens/shingles para um motor. (+29 more)
### Community 120 - "process_batch"
Cohesion: 0.10
Nodes (26): atomic_save_json(), process_batch(), Path, Salva dados em JSON de forma atômica utilizando arquivo temporário e rename., Lê o JSON de entrada, valida a estrutura, processa todos os artigos e grava o…, Path, CT-012: A entrada já contém selected_extractor -> Recalcular e substituir…, CT-013: articles está vazio -> Gerar saída válida com articles vazio. (+18 more)
### Community 122 - "test_select_article_extractor.py"
Cohesion: 0.21
Nodes (14): generate_shingles(), normalize_text(), Executa a normalização determinística para comparação: 1. Decodificar entidades…, Gera conjunto de shingles ordenados de tamanho window_size (padrão 5). - Se…, Suíte de Testes Automatizados para o Seletor Determinístico de Extrator. Cobre…, Garante que marcação de imagem Markdown ![alt](url) seja descartada e link…, test_generate_shingles_empty(), test_generate_shingles_short_text() (+6 more)
### Community 123 - "Feature Specification: Deterministic Content Selection"
Cohesion: 0.17
Nodes (12): Assumptions, Edge Cases, Feature Specification: Deterministic Content Selection, Functional Requirements, Key Entities *(include if feature involves data)*, Measurable Outcomes, Requirements *(mandatory)*, Success Criteria *(mandatory)* (+4 more)
### Community 124 - "2. Entity Descriptions & Fields"
Cohesion: 0.18
Nodes (11): 1. Domain Entities & Value Types, 2. Entity Descriptions & Fields, 3. JSON Schema Mapping, `ArticleSelectionResult` (Dataclass), `BatchProcessingResult` (Dataclass), `CandidateStatus` (Enum), Data Model: Deterministic Content Selection, Entrada (+3 more)
### Community 125 - "Implementation Plan: Deterministic Article Content Selection"
Cohesion: 0.18
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Deterministic Article Content Selection, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Next Step via `/speckit-tasks`)*, Project Structure (+3 more)
### Community 126 - "Deterministic Content Selection Checklist: End-to-End Requirements Quality"
Cohesion: 0.25
Nodes (7): Candidate State Transitions & Resilience, Decision & Tie-Breaking Hierarchy, Deterministic Content Selection Checklist: End-to-End Requirements Quality, JSON Schema Integrity & Atomic I/O, Notes, Shingles & Consensus Metric Formulation, Text Normalization & Tokenization Quality
### Community 127 - "Quickstart: Deterministic Article Content Selection"
Cohesion: 0.29
Nodes (7): 1. Pré-requisitos, 2. Execução Rápida via CLI, 3. Execução dos Testes Automatizados, 4. Validação Programática / Uso como Módulo Python, Cenário 1: Selecionar o melhor extrator para uma extração existente, Cenário 2: Especificar caminho de saída customizado e modo verboso, Quickstart: Deterministic Article Content Selection
### Community 128 - "Specification Quality Checklist: Deterministic Content Selection"
Cohesion: 0.33
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Deterministic Content Selection
### Community 129 - "CLI Interface Contract: Deterministic Article Content Selection"
Cohesion: 0.33
Nodes (5): 1. Command Syntax, 2. Arguments and Flags, 3. Standard Streams (I/O), 4. Exit Codes, CLI Interface Contract: Deterministic Article Content Selection
### Community 131 - "JSON Schema Contract: Deterministic Article Content Selection"
Cohesion: 0.50
Nodes (3): 1. Input JSON Schema, 2. Output JSON Schema, JSON Schema Contract: Deterministic Article Content Selection
## Knowledge Gaps ## Knowledge Gaps
- **465 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+460 more) - **562 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+557 more)
These have ≤1 connection - possible missing edges or undocumented components. These have ≤1 connection - possible missing edges or undocumented components.
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes. - **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions ## Suggested Questions
_Questions this graph is uniquely positioned to answer:_ _Questions this graph is uniquely positioned to answer:_
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `models.py`, `classifier.py`, `InherenceClassifier`, `detect_language`?** - **Why does `ECPSnapshot` connect `ECPSnapshot` to `models.py`, `LocalEmbeddingsAdapter`, `ClassificationResult`, `InherenceClassifier`, `classifier.py`, `main`?**
_High betweenness centrality (0.006) - this node is a cross-community bridge._
- **Why does `InherenceClassifier` connect `InherenceClassifier` to `main`, `ECPSnapshot`, `models.py`, `classifier.py`, `detect_language`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._ _High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?** - **Why does `LLMFallbackAdapter` connect `ClassificationResult` to `LocalEmbeddingsAdapter`, `ECPSnapshot`, `InherenceClassifier`, `classifier.py`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Why does `Research & Architectural Decisions: Deterministic Content Selection` connect `1. Text Normalization Pipeline` to `004-deterministic-content-selection/spec.md`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._ _High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?** - **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._ _`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?** - **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._ _`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
- **Are the 12 inferred relationships involving `ExtractorName` (e.g. with `test_article_1_regression_technical_tie_markdown_images()` and `test_ct_001_three_candidates_clear_winner()`) actually correct?**
_`ExtractorName` has 12 INFERRED edges - model-reasoned connections that need verification._
- **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?** - **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?**
_`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._ _`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"nodes": [{"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "label": "cli-contract.md", "file_type": "document", "node_kind": "page", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "label": "CLI Interface Contract: Deterministic Article Content Selection", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_1_command_syntax", "label": "1. Command Syntax", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_2_arguments_and_flags", "label": "2. Arguments and Flags", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L15"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_3_standard_streams_i_o", "label": "3. Standard Streams (I/O)", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L26"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_4_exit_codes", "label": "4. Exit Codes", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L48"}], "edges": [{"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L3", "weight": 1.0, "target_file": "$graphify-root$/specs/004-deterministic-content-selection/spec.md"}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_1_command_syntax", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_2_arguments_and_flags", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L15", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_3_standard_streams_i_o", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L26", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_4_exit_codes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L48", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"nodes": [{"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "label": "cli-contract.md", "file_type": "document", "node_kind": "page", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "label": "CLI Interface Contract: Deterministic Article Content Selection", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_1_command_syntax", "label": "1. Command Syntax", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_2_arguments_and_flags", "label": "2. Arguments and Flags", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L15"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_3_standard_streams_i_o", "label": "3. Standard Streams (I/O)", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L26"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_4_exit_codes", "label": "4. Exit Codes", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L48"}], "edges": [{"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L3", "weight": 1.0, "target_file": "$graphify-root$/specs/004-deterministic-content-selection/spec.md"}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_1_command_syntax", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_2_arguments_and_flags", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L15", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_3_standard_streams_i_o", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L26", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_cli_interface_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_cli_contract_4_exit_codes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/cli-contract.md", "source_location": "L48", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"nodes": [{"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_md", "label": "requirements.md", "file_type": "document", "node_kind": "page", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "label": "Specification Quality Checklist: Deterministic Content Selection", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_content_quality", "label": "Content Quality", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_requirement_completeness", "label": "Requirement Completeness", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L14"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_feature_readiness", "label": "Feature Readiness", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L25"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_notes", "label": "Notes", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L32"}], "edges": [{"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L5", "weight": 1.0, "target_file": "$graphify-root$/specs/004-deterministic-content-selection/spec.md"}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_content_quality", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_requirement_completeness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L14", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_feature_readiness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L25", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_specification_quality_checklist_deterministic_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_checklists_requirements_notes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/checklists/requirements.md", "source_location": "L32", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"nodes": [{"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_md", "label": "json-schema.md", "file_type": "document", "node_kind": "page", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_json_schema_contract_deterministic_article_content_selection", "label": "JSON Schema Contract: Deterministic Article Content Selection", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_1_input_json_schema", "label": "1. Input JSON Schema", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_2_output_json_schema", "label": "2. Output JSON Schema", "file_type": "document", "node_kind": "heading", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L52"}], "edges": [{"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_json_schema_contract_deterministic_article_content_selection", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_md", "target": "$graphify-root$_specs_004_deterministic_content_selection_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L3", "weight": 1.0, "target_file": "$graphify-root$/specs/004-deterministic-content-selection/spec.md"}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_json_schema_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_1_input_json_schema", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_json_schema_contract_deterministic_article_content_selection", "target": "$graphify-root$_specs_004_deterministic_content_selection_contracts_json_schema_2_output_json_schema", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/004-deterministic-content-selection/contracts/json-schema.md", "source_location": "L52", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+1 -1
View File
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+7053 -585
View File
File diff suppressed because it is too large Load Diff
+81 -3
View File
@@ -654,9 +654,9 @@
"semantic_hash": "" "semantic_hash": ""
}, },
"README.md": { "README.md": {
"mtime": 1787264334.4498444, "mtime": 1787274445.4631622,
"seen": 1787264519.0577974, "seen": 1787274455.3244681,
"ast_hash": "196db110cfde06179fb750fc75afb6fe", "ast_hash": "5270e11cfba31a235ba379effb2db0d5",
"semantic_hash": "" "semantic_hash": ""
}, },
"scripts/extract_article_contents.py": { "scripts/extract_article_contents.py": {
@@ -736,5 +736,83 @@
"seen": 1787240516.6008823, "seen": 1787240516.6008823,
"ast_hash": "d95c4763728496703cd89590288a2bed", "ast_hash": "d95c4763728496703cd89590288a2bed",
"semantic_hash": "" "semantic_hash": ""
},
"scripts/select_article_extractor.py": {
"mtime": 1787273522.903319,
"seen": 1787273548.728257,
"ast_hash": "3e6938964c7e4549051cb1637207698d",
"semantic_hash": ""
},
"tests/test_select_article_extractor.py": {
"mtime": 1787274189.7777488,
"seen": 1787274207.7989414,
"ast_hash": "9f91e09d47f06140b243b6f516d14c60",
"semantic_hash": ""
},
"docs/prd_deterministic_content_selection.md": {
"mtime": 1787264674.2038286,
"seen": 1787272611.6604555,
"ast_hash": "2c421402df8b6469ce60eaa984dce666",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/checklists/deterministic-selection.md": {
"mtime": 1787272346.8692143,
"seen": 1787272611.6713927,
"ast_hash": "6d9bcda99dbd284ffcbce372102fd868",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/checklists/requirements.md": {
"mtime": 1787264728.7202728,
"seen": 1787272611.6713977,
"ast_hash": "074efe6ba551e93f7d005c5462805693",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/contracts/cli-contract.md": {
"mtime": 1787274329.2175446,
"seen": 1787274402.0944166,
"ast_hash": "2cd795eaed6198dc343424bd20673728",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/contracts/json-schema.md": {
"mtime": 1787271837.7328393,
"seen": 1787272611.6714034,
"ast_hash": "155f0bccb7bcffabcedb6ef86b8003c5",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/data-model.md": {
"mtime": 1787271825.5970395,
"seen": 1787272611.671406,
"ast_hash": "23ec4a144f4aeeed9fee259b6165bea3",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/plan.md": {
"mtime": 1787271850.1744206,
"seen": 1787272611.671409,
"ast_hash": "3bebc2a73a9632d42ccca646af592133",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/quickstart.md": {
"mtime": 1787274324.1467915,
"seen": 1787274402.0946271,
"ast_hash": "a7886ca618af43f33971213e2255b3c6",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/research.md": {
"mtime": 1787274316.8544433,
"seen": 1787274402.0946283,
"ast_hash": "73ac1637d31472e25b0a1709cd137e5c",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/spec.md": {
"mtime": 1787274311.1564865,
"seen": 1787274402.0946293,
"ast_hash": "b2c994a14871c2e6e1cf4a11db7dbeb5",
"semantic_hash": ""
},
"specs/004-deterministic-content-selection/tasks.md": {
"mtime": 1787272595.8186145,
"seen": 1787272611.6714194,
"ast_hash": "3a2617478f44c6400ed57088d2039f93",
"semantic_hash": ""
} }
} }
+603
View File
@@ -0,0 +1,603 @@
#!/usr/bin/env python3
"""
Seletor Determinístico de Extrator de Conteúdo de Artigos.
Compara deterministicamente as saídas de Trafilatura, Newspaper4k e Readability
a partir de um arquivo JSON consolidado de extrações, calculando métricas de consenso
(shingles de 5-tokens, cobertura, suporte, F1-score) e aplicando regras rígidas de
desempate técnico e hierárquico, enriquecendo o JSON exclusivamente com a chave
`selected_extractor` de forma não-destrutiva e atômica.
"""
from __future__ import annotations
import argparse
import html
import json
import os
import re
import sys
import tempfile
import unicodedata
from collections import Counter
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any
# ==============================================================================
# Modelos de Dados e Enumerações
# ==============================================================================
class ExtractorName(str, Enum):
"""Nomes canônicos e catálogo fechado dos motores de extração."""
TRAFILATURA = "trafilatura"
NEWSPAPER4K = "newspaper4k"
READABILITY = "readability"
class CandidateStatus(str, Enum):
"""Estado de viabilidade do candidato para formação do conjunto ativo."""
USABLE = "usable"
DEGRADED = "degraded"
UNAVAILABLE = "unavailable"
# Ordem estrita de prioridade de desempate final (PRD §7.1 item 4, §7.5 item 6, §7.6 item 4)
FALLBACK_PRIORITY: list[ExtractorName] = [
ExtractorName.NEWSPAPER4K,
ExtractorName.READABILITY,
ExtractorName.TRAFILATURA,
]
# Margem de empate técnico (PRD §7.5 item 3)
TECHNICAL_TIE_THRESHOLD: float = 0.03
FLOAT_EPSILON: float = 1e-9
@dataclass
class ExtractorCandidate:
"""Representação e métricas de um candidato a extrator em um artigo."""
name: ExtractorName
raw_text: str | None = None
error: str | None = None
status: CandidateStatus = CandidateStatus.UNAVAILABLE
tokens: list[str] = field(default_factory=list)
shingles: set[tuple[str, ...]] = field(default_factory=set)
shingle_count: int = 0
coverage: float = 0.0
support: float = 0.0
score: float = 0.0
@dataclass
class ArticleSelectionResult:
"""Resultado detalhado da seleção para um artigo individual."""
article_index: int
selected_extractor: ExtractorName
selection_reason: str
active_candidates_count: int
consensus_shingles_count: int
candidates: dict[ExtractorName, ExtractorCandidate] = field(default_factory=dict)
@dataclass
class BatchProcessingResult:
"""Resultado consolidado do processamento em lote."""
total_articles: int
processed_count: int
selection_distribution: dict[str, int]
input_file: str
output_file: str
selections: list[ArticleSelectionResult] = field(default_factory=list)
# ==============================================================================
# Pipeline de Normalização Textual e Shingles
# ==============================================================================
def normalize_text(text: Any) -> list[str]:
"""
Executa a normalização determinística para comparação:
1. Decodificar entidades HTML.
2. Remover marcação HTML e Markdown, preservando o texto visível.
3. Em links, preservar o texto e remover o endereço.
4. Aplicar normalização Unicode NFKC.
5. Converter o texto para minúsculas.
6. Substituir toda sequência de espaços, tabulações ou quebras de linha por um único espaço.
7. Tokenizar mantendo letras e números Unicode.
8. Desconsiderar pontuação.
"""
if not text or not isinstance(text, str):
return []
# 1. Decodificar entidades HTML
s = html.unescape(text)
# 2. Remover imagens Markdown ![alt](url) -> '' (marcação de mídia não é texto visível)
s = re.sub(r"!\s*\[[^\]]*\]\([^)]*\)", " ", s)
# 3. Em links Markdown [texto](url), preservar texto âncora e remover endereço
s = re.sub(r"\[([^\]]+)\]\([^)]+\)", r" \1 ", s)
# 2. Remover tags HTML mantendo espaço entre palavras
s = re.sub(r"<[^>]+>", " ", s)
# 4. Normalização Unicode NFKC
s = unicodedata.normalize("NFKC", s)
# 5. Minúsculas
s = s.lower()
# 6. Colapso de múltiplos espaços em branco
s = re.sub(r"\s+", " ", s).strip()
# 7 & 8. Tokenizar mantendo letras e números Unicode, ignorando pontuações
tokens = re.findall(r"[\w]+", s, flags=re.UNICODE)
return tokens
def generate_shingles(tokens: list[str], window_size: int = 5) -> set[tuple[str, ...]]:
"""
Gera conjunto de shingles ordenados de tamanho window_size (padrão 5).
- Se len(tokens) >= window_size: todas as janelas consecutivas de 5 tokens.
- Se 1 <= len(tokens) < window_size: sequência completa como um único shingle.
- Se tokens vazio: conjunto vazio.
"""
n = len(tokens)
if n == 0:
return set()
if n < window_size:
return {tuple(tokens)}
return {tuple(tokens[i : i + window_size]) for i in range(n - window_size + 1)}
# ==============================================================================
# Classificação de Candidatos e Formação do Conjunto Ativo
# ==============================================================================
def extract_candidate_data(article_dict: dict[str, Any], name: ExtractorName) -> ExtractorCandidate:
"""Extrai campo de texto, erro e calcula tokens/shingles para um motor."""
lib_data = article_dict.get(name.value)
if not isinstance(lib_data, dict):
return ExtractorCandidate(name=name, status=CandidateStatus.UNAVAILABLE)
# Campo usado na comparação (PRD §5.2)
if name == ExtractorName.TRAFILATURA:
raw_text = lib_data.get("text")
elif name == ExtractorName.NEWSPAPER4K:
raw_text = lib_data.get("text")
elif name == ExtractorName.READABILITY:
raw_text = lib_data.get("cleaned_text")
else:
raw_text = None
raw_error = lib_data.get("error")
# Tratar erro vazio/nulo
error_val = str(raw_error) if raw_error is not None and str(raw_error).strip() else None
# Normalizar texto para obter tokens
tokens = normalize_text(raw_text)
shingles = generate_shingles(tokens)
shingle_count = len(shingles)
# Determinar status do candidato (PRD §5.3)
if not tokens or shingle_count == 0:
status = CandidateStatus.UNAVAILABLE
elif error_val is None:
status = CandidateStatus.USABLE
else:
status = CandidateStatus.DEGRADED
return ExtractorCandidate(
name=name,
raw_text=raw_text if isinstance(raw_text, str) else None,
error=error_val,
status=status,
tokens=tokens,
shingles=shingles,
shingle_count=shingle_count,
)
def form_active_set(
candidates: dict[ExtractorName, ExtractorCandidate],
) -> list[ExtractorCandidate]:
"""
Forma o conjunto ativo de candidatos conforme PRD §7.1:
1. Se existir pelo menos um utilizável, considerar somente os utilizáveis.
2. Se não existir utilizável, considerar os candidatos degradados.
3. Se não existir utilizável nem degradado, retorna lista vazia.
"""
usables = [c for c in candidates.values() if c.status == CandidateStatus.USABLE]
if usables:
return usables
degradeds = [c for c in candidates.values() if c.status == CandidateStatus.DEGRADED]
if degradeds:
return degradeds
return []
# ==============================================================================
# Cálculo de Consenso e Métricas F1
# ==============================================================================
def calculate_consensus_metrics(
active_candidates: list[ExtractorCandidate],
) -> set[tuple[str, ...]]:
"""
Constrói o conjunto de consenso (shingles presentes em >= 2 candidatos ativos)
e calcula Cobertura, Suporte e F1-score para cada candidato ativo (PRD §7.4).
"""
if len(active_candidates) < 2:
for c in active_candidates:
c.coverage = 1.0 if c.shingle_count > 0 else 0.0
c.support = 1.0 if c.shingle_count > 0 else 0.0
c.score = 1.0 if c.shingle_count > 0 else 0.0
return set()
# Contar frequência de cada shingle entre os candidatos ativos
shingle_counts: Counter[tuple[str, ...]] = Counter()
for c in active_candidates:
for s in c.shingles:
shingle_counts[s] += 1
consensus_shingles = {s for s, cnt in shingle_counts.items() if cnt >= 2}
total_consensus = len(consensus_shingles)
for c in active_candidates:
if total_consensus == 0 or c.shingle_count == 0:
c.coverage = 0.0
c.support = 0.0
c.score = 0.0
continue
common_shingles = len(c.shingles & consensus_shingles)
c.coverage = common_shingles / total_consensus
c.support = common_shingles / c.shingle_count
denominator = c.coverage + c.support
if denominator > 0:
c.score = (2.0 * c.coverage * c.support) / denominator
else:
c.score = 0.0
return consensus_shingles
# ==============================================================================
# Algoritmos de Seleção e Desempate
# ==============================================================================
def break_priority_tie(candidates: list[ExtractorCandidate]) -> ExtractorCandidate:
"""Aplica a prioridade final estrita: newspaper4k > readability > trafilatura."""
for priority_name in FALLBACK_PRIORITY:
for c in candidates:
if c.name == priority_name:
return c
return candidates[0]
def select_with_consensus(
active_candidates: list[ExtractorCandidate], consensus_shingles: set[tuple[str, ...]]
) -> tuple[ExtractorName, str]:
"""
Seleciona o melhor candidato quando existe consenso (PRD §7.5):
1. Ordenar por score decrescente.
2. Identificar candidatos no empate técnico (diferença para maior score <= 0.03).
3. Se houver 1 candidato no empate técnico, selecioná-lo.
4. Se houver empate técnico, selecionar o candidato com menor quantidade de shingles.
5. Se empatar na quantidade de shingles, aplicar prioridade final newspaper4k > readability > trafilatura.
"""
# Ordenar por score decrescente
sorted_by_score = sorted(active_candidates, key=lambda c: c.score, reverse=True)
max_score = sorted_by_score[0].score
# Grupo de empate técnico (score >= max_score - 0.03 com tolerância de precisão float)
technical_tie_pool = [
c
for c in sorted_by_score
if (max_score - c.score) <= (TECHNICAL_TIE_THRESHOLD + FLOAT_EPSILON)
]
if len(technical_tie_pool) == 1:
return technical_tie_pool[0].name, "highest_score"
# Selecionar o candidato com menor quantidade de shingles no grupo de empate
min_shingles = min(c.shingle_count for c in technical_tie_pool)
shingle_tie_pool = [c for c in technical_tie_pool if c.shingle_count == min_shingles]
if len(shingle_tie_pool) == 1:
return shingle_tie_pool[0].name, "technical_tie_smallest_shingles"
# Desempate final de prioridade
winner = break_priority_tie(shingle_tie_pool)
return winner.name, "technical_tie_priority_fallback"
def select_without_consensus(
active_candidates: list[ExtractorCandidate],
) -> tuple[ExtractorName, str]:
"""
Seleciona o candidato quando NÃO existe consenso (PRD §7.6):
- Com 3 candidatos ativos: selecionar o candidato com a quantidade mediana de shingles.
- Com 2 candidatos ativos: selecionar o candidato com a maior quantidade de shingles.
- Com 1 candidato ativo: selecionar o único candidato.
- Em empate de quantidade: aplicar prioridade newspaper4k > readability > trafilatura.
- Sem candidato ativo: selecionar newspaper4k.
"""
k = len(active_candidates)
if k == 0:
return ExtractorName.NEWSPAPER4K, "fallback_all_unavailable"
if k == 1:
return active_candidates[0].name, "single_active_candidate"
if k == 2:
c1, c2 = active_candidates[0], active_candidates[1]
if c1.shingle_count > c2.shingle_count:
return c1.name, "no_consensus_max_shingles"
elif c2.shingle_count > c1.shingle_count:
return c2.name, "no_consensus_max_shingles"
else:
winner = break_priority_tie([c1, c2])
return winner.name, "no_consensus_tie_priority_fallback"
if k == 3:
# 3 candidatos ativos -> quantidade mediana de shingles
# Ordenar por shingle_count crescente
sorted_by_shingles = sorted(active_candidates, key=lambda c: c.shingle_count)
s0, s1, s2 = (
sorted_by_shingles[0].shingle_count,
sorted_by_shingles[1].shingle_count,
sorted_by_shingles[2].shingle_count,
)
# Se todos tiverem contagens distintas (ex: 10, 20, 30), a mediana é o elemento do meio (20)
if s0 < s1 < s2:
return sorted_by_shingles[1].name, "no_consensus_median_shingles"
# Se houver empate na mediana (ex: [10, 20, 20] ou [20, 20, 30] ou [20, 20, 20])
# Os candidatos cujo shingle_count é igual ao valor mediano (s1) entram no pool de desempate
median_value = s1
median_candidates = [c for c in active_candidates if c.shingle_count == median_value]
if len(median_candidates) == 1:
return median_candidates[0].name, "no_consensus_median_shingles"
else:
winner = break_priority_tie(median_candidates)
return winner.name, "no_consensus_median_priority_fallback"
# Fallback genérico para listas maiores (se houver)
winner = break_priority_tie(active_candidates)
return winner.name, "priority_fallback"
def select_article_extractor(
article_dict: dict[str, Any], article_index: int = 0
) -> ArticleSelectionResult:
"""
Orquestrador determinístico completo para um único artigo.
Classifica candidatos, forma conjunto ativo, calcula métricas e aplica regras de decisão.
"""
candidates: dict[ExtractorName, ExtractorCandidate] = {
name: extract_candidate_data(article_dict, name) for name in ExtractorName
}
active_set = form_active_set(candidates)
if not active_set:
# Nenhum candidato utilizável nem degradado -> Fallback compulsório para newspaper4k (PRD §7.1 item 4)
return ArticleSelectionResult(
article_index=article_index,
selected_extractor=ExtractorName.NEWSPAPER4K,
selection_reason="fallback_all_unavailable",
active_candidates_count=0,
consensus_shingles_count=0,
candidates=candidates,
)
if len(active_set) == 1:
# Apenas 1 candidato ativo -> selecioná-lo imediatamente (PRD §7.1 item 5)
winner = active_set[0]
return ArticleSelectionResult(
article_index=article_index,
selected_extractor=winner.name,
selection_reason="single_usable_candidate"
if winner.status == CandidateStatus.USABLE
else "single_degraded_candidate",
active_candidates_count=1,
consensus_shingles_count=0,
candidates=candidates,
)
# Calcular consenso e métricas F1
consensus_shingles = calculate_consensus_metrics(active_set)
if consensus_shingles:
winner_name, reason = select_with_consensus(active_set, consensus_shingles)
else:
winner_name, reason = select_without_consensus(active_set)
return ArticleSelectionResult(
article_index=article_index,
selected_extractor=winner_name,
selection_reason=reason,
active_candidates_count=len(active_set),
consensus_shingles_count=len(consensus_shingles),
candidates=candidates,
)
# ==============================================================================
# Processamento em Lote e I/O Atômico
# ==============================================================================
def atomic_save_json(data: Any, target_path: Path, indent: int = 2) -> None:
"""Salva dados em JSON de forma atômica utilizando arquivo temporário e rename."""
target_path = Path(target_path).resolve()
target_path.parent.mkdir(parents=True, exist_ok=True)
temp_fd, temp_file_path = tempfile.mkstemp(
dir=target_path.parent, prefix=f".{target_path.name}.tmp_", text=True
)
try:
with open(temp_fd, "w", encoding="utf-8") as f:
if indent > 0:
json.dump(data, f, ensure_ascii=False, indent=indent)
else:
json.dump(data, f, ensure_ascii=False, separators=(",", ":"))
f.write("\n")
os.replace(temp_file_path, target_path)
except Exception:
if os.path.exists(temp_file_path):
os.remove(temp_file_path)
raise
def process_batch(
input_path: Path | str,
output_path: Path | str | None = None,
indent: int = 2,
verbose: bool = False,
) -> BatchProcessingResult:
"""
Lê o JSON de entrada, valida a estrutura, processa todos os artigos e grava o arquivo de saída.
Preserva 100% dos dados originais e a ordem dos artigos.
"""
in_file = Path(input_path).resolve()
if not in_file.is_file():
raise FileNotFoundError(f"Arquivo de entrada não encontrado: {in_file}")
try:
with open(in_file, "r", encoding="utf-8") as f:
data = json.load(f)
except json.JSONDecodeError as exc:
raise ValueError(f"JSON inválido em '{in_file}': {exc}") from exc
if not isinstance(data, dict):
raise ValueError("A raiz do JSON de entrada deve ser um objeto.")
articles = data.get("articles")
if articles is None or not isinstance(articles, list):
raise ValueError("A chave 'articles' é obrigatória e deve ser uma lista.")
if output_path is None:
# Padrão: <nome_original_sem_extensao>_selected.json
out_file = in_file.parent / f"{in_file.stem}_selected.json"
else:
out_file = Path(output_path).resolve()
selections: list[ArticleSelectionResult] = []
distribution: Counter[str] = Counter()
for idx, article in enumerate(articles):
if not isinstance(article, dict):
# Tratar artigo malformado
article = {}
articles[idx] = article
res = select_article_extractor(article, article_index=idx)
selections.append(res)
distribution[res.selected_extractor.value] += 1
# Enriquecer ou recalcular chave no artigo (PRD §6.2 e §6.3)
article["selected_extractor"] = res.selected_extractor.value
if verbose:
sys.stderr.write(
f"[Artigo #{idx + 1:03d}] Extrator: {res.selected_extractor.value:<12} | "
f"Motivo: {res.selection_reason:<32} | Ativos: {res.active_candidates_count} | "
f"Consenso: {res.consensus_shingles_count}\n"
)
# Gravar arquivo de saída atomicamente
atomic_save_json(data, out_file, indent=indent)
return BatchProcessingResult(
total_articles=len(articles),
processed_count=len(selections),
selection_distribution=dict(distribution),
input_file=str(in_file),
output_file=str(out_file),
selections=selections,
)
# ==============================================================================
# Interface CLI
# ==============================================================================
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Seletor determinístico da melhor extração de conteúdo (Trafilatura / Newspaper4k / Readability)."
)
parser.add_argument(
"input_file", type=Path, help="Caminho do arquivo JSON consolidado de extrações."
)
parser.add_argument(
"-o",
"--output",
type=Path,
default=None,
help="Caminho do arquivo JSON de saída (padrão: <nome>_selected.json).",
)
parser.add_argument(
"--indent", type=int, default=2, help="Indentação do arquivo JSON de saída (padrão: 2)."
)
parser.add_argument(
"-v",
"--verbose",
action="store_true",
help="Exibe detalhes da seleção e pontuações no stderr.",
)
return parser.parse_args(argv)
def main(argv: list[str] | None = None) -> int:
args = parse_args(argv)
try:
result = process_batch(
input_path=args.input_file,
output_path=args.output,
indent=args.indent,
verbose=args.verbose,
)
output_summary = {
"status": "success",
"input_file": result.input_file,
"output_file": result.output_file,
"total_articles": result.total_articles,
"processed_count": result.processed_count,
"distribution": result.selection_distribution,
}
sys.stdout.write(json.dumps(output_summary, indent=2, ensure_ascii=False) + "\n")
return 0
except FileNotFoundError as e:
sys.stderr.write(f"ERRO DE ARQUIVO: {e}\n")
return 1
except ValueError as e:
sys.stderr.write(f"ERRO DE VALIDAÇÃO: {e}\n")
return 2
except Exception as e:
sys.stderr.write(f"ERRO INESPERADO: {e}\n")
return 1
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,54 @@
# Deterministic Content Selection Checklist: End-to-End Requirements Quality
**Purpose**: Validate the completeness, clarity, consistency, and measurability of requirements for the deterministic extractor selection pipeline
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
---
## Text Normalization & Tokenization Quality
- [x] CHK001 Are Unicode NFKC normalization rules explicitly specified for multilingual text content? [Completeness, Spec §FR-007]
- [x] CHK002 Is HTML and Markdown tag stripping behavior defined to prevent accidental concatenation of neighboring words? [Clarity, Spec §FR-007]
- [x] CHK003 Are anchor text extraction rules for Markdown and HTML links documented unambiguously? [Clarity, Spec §FR-007]
- [x] CHK004 Is the tokenization behavior (Unicode alphanumeric tokens, punctuation exclusion, lowercase) completely specified? [Completeness, Spec §FR-007]
## Shingles & Consensus Metric Formulation
- [x] CHK005 Is the sliding window shingle size (5-tokens) and the fallback rule for short texts (< 5 tokens) explicitly defined? [Clarity, Spec §FR-008]
- [x] CHK006 Are the mathematical formulas for Coverage, Support, and F1 Score defined with explicit zero-division handling? [Measurability, Spec §FR-010]
- [x] CHK007 Is the threshold for a shingle to enter the Consensus set (presence in $\ge 2$ active candidates) unambiguously stated? [Clarity, Spec §FR-009]
## Decision & Tie-Breaking Hierarchy
- [x] CHK008 Is the technical tie threshold ($\le 0.03$) quantified with exact comparison semantics? [Clarity, Spec §FR-011]
- [x] CHK009 Is the tie-breaker preference for the smaller candidate (fewest shingles) explicitly constrained to candidates within the technical tie pool? [Consistency, Spec §FR-011]
- [x] CHK010 Is the zero-consensus fallback hierarchy (median of 3, maximum of 2, single candidate) completely specified without ambiguous gaps? [Coverage, Spec §FR-012]
- [x] CHK011 Is the final mandatory priority order (`newspaper4k` > `readability` > `trafilatura`) consistent across all tie scenarios? [Consistency, Spec §FR-011, §FR-012]
## Candidate State Transitions & Resilience
- [x] CHK012 Are the criteria distinguishing Usable, Degraded, and Unavailable candidates defined unambiguously? [Completeness, Spec §FR-005]
- [x] CHK013 Does the spec define the exact behavior and fallback when all 3 extractors are Unavailable? [Edge Case, Spec §FR-013]
- [x] CHK014 Does the spec define what occurs when degraded candidates exist but no usable candidates are present? [Coverage, Spec §FR-006]
## JSON Schema Integrity & Atomic I/O
- [x] CHK015 Are requirements explicit that 100% of pre-existing fields, structures, and article order must be preserved unchanged? [Completeness, Spec §FR-003, §FR-010]
- [x] CHK016 Is the output filename pattern `<original_name_without_extension>_selected.json` specified for default CLI execution? [Clarity, Spec §FR-016]
- [x] CHK017 Are atomic write requirements (temporary file + atomic replacement) defined to prevent partial or corrupted files on disk? [Non-Functional, Spec §FR-002, §FR-016]
- [x] CHK018 Is the behavior for recalculating an already present `selected_extractor` key explicitly specified? [Clarity, Spec §FR-014]
---
## Notes
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied.
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation.
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers.
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`.
- Items are numbered sequentially (CHK001 - CHK018) for easy reference.
@@ -0,0 +1,36 @@
# Specification Quality Checklist: Deterministic Content Selection
**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md)
## Content Quality
- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed
## Requirement Completeness
- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified
## Feature Readiness
- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification
## Notes
- All requirements are derived directly from PRD `docs/prd_deterministic_content_selection.md`.
- No ambiguity remains; all edge cases and tie-breaking hierarchies are fully specified.
- Ready for `/speckit-plan`.
@@ -0,0 +1,54 @@
# CLI Interface Contract: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](../spec.md)
---
## 1. Command Syntax
```bash
python scripts/select_article_extractor.py <input_file> [-o OUTPUT] [--indent INDENT] [--verbose]
```
---
## 2. Arguments and Flags
| Argumento / Flag | Tipo | Obrigatório | Padrão | Descrição |
|---|---|:---:|---|---|
| `input_file` | `Path` (Posicional) | Sim | - | Caminho para o arquivo JSON contendo a coleção `articles` extraída. |
| `-o`, `--output` | `Path` | Não | `<input_file_without_ext>_selected.json` | Caminho do arquivo JSON de destino. Se omitido, grava no mesmo diretório com sufixo `_selected.json`. |
| `--indent` | `int` | Não | `2` | Número de espaços para indentação do JSON de saída. Use `0` para JSON compacto em linha única. |
| `-v`, `--verbose` | `flag` | Não | `False` | Exibe no `stderr` detalhes da pontuação e justificativa de escolha por artigo. |
---
## 3. Standard Streams (I/O)
- **`stdout`**:
- Emite o sumário operacional em JSON ou texto resumido ao término da execução:
```json
{
"status": "success",
"input_file": "out/river_plate_extracted.json",
"output_file": "out/river_plate_extracted_selected.json",
"total_articles": 20,
"distribution": {
"newspaper4k": 9,
"readability": 9,
"trafilatura": 2
}
}
```
- **`stderr`**:
- Mensagens de log, progresso da barra/processamento de artigos e erros de validação ou exceções.
---
## 4. Exit Codes
| Código | Significado | Comportamento |
|:---:|---|---|
| `0` | **Sucesso** | Todos os artigos foram processados e o arquivo final foi gravado atomicamente com sucesso. |
| `1` | **Erro de I/O ou JSON Inválido** | Arquivo não encontrado, JSON malformado ou permissão negada. Nenhum arquivo de saída é gerado. |
| `2` | **Erro de Validação de Estrutura** | Raiz não é objeto ou chave `articles` não é uma lista. Nenhum arquivo de saída é gerado. |
@@ -0,0 +1,80 @@
# JSON Schema Contract: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](../spec.md)
---
## 1. Input JSON Schema
O arquivo de entrada deve conter uma lista de artigos sob a chave `articles`.
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["articles"],
"properties": {
"articles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"trafilatura": {
"type": "object",
"properties": {
"text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
},
"newspaper4k": {
"type": "object",
"properties": {
"text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
},
"readability": {
"type": "object",
"properties": {
"cleaned_text": { "type": ["string", "null"] },
"error": { "type": ["string", "null"] }
}
}
}
}
}
}
}
```
---
## 2. Output JSON Schema
O arquivo de saída mantém todos os campos, metadados e ordem originais, adicionando obrigatoriamente `selected_extractor`.
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"required": ["articles"],
"properties": {
"articles": {
"type": "array",
"items": {
"type": "object",
"required": ["selected_extractor"],
"properties": {
"selected_extractor": {
"type": "string",
"enum": ["trafilatura", "newspaper4k", "readability"]
},
"trafilatura": { "type": "object" },
"newspaper4k": { "type": "object" },
"readability": { "type": "object" }
}
}
}
}
}
```
@@ -0,0 +1,137 @@
# Data Model: Deterministic Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Domain Entities & Value Types
```mermaid
classDiagram
class ExtractorName {
<<enumeration>>
TRAFILATURA = "trafilatura"
NEWSPAPER4K = "newspaper4k"
READABILITY = "readability"
}
class CandidateStatus {
<<enumeration>>
USABLE
DEGRADED
UNAVAILABLE
}
class ExtractorCandidate {
+ExtractorName name
+str raw_text
+str error
+CandidateStatus status
+List~str~ tokens
+Set~Tuple~ shingles
+int shingle_count
+float coverage
+float support
+float score
}
class ArticleSelectionResult {
+int article_index
+ExtractorName selected_extractor
+str selection_reason
+int active_candidates_count
+int consensus_shingles_count
+Dict~ExtractorName, ExtractorCandidate~ candidates
}
class BatchProcessingResult {
+int total_articles
+int processed_count
+Dict~str, int~ selection_distribution
+str input_file
+str output_file
}
ExtractorCandidate --> ExtractorName
ExtractorCandidate --> CandidateStatus
ArticleSelectionResult --> ExtractorName
ArticleSelectionResult --> ExtractorCandidate
BatchProcessingResult --> ArticleSelectionResult
```
---
## 2. Entity Descriptions & Fields
### `ExtractorName` (Enum / Literal)
Enumeração estrita com os três motores de extração suportados:
- `"trafilatura"`
- `"newspaper4k"`
- `"readability"`
### `CandidateStatus` (Enum)
Classificação do estado de cada extrator em um dado artigo:
- `USABLE`: Campo de texto contém string não-vazia após normalização e campo `error` é nulo/vazio.
- `DEGRADED`: Campo de texto contém string não-vazia após normalização, porém campo `error` não é nulo.
- `UNAVAILABLE`: Campo de texto é ausente, nulo, tipo diferente de string ou vazio após normalização.
### `ExtractorCandidate` (Dataclass)
Representação estruturada de um candidato durante o cálculo:
| Campo | Tipo | Descrição |
|---|---|---|
| `name` | `ExtractorName` | Identificador do motor de extração (`trafilatura`, `newspaper4k`, `readability`). |
| `raw_text` | `str \| None` | Texto bruto obtido do campo correspondente no JSON (`trafilatura.text`, `newspaper4k.text`, `readability.cleaned_text`). |
| `error` | `str \| None` | Mensagem de erro do motor, se houver (`trafilatura.error`, etc.). |
| `status` | `CandidateStatus` | Estado de viabilidade do candidato (`USABLE`, `DEGRADED`, `UNAVAILABLE`). |
| `tokens` | `list[str]` | Sequência ordenada de tokens alfanuméricos minúsculos após normalização NFKC. |
| `shingles` | `set[tuple[str, ...]]` | Conjunto de n-grams consecutivos de 5 tokens (ou 1 n-gram se $1 \le \text{tokens} \le 4$). |
| `shingle_count` | `int` | Quantidade total de shingles gerados (`len(shingles)`). |
| `coverage` | `float` | Proporção de shingles do consenso presentes no candidato ($[0.0, 1.0]$). |
| `support` | `float` | Proporção de shingles do candidato que pertencem ao consenso ($[0.0, 1.0]$). |
| `score` | `float` | Pontuação $F_1$ baseada em cobertura e suporte ($[0.0, 1.0]$). |
---
### `ArticleSelectionResult` (Dataclass)
Resultado detalhado da avaliação para um único artigo:
| Campo | Tipo | Descrição |
|---|---|---|
| `article_index` | `int` | Posição ordinal do artigo no array `articles` original (0-indexed). |
| `selected_extractor` | `ExtractorName` | Vencedor da seleção determinística (`trafilatura`, `newspaper4k`, `readability`). |
| `selection_reason` | `str` | Justificativa rastreável da escolha (ex: `"highest_score"`, `"technical_tie_smallest_shingles"`, `"no_consensus_median_shingles"`, `"single_usable_candidate"`, `"fallback_all_unavailable"`). |
| `active_candidates_count` | `int` | Número de candidatos que formaram o conjunto ativo avaliado. |
| `consensus_shingles_count` | `int` | Quantidade de shingles no conjunto de consenso. |
| `candidates` | `dict[ExtractorName, ExtractorCandidate]` | Dicionário com o detalhamento de cada um dos 3 motores. |
---
### `BatchProcessingResult` (Dataclass)
Sumário da execução do lote:
| Campo | Tipo | Descrição |
|---|---|---|
| `total_articles` | `int` | Total de artigos encontrados no arquivo de entrada. |
| `processed_count` | `int` | Total de artigos processados e enriquecidos com sucesso. |
| `selection_distribution` | `dict[str, int]` | Contagem de seleções por motor (`{"trafilatura": X, "newspaper4k": Y, "readability": Z}`). |
| `input_file` | `str` | Caminho do arquivo lido. |
| `output_file` | `str` | Caminho do arquivo gerado de forma atômica. |
---
## 3. JSON Schema Mapping
### Entrada
- Raiz: Objeto contendo chave `articles: list[dict]`.
- Cada item em `articles`:
- `trafilatura` (objeto opcional): `{ "text": str | null, "error": str | null, ... }`
- `newspaper4k` (objeto opcional): `{ "text": str | null, "error": str | null, ... }`
- `readability` (objeto opcional): `{ "cleaned_text": str | null, "error": str | null, ... }`
### Saída
- Mesma estrutura exata da entrada, preservando 100% dos dados anteriores e ordem da lista `articles`.
- Em cada item de `articles`, adição/atualização da chave:
```json
"selected_extractor": "newspaper4k" | "readability" | "trafilatura"
```
@@ -0,0 +1,93 @@
# Implementation Plan: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
**Input**: Feature specification from `specs/004-deterministic-content-selection/spec.md`
---
## Summary
Implementação do motor determinístico de seleção de extratores (`scripts/select_article_extractor.py`), capaz de consumir arquivos JSON consolidados com saídas do **Trafilatura**, **Newspaper4k** e **Readability**, aplicar normalização de texto, geração de shingles (5-tokens), pontuação $F_1$ baseada em consenso e regras de desempate técnico / hierárquico estritas, gerando um novo arquivo JSON enriquecido exclusivamente com a chave `selected_extractor` em cada artigo de forma não-destrutiva e atômica.
---
## Technical Context
**Language/Version**: Python 3.10+
**Primary Dependencies**: Standard Library (`json`, `re`, `unicodedata`, `html`, `argparse`, `dataclasses`, `pathlib`, `tempfile`, `os`)
**Storage**: Arquivos JSON locais no diretório `out/`
**Testing**: `pytest` com testes unitários e de integração cobrindo 100% dos casos de teste obrigatórios (CT-001 a CT-014)
**Target Platform**: Windows / Linux / macOS (Terminal CLI & Módulo Python)
**Project Type**: CLI tool & modular selection engine
**Performance Goals**: Processamento em lote de centenas de artigos em menos de 1 segundo (complexidade linear $O(N)$ em memória)
**Constraints**: 100% determinístico, 0 chamadas de rede, sem uso de LLMs ou embeddings, escrita atômica em disco
**Scale/Scope**: Lotes de 1 a 10.000+ artigos
---
## Constitution Check
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
| Princípio | Avaliação | Status |
|---|---|---|
| **I. Library / Modular Design** | Módulo estruturado com funções puras e dataclasses desacopladas (`normalize_text`, `generate_shingles`, `calculate_consensus_metrics`, `select_best_candidate`, `process_batch`). | ✅ Aprovado |
| **II. CLI Interface** | CLI via `scripts/select_article_extractor.py` com flags descritivas, streams padronizados (`stdout` para resumo e `stderr` para logs/erros) e códigos de saída específicos. | ✅ Aprovado |
| **III. Test-First (NON-NEGOTIABLE)** | TDD com suíte automatizada em `tests/test_select_article_extractor.py` cobrindo todos os cenários (CT-001 a CT-014) e validação end-to-end com o arquivo real `out/river_plate_extracted.json`. | ✅ Aprovado |
| **IV. Integration Testing** | Testes de integração validando leitura, enriquecimento de `selected_extractor`, não-destrutividade de campos e escrita atômica. | ✅ Aprovado |
| **V. Simplicity & YAGNI** | Uso exclusivo da biblioteca padrão do Python, sem dependências adicionais pesadas. | ✅ Aprovado |
---
## Project Structure
### Documentation (this feature)
```text
specs/004-deterministic-content-selection/
├── spec.md # Especificação de requisitos funcionais e critérios
├── plan.md # Este plano de implementação (/speckit-plan)
├── research.md # Decisões técnicas e algoritmos (Phase 0)
├── data-model.md # Entidades e modelos de dados (Phase 1)
├── quickstart.md # Guia de validação e execução (Phase 1)
├── contracts/
│ ├── cli-contract.md # Contrato de linha de comando
│ └── json-schema.md # Esquemas JSON de entrada e saída
└── checklists/
└── requirements.md # Checklist de validação da especificação
```
### Source Code Layout
```text
scripts/
├── extract_google_news.py # Extrator RSS do Google News
├── extract_article_contents.py # Extrator multimotor de artigos
└── select_article_extractor.py # [NEW] Seletor determinístico de extrator por artigo
tests/
├── test_extract_google_news.py # Testes do extrator Google News
├── test_extract_article_contents.py # Testes do extrator multimotor
└── test_select_article_extractor.py # [NEW] Testes unitários e de integração do seletor
```
**Structure Decision**: Criação de `scripts/select_article_extractor.py` como ferramenta CLI e biblioteca modular autônoma, e `tests/test_select_article_extractor.py` contendo a suíte de testes de alta fidelidade aos requisitos do PRD.
---
## Implementation Phases
### Phase 0: Outline & Research *(Completed)*
- Normalização de texto via biblioteca padrão (`html.unescape`, `unicodedata.normalize('NFKC')`, regex Unicode).
- Estratégia de geração de shingles de 5 tokens e cálculo de $F_1$ sobre consenso compartilhado por $\ge 2$ motores.
- Regras de desempate técnico (`<= 0.03`), desempate sem consenso (mediana/máximo) e fallback prioritário (`newspaper4k` > `readability` > `trafilatura`).
- Documentado em [research.md](research.md).
### Phase 1: Design & Contracts *(Completed)*
- Modelos de dados e dataclasses estruturados em [data-model.md](data-model.md).
- Contratos de linha de comando e JSON schema definidos em [contracts/](contracts/).
- Guia prático de execução e validação estruturado em [quickstart.md](quickstart.md).
### Phase 2: Tasks & Execution *(Next Step via `/speckit-tasks`)*
- Criação das tarefas de implementação e testes orientados a TDD em `tasks.md`.
@@ -0,0 +1,66 @@
# Quickstart: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Pré-requisitos
- Python 3.10+
- Ambiente virtual configurado com dependências do projeto instaladas (`pip install -r requirements.txt`).
---
## 2. Execução Rápida via CLI
### Cenário 1: Selecionar o melhor extrator para uma extração existente
```bash
python scripts/select_article_extractor.py out/river_plate_extracted.json
```
**Resultado esperado**:
- Arquivo `out/river_plate_extracted_selected.json` gerado contendo todos os 20 artigos com a chave `selected_extractor` devidamente preenchida (`trafilatura`, `newspaper4k` ou `readability`).
- O arquivo original `out/river_plate_extracted.json` permanece inalterado.
### Cenário 2: Especificar caminho de saída customizado e modo verboso
```bash
python scripts/select_article_extractor.py out/river_plate_extracted.json -o out/meu_resultado.json --verbose
```
**Resultado esperado**:
- Logs detalhados no `stderr` mostrando as pontuações e a regra acionada (ex: `highest_score`, `technical_tie`, etc.).
---
## 3. Execução dos Testes Automatizados
Para rodar a suíte completa de testes unitários e de integração (cobrindo os casos CT-001 a CT-014):
```bash
pytest tests/test_select_article_extractor.py -v
```
---
## 4. Validação Programática / Uso como Módulo Python
```python
from scripts.select_article_extractor import select_article_extractor
article_data = {
"trafilatura": {"text": "El Club Atlético River Plate venció 2-0 anoche.", "error": None},
"newspaper4k": {
"text": "El Club Atlético River Plate venció 2-0 anoche en el Monumental.",
"error": None,
},
"readability": {
"cleaned_text": "El Club Atlético River Plate venció 2-0 anoche.",
"error": None,
},
}
result = select_article_extractor(article_data)
print("Extrator selecionado:", result.selected_extractor.value)
print("Motivo da escolha:", result.selection_reason)
# Output esperado:
# Extrator selecionado: newspaper4k (ou readability dependendo do desempate de shingles)
# Motivo da escolha: technical_tie_smallest_shingles (ou highest_score)
```
@@ -0,0 +1,107 @@
# Research & Architectural Decisions: Deterministic Content Selection
**Branch**: `004-deterministic-content-selection` | **Date**: 2026-08-20 | **Spec**: [spec.md](spec.md)
---
## 1. Text Normalization Pipeline
### Context
Cada extrator (Trafilatura, Newspaper4k e Readability) gera textos com diferentes resíduos de formatação (entidades HTML como `&amp;` ou `&nbsp;`, links no formato markdown `[texto](url)` ou tags `<a href="...">texto</a>`, variações de quebras de linha e pontuações). A comparação textual para consenso exige uma normalização uniforme, determinística e de alta performance.
### Decisions
1. **Decodificação de entidades HTML**: Utilizar `html.unescape()` da biblioteca padrão do Python.
2. **Remoção de imagens Markdown**: Expressão regular `re.compile(r'!\s*\[[^\]]*\]\([^)]*\)')` substituindo imagens Markdown por espaços para descartar marcação de mídia não-textual e evitar falsos consensos com legendas.
3. **Preservação de texto de links Markdown**: Expressão regular `re.compile(r'\[([^\]]+)\]\([^)]+\)')` substituindo links Markdown pelo texto âncora `\1`.
4. **Remoção de tags HTML**: Expressão regular `re.compile(r'<[^>]+>')` substituindo tags por espaços para evitar fusão acidental de palavras vizinhas.
5. **Normalização Unicode**: `unicodedata.normalize('NFKC', text)` para uniformizar caracteres compostos, ligaduras e variantes tipográficas.
6. **Conversão para minúsculas**: `.lower()` após NFKC.
7. **Colapso de espaços em branco**: `re.sub(r'\s+', ' ', text).strip()`.
8. **Tokenização**: Extração de sequências alfanuméricas com `re.findall(r'[\w]+', text, flags=re.UNICODE)`. Pontuações são descartadas naturalmente sem remoção semântica de palavras.
### Rationale
- 100% implementável com módulos padrão do Python (`re`, `unicodedata`, `html`), garantindo portabilidade em qualquer ambiente sem novas dependências externas.
- Complexidade linear $O(N)$ no tamanho do texto, com execução em frações de milissegundo por artigo.
### Alternatives Considered
- `BeautifulSoup` para strip de tags: Rejeitado por ser mais lento e desnecessário para textos já extraídos.
- `nltk` ou `spacy`: Rejeitados por adicionarem dependências pesadas, download de modelos e lentidão desnecessária para uma tarefa de tokenização alfanumérica pura.
---
## 2. 5-Token Shingles & Consensus Metrics
### Context
O algoritmo compara a sobreposição textual entre os candidatos ativos através de janelas deslizantes consecutivas de 5 tokens (shingles).
### Decisions
1. **Geração de Shingles**:
- Para um candidato com $T$ tokens ordenados $[t_0, t_1, \dots, t_{T-1}]$:
- Se $T \ge 5$: conjunto de tuplas de 5 tokens $\{ (t_i, t_{i+1}, t_{i+2}, t_{i+3}, t_{i+4}) \mid 0 \le i \le T-5 \}$.
- Se $1 \le T \le 4$: conjunto contendo uma única tupla com todos os tokens $\{ (t_0, \dots, t_{T-1}) \}$.
- Se $T = 0$: conjunto vazio $\emptyset$.
2. **Construção do Consenso**:
- Para cada shingle único observado nos candidatos ativos, conta-se em quantos candidatos distintos ele aparece.
- $\text{Consenso} = \{ s \mid \text{contagem}(s) \ge 2 \}$.
3. **Métricas por Candidato Ativo $C$**:
- $\text{cobertura}(C) = \frac{|C_{\text{shingles}} \cap \text{Consenso}|}{|\text{Consenso}|}$
- $\text{suporte}(C) = \frac{|C_{\text{shingles}} \cap \text{Consenso}|}{|C_{\text{shingles}}|}$
- $\text{score}(C) = \frac{2 \times \text{cobertura}(C) \times \text{suporte}(C)}{\text{cobertura}(C) + \text{suporte}(C)}$ (se denominador for zero, $\text{score} = 0.0$).
### Rationale
- A métrica de pontuação $F_1$ penaliza tanto extratores que perderam conteúdo essencial (baixa cobertura) quanto extratores que trouxeram excesso de lixo/boilerplate do site (baixo suporte).
- A representação por `set` de tuplas em Python permite operações de intersecção (`&`) com complexidade ótima de tempo $O(|C|)$.
---
## 3. Regras de Decisão, Empate Técnico e Desempate Hierárquico
### Context
O sistema precisa garantir uma escolha única e determinística em todas as variações possíveis de entrada.
### Decisions
1. **Formação do Conjunto Ativo**:
- Classificação:
- `Usável`: `text` é string não vazia após normalização e `error` é `None`/vazio.
- `Degradado`: `text` é string não vazia após normalização, mas `error` não é `None`.
- `Indisponível`: `text` é nulo, ausente, não-string ou vazio.
- Se houver $\ge 1$ Usável $\to$ Ativos = Usáveis.
- Senão, se houver $\ge 1$ Degradado $\to$ Ativos = Degradados.
- Senão $\to$ Seleciona `newspaper4k` diretamente (Fallback Final).
- Se $|\text{Ativos}| = 1 \to$ Seleciona o único candidato ativo imediatamente.
2. **Seleção Com Consenso ($|\text{Consenso}| > 0$)**:
- Maior score $S_{\max} = \max_{C \in \text{Ativos}} \text{score}(C)$.
- Grupo de empate técnico: $\{ C \in \text{Ativos} \mid S_{\max} - \text{score}(C) \le 0.03 + 10^{-9} \}$.
- Se grupo tiver 1 candidato $\to$ Seleciona ele.
- Se grupo tiver $\ge 2$ candidatos $\to$ Seleciona o candidato com menor $|C_{\text{shingles}}|$ (menor conteúdo excedente).
- Se ainda houver empate no número de shingles $\to$ Desempate por prioridade fixa: `newspaper4k` > `readability` > `trafilatura`.
3. **Seleção Sem Consenso ($|\text{Consenso}| = 0$)**:
- Se $|\text{Ativos}| = 3 \to$ Seleciona candidato com quantidade **mediana** de shingles.
- Se $|\text{Ativos}| = 2 \to$ Seleciona candidato com **maior** quantidade de shingles.
- Se $|\text{Ativos}| = 1 \to$ Seleciona o único candidato.
- Empates na quantidade de shingles $\to$ Prioridade fixa: `newspaper4k` > `readability` > `trafilatura`.
### Rationale
- Total aderência às seções 7.1 a 7.6 do PRD. A tolerância de $10^{-9}$ evita imprecisões de ponto flutuante em comparações `<= 0.03`.
---
## 4. Estratégia de I/O Não Destrutiva e Escrita Atômica
### Context
O processamento em lote deve preservar a ordem dos artigos e todos os campos originais do JSON, gravando o resultado sem risco de corrupção de arquivos em caso de interrupção.
### Decisions
1. **Entrada e Saída**:
- Nome padrão de saída: `<nome_original_sem_extensão>_selected.json`.
- Suporte a argumento opcional de saída `--output / -o`.
2. **Gravação Atômica**:
- Gravar os dados em um arquivo temporário no mesmo diretório (`<saida>.tmp.<pid>`).
- Executar substituição atômica via `os.replace(temp_path, target_path)`.
3. **Preservação de Conteúdo**:
- Carregar o JSON original em estruturas nativas de dicionário/lista.
- Inserir a chave `selected_extractor` diretamente em cada dicionário de artigo.
- Se `selected_extractor` já existir na entrada, sobrescrever com o novo valor recalculado.
### Rationale
- Garante integridade absoluta dos dados contra falhas de disco ou encerramentos abruptos.
@@ -0,0 +1,135 @@
# Feature Specification: Deterministic Content Selection
**Feature Branch**: `004-deterministic-content-selection`
**Created**: 2026-08-20
**Status**: Draft
**Input**: User description: "usando o PRD: docs/prd_deterministic_content_selection.md"
## User Scenarios & Testing *(mandatory)*
### User Story 1 - Deterministic Selection with Text Consensus (Priority: P1)
As a data pipeline consumer or analyst, I want the system to automatically analyze the extracted text from Trafilatura, Newspaper4k, and Readability for each article and pick the single best extractor using consensus and coverage scoring, so that our dataset has high-quality, standardized content without human review.
**Why this priority**: Core value of the feature. Resolves the primary dilemma of choosing between 3 extractor outputs per article based on mutual agreement (consensus shingles) and concise content.
**Independent Test**: Can be tested independently by running the selection algorithm on articles where extractors have high agreement or partial variations, verifying that the extractor with highest F1 score (or closest score with fewest excess shingles) is selected.
**Acceptance Scenarios**:
1. **Given** an article with usable extracts from all 3 libraries where 2 or 3 libraries agree closely, **When** selection is evaluated, **Then** the library with the highest consensus F1-score (or the more concise candidate within a 0.03 technical tie margin) is set in `selected_extractor`.
2. **Given** an extractor with excess boilerplate/noise and two extractors with clean common content, **When** selection is evaluated, **Then** the noisy extractor suffers lower support score and the clean agreeing extractor is selected.
3. **Given** an extractor with only a small snippet and two extractors with complete text, **When** selection is evaluated, **Then** the short snippet loses due to low consensus coverage.
---
### User Story 2 - Resilient Decision Under Total Disagreement or Degradation (Priority: P2)
As a pipeline maintainer, I want the selection algorithm to make a deterministic and sensible fallback choice even when extractors completely disagree, produce errors, or return empty/degraded content, so that the pipeline never halts or leaves an article without a chosen extractor.
**Why this priority**: Essential for pipeline stability. The system must guarantee that every article gets an unambiguous winner without throwing runtime exceptions or generating `null`/`ambiguous` states.
**Independent Test**: Can be tested with synthetic articles representing edge cases: all extractors returning non-overlapping text, extractors reporting errors, or all extractors failing.
**Acceptance Scenarios**:
1. **Given** 3 active candidates with 0 consensus shingles, **When** selection runs, **Then** the candidate with the median shingle length is selected.
2. **Given** 2 active candidates with 0 consensus shingles, **When** selection runs, **Then** the candidate with the larger shingle count is selected.
3. **Given** an article where all usable candidates are absent but degraded candidates exist, **When** selection runs, **Then** the algorithm evaluates only the degraded candidates.
4. **Given** an article where all 3 extractors failed or returned empty content, **When** selection runs, **Then** `newspaper4k` is selected via the mandatory final fallback rule.
---
### User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
As a system operator, I want to pass a JSON file with an `articles` array, execute the deterministic selector, and receive a new file `<original_name>_selected.json` with all original data and order intact plus the `selected_extractor` field, leaving the original file completely untouched.
**Why this priority**: Guarantees data preservation, idempotency, and clean pipeline integration.
**Independent Test**: Can be tested by running the process on a full batch JSON file (such as `out/river_plate_extracted.json`) and comparing input vs output keys, element counts, article order, and field contents.
**Acceptance Scenarios**:
1. **Given** a valid JSON file with $N$ articles, **When** the batch selection is executed, **Then** a new file `<original_name>_selected.json` is generated containing exactly $N$ articles in identical order, each with all original fields plus `selected_extractor`.
2. **Given** an input JSON file where `selected_extractor` already exists, **When** the batch selection is executed, **Then** `selected_extractor` is recalculated and updated.
3. **Given** an invalid JSON file or a file where `articles` is not a list, **When** execution runs, **Then** the process terminates with an error and does not produce a partial or corrupted output file.
---
### Edge Cases
- **Empty `articles` list (`[]`)**: Produces a valid output JSON containing an empty `articles: []` list without errors.
- **Exact score & shingle count tie**: Resolved deterministically by the strict fallback hierarchy: `newspaper4k` > `readability` > `trafilatura`.
- **Single active candidate**: When only 1 library produces usable output, it is selected immediately without computing consensus.
- **Short texts (< 5 tokens)**: When candidate text has between 1 and 4 tokens, the entire token sequence forms a single shingle.
- **Malformed fields / Type mismatch**: If a content field is not a string or missing, the candidate is classified as unavailable.
- **Missing library block**: If an article does not contain a `trafilatura`, `newspaper4k`, or `readability` block, that candidate is treated as unavailable.
## Requirements *(mandatory)*
### Functional Requirements
- **FR-001**: System MUST accept a valid JSON file path containing a root object with an `articles` array.
- **FR-002**: System MUST validate input structure (root is object, `articles` is list) and terminate immediately without creating an output file if validation fails.
- **FR-003**: System MUST process all articles in `articles`, preserving their exact sequence and all existing fields and values without modification.
- **FR-004**: System MUST evaluate extractor candidates using exclusively:
- `trafilatura.text` for Trafilatura
- `newspaper4k.text` for Newspaper4k
- `readability.cleaned_text` for Readability
- **FR-005**: System MUST categorize each candidate into one of three states:
- *Usable*: Content is non-empty string after normalization and extractor `error` is null/empty.
- *Degraded*: Content is non-empty string after normalization but extractor `error` is non-null.
- *Unavailable*: Content is missing, not a string, or empty after normalization.
- **FR-006**: System MUST form the active candidate set per article: Usable candidates if any exist; otherwise Degraded candidates if any exist; otherwise trigger final fallback.
- **FR-007**: System MUST perform deterministic in-memory normalization for candidate comparisons:
1. Decode HTML entities.
2. Strip Markdown images (`![alt](url)`), removing non-textual media embeds.
3. In Markdown links (`[text](url)`), preserve anchor text and strip URL targets.
4. Strip HTML tags, maintaining spacing between adjacent words.
5. Apply Unicode NFKC normalization.
6. Convert to lowercase.
7. Collapse multiple whitespace/newlines/tabs into a single space.
8. Tokenize retaining Unicode letters and digits.
9. Ignore punctuation symbols.
- **FR-008**: System MUST generate 5-token sliding window shingles from the ordered token sequence of each candidate (or single $N$-token shingle if $1 \le N \le 4$).
- **FR-009**: System MUST construct the consensus shingle set (shingles appearing in at least 2 active candidates).
- **FR-010**: System MUST compute `coverage`, `support`, and `score` ($F_1 = 2 \times \text{coverage} \times \text{support} / (\text{coverage} + \text{support})$) for each active candidate against the consensus shingles (or 0 if denominator is 0).
- **FR-011**: When consensus shingles exist, the system MUST:
1. Sort candidates descending by score.
2. Identify all candidates within a `0.03` difference from the top score (technical tie pool).
3. If technical tie pool has 1 candidate, select it.
4. If multiple candidates are in technical tie, select the one with the smallest total shingle count (least surplus).
5. If shingle count is also tied, apply priority hierarchy: `newspaper4k` > `readability` > `trafilatura`.
- **FR-012**: When 0 consensus shingles exist, the system MUST:
- With 3 active candidates: select candidate with median shingle count.
- With 2 active candidates: select candidate with maximum shingle count.
- With 1 active candidate: select that single candidate.
- In shingle count ties: apply priority hierarchy (`newspaper4k` > `readability` > `trafilatura`).
- **FR-013**: When 0 active candidates exist (all unavailable), system MUST assign `newspaper4k`.
- **FR-014**: System MUST inject or replace `selected_extractor` in each article item with exactly one value from `{"trafilatura", "newspaper4k", "readability"}`.
- **FR-015**: System MUST never output `null`, empty string, `ambiguous`, or leave an article without a selection.
- **FR-016**: System MUST write the result atomically to `<original_name_without_extension>_selected.json` in the same directory or specified target, leaving the input file unchanged.
- **FR-017**: System MUST produce 100% deterministic and identical outputs across repeated runs with identical inputs.
### Key Entities *(include if feature involves data)*
- **Article Input Batch**: Root JSON container with metadata and an ordered list of `articles`.
- **Article Record**: Object representing an article, containing source metadata, extraction results from the 3 extractors (`trafilatura`, `newspaper4k`, `readability`), and the resulting `selected_extractor` tag.
- **Extractor Candidate**: Evaluation model for an individual extractor containing raw text, error state, candidate usability state (`Usable`, `Degraded`, `Unavailable`), normalized token stream, 5-token shingles, and computed metrics (`coverage`, `support`, `score`, `shingle_count`).
- **Consensus Shingle Set**: Set of unique 5-token shingles shared by 2 or more active extractor candidates.
## Success Criteria *(mandatory)*
### Measurable Outcomes
- **SC-001**: **100% Selection Completeness**: 100% of articles in the input collection receive a valid `selected_extractor` value from the closed set `['trafilatura', 'newspaper4k', 'readability']`.
- **SC-002**: **0% Ambiguity**: Exactly 0 articles result in `null`, missing, empty, or ambiguous selection states.
- **SC-003**: **100% Deterministic Reproducibility**: 100% identical `selected_extractor` values when executing across multiple runs on identical input datasets.
- **SC-004**: **100% Non-Destructive Integrity**: 100% of pre-existing keys, nested objects, article counts, and article ordering are preserved identically in the output JSON.
- **SC-005**: **100% Test Case Coverage**: Passes 100% of defined mandatory test cases (CT-001 through CT-014).
- **SC-006**: **Atomic Operation**: 0 partial or corrupted output files generated on process failure or invalid JSON inputs.
## Assumptions
- The input JSON is generated by the extraction pipeline and contains `articles` where each item may have `trafilatura`, `newspaper4k`, and `readability` sub-objects.
- All three extraction libraries operated on the exact same HTML source document.
- No external dependencies (LLM APIs, embedding services, or network calls) are permitted during the selection process.
- Unicode NFKC normalization and standard tokenization cover multilingual article content (e.g. Portuguese, Spanish, English).
- Default output file path naming convention `<name>_selected.json` is sufficient, with CLI support for optional custom destination.
@@ -0,0 +1,148 @@
# Tasks: Deterministic Article Content Selection
**Branch**: `004-deterministic-content-selection` | **Spec**: [spec.md](spec.md) | **Plan**: [plan.md](plan.md)
---
## Phase 1: Setup (Shared Infrastructure)
**Purpose**: Project initialization and test harness setup
- [X] T001 Initialize script entrypoint and test suite structure in `scripts/select_article_extractor.py` and `tests/test_select_article_extractor.py`
---
## Phase 2: Foundational (Data Structures & Normalization Engine)
**Purpose**: Core data models, text normalization, and shingle generation that all user stories depend upon
**⚠️ CRITICAL**: Must be completed before user story implementation begins
- [X] T002 [P] Implement dataclasses and enumerations (`ExtractorName`, `CandidateStatus`, `ExtractorCandidate`, `ArticleSelectionResult`, `BatchProcessingResult`) in `scripts/select_article_extractor.py`
- [X] T003 [P] Implement text normalization pipeline (`normalize_text`, HTML entities unescape, HTML/Markdown tag strip, NFKC, lowercase, Unicode tokenization) in `scripts/select_article_extractor.py`
- [X] T004 Implement 5-token sliding window and short text shingle generator (`generate_shingles`) in `scripts/select_article_extractor.py`
- [X] T005 Implement unit tests for normalization, tokenization, and shingle generation in `tests/test_select_article_extractor.py`
**Checkpoint**: Foundation ready — text normalization and shingle generator fully operational and tested.
---
## Phase 3: User Story 1 - Deterministic Selection with Text Consensus (Priority: P1) 🎯 MVP
**Goal**: Calculate consensus shingles ($\ge 2$ active extractors), compute Coverage, Support, and $F_1$ score, apply technical tie margin ($\le 0.03$), and select the best extractor based on agreement and conciseness.
**Independent Test**: Execute tests with synthetic and real articles where 2 or 3 extractors agree (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) and verify the winner matches expected score / tie-breaker.
### Tests for User Story 1 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T006 [P] [US1] Write unit tests for consensus scoring, coverage/support F1 calculation, and technical tie-breaking (CT-001, CT-002, CT-003, CT-009, CT-010, CT-011) in `tests/test_select_article_extractor.py`
### Implementation for User Story 1
- [X] T007 [US1] Implement consensus shingle builder and metric calculator (`calculate_consensus_metrics`) in `scripts/select_article_extractor.py`
- [X] T008 [US1] Implement consensus decision selector with technical tie pool ($\le 0.03$), smallest shingle count preference, and final priority fallback (`select_with_consensus`) in `scripts/select_article_extractor.py`
**Checkpoint**: User Story 1 (MVP) is fully functional and independently testable for all consensus scenarios.
---
## Phase 4: User Story 2 - Resilient Decision Under Disagreement or Degradation (Priority: P2)
**Goal**: Ensure zero unassigned or ambiguous selections by handling zero-consensus articles (median of 3, max of 2, single), candidate degradation, all-unavailable extractors, and strict tie-breaking priority (`newspaper4k` > `readability` > `trafilatura`).
**Independent Test**: Execute tests for zero consensus, degraded errors, single usable candidate, and total extraction failure (CT-004, CT-005, CT-006, CT-007, CT-008).
### Tests for User Story 2 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T009 [P] [US2] Write unit tests for zero-consensus, degraded candidate handling, single candidate, and total unavailability fallback (CT-004, CT-005, CT-006, CT-007, CT-008) in `tests/test_select_article_extractor.py`
### Implementation for User Story 2
- [X] T010 [US2] Implement candidate status classifier (`classify_candidate_status`) and active candidate set builder (`form_active_set`) in `scripts/select_article_extractor.py`
- [X] T011 [US2] Implement zero-consensus decision logic (median of 3, max of 2, single candidate, priority hierarchy) (`select_without_consensus`) in `scripts/select_article_extractor.py`
- [X] T012 [US2] Implement single article selector orchestrator (`select_article_extractor`) integrating Usable, Degraded, Consensus, and Non-Consensus decision branches in `scripts/select_article_extractor.py`
**Checkpoint**: User Stories 1 AND 2 are fully functional and handle 100% of single-article decision paths.
---
## Phase 5: User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)
**Goal**: Read JSON batch files containing `articles`, preserve all original fields and article order, recalculate existing `selected_extractor` values, validate schema, and write output atomically to `<name>_selected.json`.
**Independent Test**: Execute CLI and batch tests (CT-012, CT-013, CT-014), verify non-destructive field preservation, and run end-to-end processing on `out/river_plate_extracted.json`.
### Tests for User Story 3 🧪
> **NOTE: Write these tests FIRST, ensure they FAIL before implementation**
- [X] T013 [P] [US3] Write integration tests for JSON schema validation, error handling, key preservation, recalculation of existing key, and atomic I/O (CT-012, CT-013, CT-014) in `tests/test_select_article_extractor.py`
### Implementation for User Story 3
- [X] T014 [US3] Implement batch processor (`process_batch`) and atomic file saver (`atomic_save_json`) in `scripts/select_article_extractor.py`
- [X] T015 [US3] Implement CLI interface (`main`) with `argparse`, options (`-o`, `--indent`, `--verbose`), exit codes (0, 1, 2), `stderr` logs, and `stdout` JSON summary in `scripts/select_article_extractor.py`
**Checkpoint**: Full end-to-end batch processing operational and tested against contracts and real datasets.
---
## Phase 6: Polish & Cross-Cutting Concerns
**Purpose**: Validation, performance checks, and documentation verification
- [X] T016 [P] Execute quickstart validation scenarios on `out/river_plate_extracted.json` per `specs/004-deterministic-content-selection/quickstart.md`
- [X] T017 Run full test suite with coverage via `pytest tests/test_select_article_extractor.py -v` ensuring all 14 mandatory test cases (CT-001 to CT-014) pass
---
## Dependencies & Execution Order
```mermaid
graph TD
T001[T001: Setup Harness] --> T002[T002: Data Models]
T001 --> T003[T003: Text Normalization]
T002 --> T004[T004: Shingle Generator]
T003 --> T004
T004 --> T005[T005: Foundation Tests]
T005 --> T006[T006: US1 Tests]
T006 --> T007[T007: US1 Consensus Metrics]
T007 --> T008[T008: US1 Consensus Selection]
T008 --> T009[T009: US2 Tests]
T009 --> T010[T010: US2 Candidate Classifier]
T010 --> T011[T011: US2 Zero-Consensus Logic]
T011 --> T012[T012: US2 Orchestrator]
T012 --> T013[T013: US3 Integration Tests]
T013 --> T014[T014: US3 Batch & Atomic I/O]
T014 --> T015[T015: US3 CLI Interface]
T015 --> T016[T016: Quickstart Validation]
T016 --> T017[T017: Full Test Suite CT-001..CT-014]
```
---
## Parallel Opportunities
- **Phase 2 (Foundations)**: `T002` (Models) and `T003` (Normalization) can be developed in parallel.
- **Phase 3 (User Story 1)**: `T006` (Unit tests) can be written in parallel with data model test setups.
- **Phase 4 (User Story 2)**: `T009` (Unit tests) can be written in parallel with candidate state transition logic.
- **Phase 5 (User Story 3)**: `T013` (Integration tests) can be developed alongside CLI option parser definitions.
---
## Implementation Strategy
### MVP First (Phases 1, 2 & 3)
1. Complete Setup and Foundations (`T001` - `T005`).
2. Implement User Story 1 (`T006` - `T008`): text consensus and scoring engine.
3. Validate MVP: test consensus-based decisions on multi-extractor articles.
### Incremental Delivery (Phases 4, 5 & 6)
4. Add User Story 2 (`T009` - `T012`): resilient fallbacks, zero consensus, degraded candidate evaluation.
5. Add User Story 3 (`T013` - `T015`): non-destructive batch JSON processing, atomic save, and CLI tool.
6. Polish & Final Verification (`T016` - `T017`): run quickstart validation and full test suite covering CT-001 to CT-014.
+615
View File
@@ -0,0 +1,615 @@
"""
Suíte de Testes Automatizados para o Seletor Determinístico de Extrator.
Cobre 100% dos Casos de Teste Obrigatórios do PRD (CT-001 a CT-014), testes unitários
de normalização e shingles, testes de integração de lote e testes E2E via subprocess.
"""
from __future__ import annotations
import json
import subprocess
import sys
from pathlib import Path
import pytest
from scripts.select_article_extractor import (
ExtractorName,
generate_shingles,
normalize_text,
process_batch,
select_article_extractor,
)
# ==============================================================================
# Testes Unitários de Normalização e Tokenização
# ==============================================================================
def test_normalize_text_empty_and_invalid():
assert normalize_text(None) == []
assert normalize_text("") == []
assert normalize_text(" \n\t ") == []
assert normalize_text(12345) == []
def test_normalize_text_html_entities_and_tags():
raw = "<p>El &amp; <b>futebol</b> mundial &quot;está&quot; mudando.</p>"
tokens = normalize_text(raw)
assert tokens == ["el", "futebol", "mundial", "está", "mudando"]
def test_normalize_text_markdown_links():
raw = "Veja mais no [Portal de Notícias](https://example.com/noticias) hoje."
tokens = normalize_text(raw)
assert tokens == ["veja", "mais", "no", "portal", "de", "notícias", "hoje"]
def test_normalize_text_markdown_images_stripped_while_links_preserved():
"""Garante que marcação de imagem Markdown ![alt](url) seja descartada e link [texto](url) seja preservado."""
raw = (
"Texto inicial do artigo. "
"![Legenda da foto e imagem](https://example.com/imagem.webp) "
"Mais texto com [link importante](https://example.com/pagina) e outra "
"![Outra foto](https://example.com/foto2.jpg) informação."
)
tokens = normalize_text(raw)
assert "legenda" not in tokens
assert "foto" not in tokens
assert "imagem" not in tokens
assert "link" in tokens
assert "importante" in tokens
assert tokens == [
"texto",
"inicial",
"do",
"artigo",
"mais",
"texto",
"com",
"link",
"importante",
"e",
"outra",
"informação",
]
def test_normalize_text_nfkc_unicode_and_punctuation():
# Caracteres combinados e pontuação
raw = "River Plate venceu por 3-0! (Com gol de pênalti & golaço de falta)."
tokens = normalize_text(raw)
assert tokens == [
"river",
"plate",
"venceu",
"por",
"3",
"0",
"com",
"gol",
"de",
"pênalti",
"golaço",
"de",
"falta",
]
# ==============================================================================
# Testes Unitários de Shingles
# ==============================================================================
def test_generate_shingles_sliding_window():
tokens = ["um", "dois", "três", "quatro", "cinco", "seis"]
shingles = generate_shingles(tokens, window_size=5)
assert len(shingles) == 2
assert ("um", "dois", "três", "quatro", "cinco") in shingles
assert ("dois", "três", "quatro", "cinco", "seis") in shingles
def test_generate_shingles_short_text():
# Entre 1 e 4 tokens deve gerar 1 único shingle com a tupla completa
tokens = ["river", "plate", "campeão"]
shingles = generate_shingles(tokens, window_size=5)
assert len(shingles) == 1
assert ("river", "plate", "campeão") in shingles
def test_generate_shingles_empty():
assert generate_shingles([]) == set()
# ==============================================================================
# Casos de Teste Obrigatórios do PRD (§12: CT-001 a CT-014)
# ==============================================================================
def test_ct_001_three_candidates_clear_winner():
"""CT-001: Três candidatos com consenso e um vencedor claro -> Selecionar o maior score."""
base = "river plate venceu o clássico ontem a noite no estádio monumental"
article = {
"trafilatura": {"text": base, "error": None},
"newspaper4k": {"text": base, "error": None},
"readability": {
"cleaned_text": "texto completamente diferente sem nenhuma relação",
"error": None,
},
}
result = select_article_extractor(article)
assert result.selected_extractor in (ExtractorName.NEWSPAPER4K, ExtractorName.TRAFILATURA)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
def test_ct_002_technical_tie_smallest_shingles():
"""CT-002: Dois ou mais candidatos dentro de 0,03 do maior score -> Selecionar o de menor quantidade de shingles."""
tokens_comuns = (
"o rio de janeiro continua lindo e sempre maravilhoso em todas as estações do ano"
)
article = {
"trafilatura": {"text": tokens_comuns, "error": None},
"readability": {
"cleaned_text": tokens_comuns + " propaganda extra adicionada no fim",
"error": None,
},
"newspaper4k": {"text": tokens_comuns, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
def test_ct_003_technical_tie_priority_fallback():
"""CT-003: Empate técnico e mesma quantidade de shingles -> Aplicar prioridade final (newspaper4k > readability > trafilatura)."""
texto = "o time jogou muito bem durante toda a partida de futebol"
article = {
"trafilatura": {"text": texto, "error": None},
"readability": {"cleaned_text": texto, "error": None},
"newspaper4k": {"text": texto, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
# Agora sem newspaper4k ativo (somente readability e trafilatura idênticos)
article_two = {
"trafilatura": {"text": texto, "error": None},
"readability": {"cleaned_text": texto, "error": None},
"newspaper4k": {"text": None, "error": "Crash"},
}
result_two = select_article_extractor(article_two)
assert result_two.selected_extractor == ExtractorName.READABILITY
def test_ct_004_three_candidates_no_consensus():
"""CT-004: Três candidatos sem consenso -> Selecionar a quantidade mediana de shingles."""
t1 = "alfa bravo charlie delta echo foxtrot golf hotel india juliet" # 10 tokens -> 6 shingles
t2 = "kilo lima mike november oscar papa quebec romeo sierra tango uniform victor" # 12 tokens -> 8 shingles (MEDIANA)
t3 = "whiskey xray yankee zulu zero one two three four five six seven eight nine" # 14 tokens -> 10 shingles
article = {
"trafilatura": {"text": t1, "error": None},
"readability": {"cleaned_text": t2, "error": None},
"newspaper4k": {"text": t3, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.READABILITY
assert result.selection_reason == "no_consensus_median_shingles"
def test_ct_005_two_candidates_no_consensus():
"""CT-005: Dois candidatos sem consenso -> Selecionar a maior quantidade de shingles."""
t_short = "alfa bravo charlie delta echo foxtrot" # 6 tokens -> 2 shingles
t_long = (
"kilo lima mike november oscar papa quebec romeo sierra tango" # 10 tokens -> 6 shingles
)
article = {
"trafilatura": {"text": t_short, "error": None},
"newspaper4k": {"text": t_long, "error": None},
"readability": {"cleaned_text": None, "error": "Not extracted"},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
assert result.selection_reason == "no_consensus_max_shingles"
def test_ct_006_single_usable_candidate():
"""CT-006: Somente um candidato utilizável -> Selecionar esse candidato."""
article = {
"trafilatura": {
"text": "conteúdo válido e utilizável extraído com sucesso aqui",
"error": None,
},
"newspaper4k": {"text": "", "error": None},
"readability": {"cleaned_text": None, "error": "Timeout error"},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.TRAFILATURA
assert result.selection_reason == "single_usable_candidate"
def test_ct_007_degraded_candidates_only():
"""CT-007: Nenhum utilizável, mas existe candidato degradado -> Executar o algoritmo somente com os degradados."""
texto_comum = "artigo relevante sobre economia global e finanças internacionais com detalhes"
article = {
"trafilatura": {"text": texto_comum, "error": "Warning: partial parse"},
"newspaper4k": {"text": texto_comum, "error": "HTTP 403 partial"},
"readability": {"cleaned_text": None, "error": "Fatal exception"},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
def test_ct_008_all_candidates_unavailable():
"""CT-008: Todos os candidatos indisponíveis -> Selecionar newspaper4k."""
article = {
"trafilatura": {"text": None, "error": "Error"},
"newspaper4k": {"text": "", "error": "Empty"},
"readability": {"cleaned_text": " ", "error": "Blank"},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
assert result.selection_reason == "fallback_all_unavailable"
def test_ct_009_small_fragment_loses_due_to_low_coverage():
"""CT-009: Readability retorna apenas um fragmento pequeno enquanto os outros concordam -> Perde por baixa cobertura."""
full_text = (
"o presidente da república anunciou novas medidas econômicas para conter a inflação "
"e estimular o crescimento industrial em todo o território nacional durante o pronunciamento oficial"
)
small_fragment = "o presidente da república anunciou"
article = {
"trafilatura": {"text": full_text, "error": None},
"newspaper4k": {"text": full_text, "error": None},
"readability": {"cleaned_text": small_fragment, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor in (ExtractorName.TRAFILATURA, ExtractorName.NEWSPAPER4K)
assert result.selected_extractor != ExtractorName.READABILITY
def test_ct_010_excessive_boilerplate_loses_due_to_low_support():
"""CT-010: Um candidato contém o conteúdo comum e muito conteúdo excedente -> Perde suporte e reduz pontuação."""
common_content = (
"notícia oficial com dados apurados sobre a operação policial realizada nesta manhã"
)
massive_boilerplate = common_content + (
" compartilhe no facebook twitter whatsapp veja também esportes receitas horóscopo política "
" e assine nossa newsletter diária para receber mais novidades sobre culinária e fofocas"
)
article = {
"trafilatura": {"text": common_content, "error": None},
"readability": {"cleaned_text": common_content, "error": None},
"newspaper4k": {"text": massive_boilerplate, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.READABILITY
def test_ct_011_partial_content_loses_due_to_low_coverage():
"""CT-011: Um candidato contém somente parte do conteúdo comum -> Perde cobertura e reduz pontuação."""
full_content = "primeiro parágrafo do artigo completo segundo parágrafo com explicações terceiro parágrafo final"
half_content = "primeiro parágrafo do artigo completo"
article = {
"newspaper4k": {"text": full_content, "error": None},
"readability": {"cleaned_text": full_content, "error": None},
"trafilatura": {"text": half_content, "error": None},
}
result = select_article_extractor(article)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
def test_ct_012_recalculate_existing_selected_extractor(tmp_path: Path):
"""CT-012: A entrada já contém selected_extractor -> Recalcular e substituir somente essa chave."""
input_data = {
"articles": [
{
"titulo": "Teste",
"selected_extractor": "trafilatura",
"trafilatura": {"text": "lixo sem sentido", "error": None},
"newspaper4k": {
"text": "conteúdo correto compartilhado por dois motores",
"error": None,
},
"readability": {
"cleaned_text": "conteúdo correto compartilhado por dois motores",
"error": None,
},
}
]
}
in_file = tmp_path / "artigos.json"
in_file.write_text(json.dumps(input_data, ensure_ascii=False), encoding="utf-8")
res = process_batch(in_file)
out_file = Path(res.output_file)
assert out_file.exists()
with open(out_file, "r", encoding="utf-8") as f:
out_data = json.load(f)
assert out_data["articles"][0]["selected_extractor"] == "newspaper4k"
def test_ct_013_empty_articles_list(tmp_path: Path):
"""CT-013: articles está vazio -> Gerar saída válida com articles vazio."""
input_data = {"metadata": "info", "articles": []}
in_file = tmp_path / "empty.json"
in_file.write_text(json.dumps(input_data), encoding="utf-8")
res = process_batch(in_file)
out_file = Path(res.output_file)
assert out_file.exists()
with open(out_file, "r", encoding="utf-8") as f:
out_data = json.load(f)
assert out_data["metadata"] == "info"
assert out_data["articles"] == []
assert res.total_articles == 0
assert res.processed_count == 0
def test_ct_014_invalid_json_fails_atomically(tmp_path: Path):
"""CT-014: JSON inválido -> Não gerar saída."""
in_file = tmp_path / "invalid.json"
in_file.write_text("{articles: [ malformed json", encoding="utf-8")
expected_out = tmp_path / "invalid_selected.json"
with pytest.raises(ValueError, match="JSON inválido"):
process_batch(in_file)
assert not expected_out.exists()
# ==============================================================================
# Testes de Integração
# ==============================================================================
def test_integration_reference_file(tmp_path: Path):
"""Testa o processamento em lote completo sobre o arquivo real out/river_plate_extracted.json."""
ref_file = Path("out/river_plate_extracted.json")
if not ref_file.exists():
pytest.skip(
"Arquivo out/river_plate_extracted.json não encontrado para teste de integração."
)
out_file = tmp_path / "river_plate_extracted_selected.json"
result = process_batch(ref_file, output_path=out_file, verbose=True)
assert result.total_articles == 20
assert result.processed_count == 20
assert out_file.exists()
with open(out_file, "r", encoding="utf-8") as f:
data = json.load(f)
assert len(data["articles"]) == 20
for art in data["articles"]:
assert "selected_extractor" in art
assert art["selected_extractor"] in ["trafilatura", "newspaper4k", "readability"]
# Validar a distribuição exata conforme o algoritmo do PRD
assert result.selection_distribution == {
"newspaper4k": 9,
"readability": 9,
"trafilatura": 2,
}
def test_article_1_regression_technical_tie_markdown_images():
"""
Teste de regressão para o Artigo 1 (Los puntajes de River vs. Independiente Santa Fe):
Valida que com o descarte de imagens Markdown ![alt](url), Readability e Newspaper4k
entram em empate técnico (diff <= 0.03) e Newspaper4k vence por possuir menor quantidade
de shingles (1009 vs 1057).
"""
ref_file = Path("out/river_plate_extracted.json")
if not ref_file.exists():
pytest.skip("Arquivo out/river_plate_extracted.json não encontrado.")
with open(ref_file, "r", encoding="utf-8") as f:
data = json.load(f)
art1 = data["articles"][0]
result = select_article_extractor(art1, article_index=0)
assert result.selected_extractor == ExtractorName.NEWSPAPER4K
assert result.selection_reason == "technical_tie_smallest_shingles"
cand_news = result.candidates[ExtractorName.NEWSPAPER4K]
cand_read = result.candidates[ExtractorName.READABILITY]
cand_traf = result.candidates[ExtractorName.TRAFILATURA]
assert cand_news.shingle_count == 1009
assert cand_read.shingle_count == 1057
assert cand_traf.shingle_count == 1091
# Diferença para o maior score <= 0.03 (empate técnico)
max_score = max(cand_news.score, cand_read.score, cand_traf.score)
assert (max_score - cand_news.score) <= 0.03
assert (max_score - cand_read.score) <= 0.03
def test_integration_large_batch_determinism(tmp_path: Path):
"""Gera um lote de 100 artigos sintéticos e verifica 100% de repetibilidade determinística entre 2 execuções."""
articles = []
for i in range(100):
if i % 4 == 0:
art = {
"trafilatura": {
"text": f"artigo numero {i} sobre futebol internacional no estadio",
"error": None,
},
"newspaper4k": {
"text": f"artigo numero {i} sobre futebol internacional no estadio",
"error": None,
},
"readability": {"cleaned_text": "sem relacao", "error": None},
}
elif i % 4 == 1:
art = {
"trafilatura": {"text": None, "error": "timeout"},
"newspaper4k": {"text": f"noticia exclusiva {i} com detalhes", "error": None},
"readability": {
"cleaned_text": f"noticia exclusiva {i} com detalhes",
"error": None,
},
}
elif i % 4 == 2:
art = {
"trafilatura": {"text": f"texto a {i}", "error": None},
"newspaper4k": {"text": f"texto b diferente {i}", "error": None},
"readability": {"cleaned_text": f"texto c terceiro {i}", "error": None},
}
else:
art = {
"trafilatura": {"text": None, "error": "err"},
"newspaper4k": {"text": "", "error": "err"},
"readability": {"cleaned_text": None, "error": "err"},
}
articles.append(art)
batch_payload = {"articles": articles}
in_file = tmp_path / "large_batch.json"
in_file.write_text(json.dumps(batch_payload, ensure_ascii=False), encoding="utf-8")
out1 = tmp_path / "large_batch_run1.json"
out2 = tmp_path / "large_batch_run2.json"
res1 = process_batch(in_file, output_path=out1)
res2 = process_batch(in_file, output_path=out2)
assert res1.processed_count == 100
assert res2.processed_count == 100
# Os resultados devem ser 100% idênticos
selections1 = [s.selected_extractor for s in res1.selections]
selections2 = [s.selected_extractor for s in res2.selections]
assert selections1 == selections2
def test_integration_pipeline_downstream_consumer(tmp_path: Path):
"""
Testa a integração end-to-end do pipeline downstream:
Lê o JSON enriquecido com selected_extractor, recupera o conteúdo do extrator vencedor e
garante que o texto está higienizado e pronto para os classificadores NLP.
"""
ref_file = Path("out/river_plate_extracted.json")
if not ref_file.exists():
pytest.skip("Arquivo out/river_plate_extracted.json não encontrado.")
out_file = tmp_path / "downstream_test.json"
process_batch(ref_file, output_path=out_file)
with open(out_file, "r", encoding="utf-8") as f:
data = json.load(f)
for idx, art in enumerate(data["articles"]):
winner = art["selected_extractor"]
assert winner in ["trafilatura", "newspaper4k", "readability"]
# Recuperar texto do extrator vencedor conforme mapeamento do PRD
if winner == "trafilatura":
chosen_text = art["trafilatura"]["text"]
elif winner == "newspaper4k":
chosen_text = art["newspaper4k"]["text"]
elif winner == "readability":
chosen_text = art["readability"]["cleaned_text"]
assert isinstance(chosen_text, str)
assert len(chosen_text.strip()) > 0
# ==============================================================================
# Testes End-to-End (E2E) via Subprocess CLI
# ==============================================================================
def test_e2e_cli_subprocess_real_execution(tmp_path: Path):
"""E2E: Executa scripts/select_article_extractor.py como subprocesso real na linha de comando."""
ref_file = Path("out/river_plate_extracted.json")
if not ref_file.exists():
pytest.skip("Arquivo out/river_plate_extracted.json não encontrado para teste E2E.")
out_file = tmp_path / "e2e_river_plate_selected.json"
script_path = Path("scripts/select_article_extractor.py").resolve()
cmd = [
sys.executable,
str(script_path),
str(ref_file.resolve()),
"-o",
str(out_file),
"--verbose",
"--indent",
"2",
]
proc = subprocess.run(cmd, capture_output=True, text=True, check=False)
assert proc.returncode == 0
assert out_file.exists()
# Validar saída estruturada JSON do stdout
summary = json.loads(proc.stdout)
assert summary["status"] == "success"
assert summary["total_articles"] == 20
assert summary["processed_count"] == 20
assert "distribution" in summary
# Validar que logs de verbose foram emitidos no stderr
assert "[Artigo #001]" in proc.stderr
assert "[Artigo #020]" in proc.stderr
def test_e2e_cli_subprocess_default_naming(tmp_path: Path):
"""E2E: Executa CLI sem a flag -o e valida criação automática de <nome>_selected.json."""
sample_data = {
"articles": [
{
"titulo": "Artigo Automático",
"trafilatura": {"text": "conteúdo padrão", "error": None},
"newspaper4k": {"text": "conteúdo padrão", "error": None},
"readability": {"cleaned_text": "conteúdo padrão", "error": None},
}
]
}
in_file = tmp_path / "my_news.json"
in_file.write_text(json.dumps(sample_data), encoding="utf-8")
expected_out = tmp_path / "my_news_selected.json"
script_path = Path("scripts/select_article_extractor.py").resolve()
cmd = [sys.executable, str(script_path), str(in_file)]
proc = subprocess.run(cmd, capture_output=True, text=True, check=False)
assert proc.returncode == 0
assert expected_out.exists()
with open(expected_out, "r", encoding="utf-8") as f:
data = json.load(f)
assert data["articles"][0]["selected_extractor"] == "newspaper4k"
def test_e2e_cli_subprocess_invalid_input(tmp_path: Path):
"""E2E: Executa CLI com JSON inválido e valida código de saída e erro no stderr."""
invalid_file = tmp_path / "broken.json"
invalid_file.write_text("not a valid json {", encoding="utf-8")
script_path = Path("scripts/select_article_extractor.py").resolve()
cmd = [sys.executable, str(script_path), str(invalid_file)]
proc = subprocess.run(cmd, capture_output=True, text=True, check=False)
assert proc.returncode == 2
assert "ERRO DE VALIDAÇÃO" in proc.stderr
def test_e2e_cli_subprocess_missing_file():
"""E2E: Executa CLI com arquivo inexistente e valida código 1 e mensagem no stderr."""
script_path = Path("scripts/select_article_extractor.py").resolve()
cmd = [sys.executable, str(script_path), "non_existent_file_12345.json"]
proc = subprocess.run(cmd, capture_output=True, text=True, check=False)
assert proc.returncode == 1
assert "ERRO DE ARQUIVO" in proc.stderr