feat(extractor): implement multi-engine article content extractor
- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability) - Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing) - Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts - Passed ruff linting/formatting and mypy type checking cleanly
This commit is contained in:
@@ -25,6 +25,10 @@
|
||||
- [Diferenciais Técnicos](#diferenciais-técnicos)
|
||||
- [Argumentos e Flags de Linha de Comando](#argumentos-e-flags-de-linha-de-comando)
|
||||
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso)
|
||||
- [3. Extrator e Parser Multimotor de Artigos](#3--extrator-e-parser-multimotor-de-artigos)
|
||||
- [Visão Geral e Tríplice Extração](#visão-geral-e-tríplice-extração)
|
||||
- [Argumentos e Flags CLI](#argumentos-e-flags-cli)
|
||||
- [Exemplos de Uso](#exemplos-de-uso)
|
||||
- [Estrutura do Projeto](#-estrutura-do-projeto)
|
||||
- [Testes e Qualidade de Código](#-testes-e-qualidade-de-código)
|
||||
- [Licença](#-licença)
|
||||
@@ -37,6 +41,7 @@ O **TextNLPClassifierApp** reúne ferramentas de engenharia de dados e processam
|
||||
|
||||
1. **`classify.py`**: Motor de classificação semântica e contextual que determina o grau de aderência e inerência de um documento Markdown em relação a uma entidade alvo definida em um **ECP Snapshot (Entity Context Profile)**.
|
||||
2. **`scripts/extract_google_news.py`**: Extrator de notícias por palavra-chave, idioma e região geográfica que utiliza o motor stealth **Foxcape** (em modo headless), decodificação paralela de URLs para os links reais dos portais de notícias e feedback em tempo real.
|
||||
3. **`scripts/extract_article_contents.py`**: Extrator e parser de artigos multimotor com navegação stealth Foxcape headless e extração combinada via **Trafilatura**, **Newspaper4k** (NLP) e **Readability**, consolidando texto higienizado, autores, datas, imagens e resumos em JSON estruturado.
|
||||
|
||||
---
|
||||
|
||||
@@ -208,6 +213,49 @@ python scripts/extract_google_news.py -q "inteligência artificial" -s | jq '.it
|
||||
|
||||
---
|
||||
|
||||
## 3. 📰 Extrator e Parser Multimotor de Artigos
|
||||
|
||||
### Visão Geral e Tríplice Extração
|
||||
|
||||
O script `scripts/extract_article_contents.py` lê os arquivos JSON gerados pelo extrator do Google News (ou qualquer lista contendo `items` com `url`), acessa cada página via **Foxcape** em modo stealth headless (reutilizando uma única sessão de navegador ativa com espera do evento `domcontentloaded`), e executa simultaneamente 3 motores especializados de extração:
|
||||
|
||||
1. **Trafilatura**: Texto principal higienizado, autores, data de publicação, categorias, tags, URL canônica e payload estruturado nativo.
|
||||
2. **Newspaper4k**: Artigo completo, autores, imagens (`top_image` e galeria), resumo automático e palavras-chave (*keywords*) extraídas por NLP nativo.
|
||||
3. **Readability (`readability-lxml`)**: Miolo limpo em HTML sem anúncios ou scripts supérfluos, títulos e texto puro formatado.
|
||||
|
||||
O JSON final consolidado é salvo em `out/` com descarte de strings HTML brutas para manter o arquivo leve e veloz.
|
||||
|
||||
### Argumentos e Flags CLI
|
||||
|
||||
| Parâmetro | Tipo | Padrão | Descrição |
|
||||
|---|---|---|---|
|
||||
| `-i, --input` | Caminho (obrigatório) | — | Arquivo JSON de busca de notícias (ex: `out/river_plate.json`). |
|
||||
| `-o, --output` | Caminho (opcional) | `<input_stem>_extracted.json` | Arquivo JSON de destino consolidado. |
|
||||
| `-l, --limit` | Inteiro (opcional) | Todos | Limita a quantidade máxima de notícias processadas. |
|
||||
| `--lang, --language` | String (opcional) | Do JSON / `en` | Sobrescreve o código de idioma para o NLP do Newspaper4k (ex: `pt`, `es`, `en`). |
|
||||
| `-t, --timeout` | Inteiro (opcional) | `30` | Timeout em segundos por página no Foxcape. |
|
||||
| `-s, --silent` | Flag booleana | `False` | Suprime logs informativos de progresso no `stderr`. |
|
||||
|
||||
### Exemplos de Uso
|
||||
|
||||
#### 1. Extração Completa Automática
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json
|
||||
# Gera automaticamente out/river_plate_extracted.json
|
||||
```
|
||||
|
||||
#### 2. Amostragem Rápida (Limit 2 Notícias)
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json --limit 2
|
||||
```
|
||||
|
||||
#### 3. Destino Customizado e Timeout Ajustado
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/petrobras_result.json -o out/petrobras_full.json --timeout 45
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📁 Estrutura do Projeto
|
||||
|
||||
```text
|
||||
@@ -215,19 +263,21 @@ TextNLPClassifierApp/
|
||||
├── classify.py # CLI principal do Classificador de Inerência
|
||||
├── scripts/
|
||||
│ ├── __init__.py # Pacote utilitário de scripts
|
||||
│ └── extract_google_news.py # CLI de Extração de Manchetes do Google News
|
||||
│ ├── extract_google_news.py # CLI de Extração de Manchetes do Google News
|
||||
│ └── extract_article_contents.py # CLI de Extração e Parsing Multimotor de Artigos
|
||||
├── src/ # Módulos centrais do classificador
|
||||
│ ├── classifier.py # Orquestrador de classificação (Tier 1, 2, 3)
|
||||
│ ├── models.py # Modelos de dados e esquemas (ECPSnapshot, Decision)
|
||||
│ ├── preprocessor.py # Normalização de texto e detecção de idioma
|
||||
│ └── adapters/ # Adaptadores opcionais de Embeddings e LLM
|
||||
├── specs/ # Especificações e planos arquiteturais (Speckit)
|
||||
│ ├── 001-nlp-classifier/ # Especificações do classificador
|
||||
│ └── 002-google-news-extractor/ # Especificações do extrator de notícias
|
||||
│ ├── 001-multilingual-entity-classifier/
|
||||
│ ├── 002-google-news-extractor/
|
||||
│ └── 003-article-content-extractor/ # Specs da feature de extração multimotor
|
||||
├── tests/ # Suíte de testes automatizados
|
||||
│ ├── fixtures/ # Amostras de ECP, Markdown e XML RSS
|
||||
│ ├── test_classifier.py # Testes unitários do classificador
|
||||
│ └── test_extract_google_news.py # Testes unitários e testes E2E ao vivo
|
||||
│ ├── test_classifier.py
|
||||
│ ├── test_extract_google_news.py
|
||||
│ └── test_extract_article_contents.py # Testes do extrator de conteúdo
|
||||
├── requirements.txt # Dependências do projeto
|
||||
├── pyproject.toml # Configurações de ferramentas (pytest, ruff, mypy)
|
||||
└── README.md # Documentação principal
|
||||
|
||||
+1
-2
@@ -4,13 +4,12 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from src import __version__
|
||||
from src.models import ECPSnapshot, ClassificationError, ErrorCode
|
||||
from src.classifier import InherenceClassifier
|
||||
from src.models import ClassificationError, ECPSnapshot, ErrorCode
|
||||
|
||||
# Ensure UTF-8 output streams across all platforms
|
||||
if hasattr(sys.stdout, "reconfigure"):
|
||||
|
||||
@@ -63,7 +63,9 @@ class ExtractNewsInputDTO(BaseModel):
|
||||
|
||||
keyword: str = Field(..., description="Palavra ou expressão de busca")
|
||||
language: str = Field(..., description="Código do idioma (ex: 'es', 'pt', 'en')")
|
||||
max_pages: int = Field(default=3, ge=1, le=10, description="Quantidade de páginas para extrair (1 a 10)")
|
||||
max_pages: int = Field(
|
||||
default=3, ge=1, le=10, description="Quantidade de páginas para extrair (1 a 10)"
|
||||
)
|
||||
```
|
||||
|
||||
| Campo | Tipo | Obrigatório | Descrição |
|
||||
@@ -90,7 +92,9 @@ class SearchQuery:
|
||||
raise InvalidSearchQueryError("A palavra-chave não pode ser vazia.")
|
||||
|
||||
if not self.language or len(self.language.strip()) < 2:
|
||||
raise InvalidSearchQueryError("O idioma deve conter pelo menos 2 caracteres (ex: 'es', 'pt', 'en').")
|
||||
raise InvalidSearchQueryError(
|
||||
"O idioma deve conter pelo menos 2 caracteres (ex: 'es', 'pt', 'en')."
|
||||
)
|
||||
|
||||
if self.max_pages < 1 or self.max_pages > 10:
|
||||
raise InvalidSearchQueryError("O número máximo de páginas deve estar entre 1 e 10.")
|
||||
@@ -123,6 +127,7 @@ class ExtractNewsUseCase:
|
||||
from googlenews_etl.infrastructure.adapters.google_news_extractor_adapter import (
|
||||
GoogleNewsExtractorAdapter,
|
||||
)
|
||||
|
||||
self.extractor = GoogleNewsExtractorAdapter()
|
||||
else:
|
||||
self.extractor = extractor
|
||||
@@ -199,7 +204,9 @@ class GoogleNewsExtractorAdapter(NewsExtractorPort):
|
||||
resolve_final_urls: bool = True,
|
||||
) -> None:
|
||||
self.impersonate = impersonate
|
||||
self.rate_limiter = rate_limiter or RateLimiterService(min_delay_seconds=0.5, max_delay_seconds=1.0)
|
||||
self.rate_limiter = rate_limiter or RateLimiterService(
|
||||
min_delay_seconds=0.5, max_delay_seconds=1.0
|
||||
)
|
||||
self.url_resolver = url_resolver or PlaywrightUrlResolverAdapter()
|
||||
self.resolve_final_urls = resolve_final_urls
|
||||
self.session = requests.Session(impersonate=self.impersonate)
|
||||
@@ -351,8 +358,12 @@ class PlaywrightUrlResolverAdapter(UrlResolverPort):
|
||||
if not any(
|
||||
x in u
|
||||
for x in [
|
||||
"google.", "gstatic.", "googleapis.", "googletagmanager.",
|
||||
"w3.org", "schema.org",
|
||||
"google.",
|
||||
"gstatic.",
|
||||
"googleapis.",
|
||||
"googletagmanager.",
|
||||
"w3.org",
|
||||
"schema.org",
|
||||
]
|
||||
):
|
||||
if not target_url or target_url == url:
|
||||
@@ -554,8 +565,8 @@ from googlenews_etl.application.use_cases.extract_news_use_case import ExtractNe
|
||||
dto_in = ExtractNewsInputDTO(keyword="inteligencia artificial", language="pt", max_pages=1)
|
||||
resultado = ExtractNewsUseCase().execute(dto_in)
|
||||
|
||||
print(resultado.total_itens) # ex: 10
|
||||
print(resultado.items[0].titulo) # título da primeira manchete
|
||||
print(resultado.items[0].url) # URL final resolvida
|
||||
print(resultado.total_itens) # ex: 10
|
||||
print(resultado.items[0].titulo) # título da primeira manchete
|
||||
print(resultado.items[0].url) # URL final resolvida
|
||||
print(resultado.items[0].quando_publicado) # data crua do RSS
|
||||
```
|
||||
@@ -0,0 +1,255 @@
|
||||
# PRD — Extrator e Parser de Artigos Multimotor (Foxcape + Trafilatura + Newspaper4k + Readability)
|
||||
|
||||
> **Status:** Proposto / Planejamento
|
||||
> **Versão:** 1.0.0
|
||||
> **Data:** 2026-08-20
|
||||
> **Autor:** Antigravity AI / DunaMedia
|
||||
> **Localização:** `docs/prd_extrator_artigos_nlp.md`
|
||||
|
||||
---
|
||||
|
||||
## 1. Visão Geral e Contexto
|
||||
|
||||
### 1.1 Objetivo do Produto
|
||||
Construir um pipeline/script autônomo e resiliente em Python para **extração profunda e enriquecimento de conteúdo de artigos de notícias** a partir de listagens JSON previamente geradas (ex.: pelo extrator do Google News).
|
||||
|
||||
O pipeline utiliza o **Foxcape** em modo *headless* para acessar furtivamente as URLs finais, aguardar a renderização completa da árvore DOM e coletar o HTML íntegro. Em seguida, processa esse HTML simultaneamente através de **três motores consagrados de extração de conteúdo (NLP / Web Content Extraction)**:
|
||||
1. **Trafilatura**
|
||||
2. **Newspaper4k**
|
||||
3. **Readability (`readability-lxml`)**
|
||||
|
||||
O resultado consolidado com o máximo de informações extraídas por cada motor é exportado em formato JSON estruturado na pasta `out/`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Personas e Casos de Uso
|
||||
|
||||
| Persona | Necessidade | Benefício |
|
||||
|---|---|---|
|
||||
| **Engenheiro de Dados / ETL** | Ingerir em lote o conteúdo completo de notícias a partir de arquivos JSON de busca. | Automação resiliente sem necessidade de criar parsers manuais para cada portal de notícias. |
|
||||
| **Cientista de Dados / NLP** | Comparar ou combinar diferentes abordagens de extração de texto, resumos e metadados. | Obter em um único JSON os outputs estruturados de 3 extratores líderes da indústria. |
|
||||
| **Operador de Terminal / Analista** | Executar o script via CLI informando o arquivo de entrada e acompanhando o progresso em tempo real. | Visibilidade clara via `stderr` sem poluir a saída JSON no `stdout`. |
|
||||
|
||||
---
|
||||
|
||||
## 3. Arquitetura e Fluxo do Sistema
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A[Arquivo JSON de Entrada\n out/river_plate.json] --> B[Script CLI / Use Case\n Carrega lista de notícias]
|
||||
B --> C[Iterador de Artigos]
|
||||
C --> D[Motor Foxcape Headless\n Navega até a URL final]
|
||||
D --> E[Aguarda DOM carregar\n domcontentloaded / load]
|
||||
E --> F[Coleta HTML Renderizado Completo]
|
||||
|
||||
F --> G1[Motor 1: Trafilatura\n Texto limpo, autor, data, tags, meta]
|
||||
F --> G2[Motor 2: Newspaper4k\n Artigo, autores, resumo NLP, keywords, top image]
|
||||
F --> G3[Motor 3: Readability\n Miolo HTML limpo, texto limpo, título]
|
||||
|
||||
G1 --> H[Agregador de Extrações]
|
||||
G2 --> H
|
||||
G3 --> H
|
||||
|
||||
H --> I[JSON Consolidado de Saída\n out/extracted_articles_*.json]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Requisitos Funcionais (FR)
|
||||
|
||||
### RF01 — Ingestão de Arquivo JSON de Notícias
|
||||
- O sistema deve aceitar como entrada um arquivo JSON localizado em `out/` (ou caminho customizado via flag CLI `--input` / `-i`).
|
||||
- Deve validar a presença da lista `items` e extrair as propriedades fundamentais de cada item (`url`, `titulo`, `subtitulo`, `quando_publicado`, `pagina`).
|
||||
- Deve suportar opções para limitar o processamento a N itens (ex.: `--limit 5` para testes rápidos).
|
||||
|
||||
### RF02 — Navegação e Coleta com Foxcape (Obrigatório)
|
||||
- O acesso às páginas **deve obrigatoriamente** ser realizado com o pacote `foxcape`, aproveitando seu motor anti-bot (Camoufox + evasão de fingerprinting).
|
||||
- Deve executar em modo `headless=True` por padrão (`FoxcapeConfig(headless=True, humanize=False)`).
|
||||
- Deve navegar até a URL de cada notícia, aguardar o evento de carregamento do DOM (`wait_until="domcontentloaded"` com fallback para timeout) e extrair a string HTML completa e íntegra da página.
|
||||
- Deve reaproveitar a mesma instância/sessão do navegador entre as requisições do lote para otimizar velocidade e consumo de memória.
|
||||
|
||||
### RF03 — Extração Máxima com Trafilatura
|
||||
- Para cada HTML coletado, executar o motor `trafilatura`.
|
||||
- Extrair todos os campos disponíveis:
|
||||
- Texto principal limpo (`raw_text` / `text`).
|
||||
- Título (`title`).
|
||||
- Autor(es) (`author`).
|
||||
- Data de publicação (`date`).
|
||||
- Descrição / Subtítulo (`description`).
|
||||
- Categorias e Tags (`categories`, `tags`).
|
||||
- URL canônica (`canonical_url`).
|
||||
- Comentários estruturados (se disponíveis).
|
||||
- Extração em formato JSON estruturado nativo do Trafilatura.
|
||||
|
||||
### RF04 — Extração Máxima com Newspaper4k
|
||||
- Para cada HTML coletado, instanciar `Article(url, input_html=html)`.
|
||||
- Executar `parse()` e os métodos de NLP:
|
||||
- Título do artigo (`title`).
|
||||
- Autores (`authors`).
|
||||
- Data de publicação (`publish_date`).
|
||||
- Texto completo limpo (`text`).
|
||||
- Resumo gerado por NLP (`summary`).
|
||||
- Palavras-chave extraídas por NLP (`keywords`).
|
||||
- Imagem de destaque (`top_image`) e lista de todas as imagens (`images`).
|
||||
- Metadados brutos OpenGraph e Schema.org (`meta_data`).
|
||||
|
||||
### RF05 — Extração com Readability (`readability-lxml`)
|
||||
- Para cada HTML coletado, processar via `Document(html)`.
|
||||
- Extrair:
|
||||
- Título limpo (`title()` e `short_title()`).
|
||||
- Conteúdo limpo em HTML sem boilerplates/anúncios (`summary()`).
|
||||
- Texto limpo puro derivado do conteúdo principal.
|
||||
|
||||
### RF06 — Consolidação e Exportação de Resultados
|
||||
- O sistema deve agregar o resultado dos 3 extratores em um único objeto por notícia.
|
||||
- Deve salvar o resultado final em formato JSON na pasta `out/` (ex.: `out/extracted_<nome_do_arquivo_origem>.json` ou caminho fornecido via `--output` / `-o`).
|
||||
- Deve incluir metadados de execução global:
|
||||
- `source_file`: caminho do JSON de entrada.
|
||||
- `processed_at`: timestamp ISO 8601 da execução.
|
||||
- `total_articles`: quantidade de artigos no arquivo de entrada.
|
||||
- `successful_articles`: quantidade processada com sucesso.
|
||||
- `failed_articles`: quantidade com falha.
|
||||
- `articles`: array com os objetos consolidados.
|
||||
|
||||
### RF07 — Interface de Linha de Comando (CLI) e Logs em Tempo Real
|
||||
- Disponibilizar interface CLI:
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json -o out/river_plate_extracted.json
|
||||
```
|
||||
- Flags suportadas:
|
||||
- `-i, --input`: Caminho do arquivo JSON de entrada (obrigatório).
|
||||
- `-o, --output`: Caminho do arquivo JSON de saída (opcional, padrão: `out/extracted_<basename>.json`).
|
||||
- `-l, --limit`: Limitar quantidade de artigos processados (opcional).
|
||||
- `-t, --timeout`: Timeout em segundos por página no Foxcape (padrão: 30s).
|
||||
- `-s, --silent`: Suprime logs visuais no `stderr`.
|
||||
- Logs informativos enviados para `sys.stderr` com emojis e timestamps:
|
||||
- `[INFO] 🚀 Iniciando extração de N artigos a partir de ...`
|
||||
- `[INFO] 🌐 [1/10] Foxcape navegando: https://...`
|
||||
- `[INFO] ⚙️ [1/10] Processando extratores (Trafilatura, Newspaper4k, Readability)...`
|
||||
- `[INFO] ✅ [1/10] Concluído com sucesso!`
|
||||
- `[INFO] 💾 Salvo com sucesso em: out/...`
|
||||
|
||||
### RF08 — Tolerância a Falhas e Resiliência
|
||||
- Se o acesso a uma URL falhar (ex.: timeout, 404, bloqueio severo), registrar o erro no item individual (`status: "failed"`, `error_message: "..."`) e prosseguir imediatamente para a próxima notícia do lote.
|
||||
- Se um dos 3 extratores falhar em um HTML específico, os outros 2 extratores devem continuar operando normalmente, registrando o erro no campo correspondente ao motor que falhou.
|
||||
|
||||
---
|
||||
|
||||
## 5. Requisitos Não Funcionais (NFR)
|
||||
|
||||
| ID | Requisito | Critério |
|
||||
|---|---|---|
|
||||
| **RNF01** | **Furtividade Anti-Bot** | Utilizar Foxcape/Camoufox para evitar bloqueios de TLS/JA3 e Cloudflare. |
|
||||
| **RNF02** | **Eficiência de Recursos** | Manter instância persistente do navegador aberta em lote em vez de instanciar/fechar o browser a cada URL. |
|
||||
| **RNF03** | **Segregação de Streams** | `sys.stderr` exclusivo para logs; `sys.stdout` reservado para output JSON caso não seja especificado arquivo. |
|
||||
| **RNF04** | **Modularidade** | Código estruturado em classes/módulos independentes para cada extrator (`TrafilaturaExtractor`, `NewspaperExtractor`, `ReadabilityExtractor`). |
|
||||
| **RNF05** | **Compatibilidade de Plataforma** | Totalmente funcional em Windows, Linux e macOS (Python 3.10+). |
|
||||
|
||||
---
|
||||
|
||||
## 6. Esquema de Dados (Contratos de Interface)
|
||||
|
||||
### 6.1 Esquema do JSON de Entrada (`Input`)
|
||||
```json
|
||||
{
|
||||
"query": "River Plate",
|
||||
"language": "es",
|
||||
"locale": "AR",
|
||||
"total_paginas": 1,
|
||||
"total_itens": 2,
|
||||
"scraped_at": "2026-08-20T14:37:38.233557+00:00",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"subtitulo": "River empató sem gols...",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 03:27:26 GMT",
|
||||
"url": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"pagina": 1
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 6.2 Esquema do JSON Consolidado de Saída (`Output`)
|
||||
```json
|
||||
{
|
||||
"source_file": "out/river_plate.json",
|
||||
"processed_at": "2026-08-20T15:00:00.000000+00:00",
|
||||
"total_articles": 1,
|
||||
"successful_articles": 1,
|
||||
"failed_articles": 0,
|
||||
"articles": [
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo_original": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"subtitulo_original": "River empató sem gols...",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 03:27:26 GMT",
|
||||
"url_original": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"pagina": 1
|
||||
},
|
||||
"extraction_status": "success",
|
||||
"error_message": null,
|
||||
"crawled_url": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"page_title": "Los puntajes de River...",
|
||||
"http_status": 200,
|
||||
"trafilatura": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe...",
|
||||
"author": "Ernesto Provitilo",
|
||||
"date": "2026-08-20",
|
||||
"description": "El análisis uno por uno...",
|
||||
"categories": ["River Plate", "Copa Sudamericana"],
|
||||
"tags": ["River", "Santa Fe"],
|
||||
"canonical_url": "https://www.tycsports.com/...",
|
||||
"text": "Franco Armani (6): Seguro en las pocas llegadas...",
|
||||
"raw_json": {}
|
||||
},
|
||||
"newspaper4k": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"authors": ["Ernesto Provitilo"],
|
||||
"publish_date": "2026-08-20T03:27:26",
|
||||
"text": "Franco Armani (6): Seguro en las pocas llegadas...",
|
||||
"summary": "Resumo gerado por NLP do newspaper...",
|
||||
"keywords": ["river", "santa fe", "puntajes", "armani"],
|
||||
"top_image": "https://media.tycsports.com/...",
|
||||
"images": ["https://media.tycsports.com/..."],
|
||||
"meta_data": {}
|
||||
},
|
||||
"readability": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"short_title": "Los puntajes de River",
|
||||
"cleaned_html": "<div><p>Franco Armani (6)...</p></div>",
|
||||
"cleaned_text": "Franco Armani (6): Seguro en las pocas llegadas..."
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Dependências do Projeto
|
||||
|
||||
As seguintes bibliotecas Python são necessárias para viabilizar este PRD:
|
||||
|
||||
```toml
|
||||
[dependencies]
|
||||
foxcape = ">=0.1.1"
|
||||
trafilatura = ">=1.8.0"
|
||||
newspaper4k = ">=0.9.3"
|
||||
readability-lxml = ">=0.8.1"
|
||||
beautifulsoup4 = ">=4.12.0"
|
||||
lxml = ">=4.9.0"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Critérios de Aceite
|
||||
|
||||
1. **Execução de ponta a ponta**: Executar o script apontando para `out/river_plate.json` e gerar com sucesso `out/river_plate_extracted.json`.
|
||||
2. **Uso Mandatório do Foxcape**: Todas as páginas HTML devem ser renderizadas e baixadas via Foxcape headless.
|
||||
3. **Completude das Extrações**:
|
||||
- O objeto `trafilatura` deve conter título, autor, data e texto limpo.
|
||||
- O objeto `newspaper4k` deve conter NLP summary, keywords, top_image e texto limpo.
|
||||
- O objeto `readability` deve conter HTML limpo e texto limpo.
|
||||
4. **Resiliência a Falhas**: Artigos com falha de conexão não devem quebrar o script e devem vir identificados como `extraction_status: "failed"`.
|
||||
5. **Logs Informativos**: O operador visualiza no terminal o progresso item a item através do `stderr`.
|
||||
@@ -44,10 +44,10 @@
|
||||
"42": "1. Input Schemas",
|
||||
"43": "2. Basic CLI Usage Examples",
|
||||
"44": "2. Standard Streams & Exit Codes",
|
||||
"45": "ClassificationResult",
|
||||
"45": "classifier.py",
|
||||
"46": "InherenceClassifier",
|
||||
"47": "detect_language",
|
||||
"48": "test_models.py",
|
||||
"48": "ClassificationError",
|
||||
"49": "content_northvolt_de.md",
|
||||
"50": "content_presal_pt.md",
|
||||
"51": "content_tangential_es.md",
|
||||
@@ -79,7 +79,7 @@
|
||||
"77": "pt/tangential.md",
|
||||
"78": "tests/__init__.py",
|
||||
"79": "text-nlp-classifier",
|
||||
"80": "get_hl_gl_ceid",
|
||||
"80": "test_extract_article_contents.py",
|
||||
"81": "Extrator de Notícias do Google News — Guia Completo de Funcionamento",
|
||||
"82": "extract_google_news.py",
|
||||
"83": "ExtractionResult",
|
||||
@@ -97,10 +97,21 @@
|
||||
"95": "CLI Contract: Google News Headlines Extractor",
|
||||
"96": "readiness.md",
|
||||
"97": "🧠 TextNLPClassifierApp",
|
||||
"98": "build_parser",
|
||||
"98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor",
|
||||
"99": "sample_rss_xml",
|
||||
"100": "classifier.py",
|
||||
"100": "main",
|
||||
"101": "ECPSnapshot",
|
||||
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"103": "main"
|
||||
"103": "4. Requisitos Funcionais (FR)",
|
||||
"104": "Tasks: Article Content Multi-Engine Extractor",
|
||||
"105": "models.py",
|
||||
"106": "Implementation Plan: Article Content Multi-Engine Extractor",
|
||||
"107": "2. Cenários de Validação",
|
||||
"108": "1. Technical Decisions & Tradeoffs",
|
||||
"109": "Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"110": "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"112": "CLI Contract: Article Content Multi-Engine Extractor",
|
||||
"113": "001-multilingual-entity-classifier/spec.md",
|
||||
"114": "JSON Schema Contract: Article Content Multi-Engine Extractor"
|
||||
}
|
||||
|
||||
@@ -1 +1 @@
|
||||
{"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "b493e66e33005daa", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "62a40bb7296050db", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "0172259097d0974e", "46": "d2ae82bc1907fef3", "47": "7be46d99756bf41e", "48": "a94841cd5849d3f5", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6", "80": "10ea6ef8e60efb58", "81": "a38f84ae3d895236", "82": "5efe702565242e04", "83": "f74492db7d32ba1e", "84": "c0d4860b83f7cb84", "85": "f8bfd0cfe9e8b478", "86": "410d15a346bd5894", "87": "6b41d288cfd834ab", "88": "5aa6db96312a8811", "89": "80225792bb62ba04", "90": "e553a45aca39b563", "91": "d4579c5b7aa2742a", "92": "7b9ba7c3bff11361", "93": "71cd9c1fa4a857f0", "94": "34cd980be3c32d21", "95": "970093453f3b7d90", "96": "9e96780a2b7c4bd6", "97": "3befc42bd6078583", "98": "32f0d209b4ebd09b", "99": "8968e9e7d55afcbe", "100": "c14800d0ab2e5026", "101": "d7e7752d280777b1", "102": "6aa00d5a83295f11", "103": "cbb35e88e6b7e1ac"}
|
||||
{"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "5b8462a3f82d188c", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "66b69189c0acc3ff", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "70b30e0a139bbd59", "46": "58f3596265f5f902", "47": "f89e819a4cc2d299", "48": "b45478cb1443df68", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6", "80": "6ff8a97e63c9a2f3", "81": "a38f84ae3d895236", "82": "08e48bd11f9714df", "83": "5095122914e83cf5", "84": "1aef305bd7d7d63f", "85": "f8bfd0cfe9e8b478", "86": "410d15a346bd5894", "87": "6b41d288cfd834ab", "88": "5aa6db96312a8811", "89": "80225792bb62ba04", "90": "fd291228c3311f40", "91": "d4579c5b7aa2742a", "92": "7b9ba7c3bff11361", "93": "71cd9c1fa4a857f0", "94": "34cd980be3c32d21", "95": "970093453f3b7d90", "96": "9e96780a2b7c4bd6", "97": "a4e57776e229c794", "98": "089ea6a55861c693", "99": "8968e9e7d55afcbe", "100": "daa6156e559cda08", "101": "14e5b6323d34c337", "102": "6aa00d5a83295f11", "103": "f58668f5b10ccdeb", "104": "4ec787414cc6f50b", "105": "de672873b6dec6e2", "106": "edcd5d9bb3c4b00f", "107": "37f2f47110fe3eaa", "108": "b7ad5abb1da8cf8d", "109": "cb48a9c4f54efa38", "110": "f6dd36fd7f3edbe5", "111": "2925b620f0b1fd17", "112": "d8b3099917c3b711", "113": "3bb61caa0302c804", "114": "0d4f1d08dd056bb9"}
|
||||
@@ -4,7 +4,7 @@
|
||||
"2": "SpecKit Utilities",
|
||||
"3": "Graphify Commands",
|
||||
"4": "speckit-analyze/SKILL.md",
|
||||
"5": "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"5": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"6": "Feature Specification Template",
|
||||
"7": "Graphify Rules",
|
||||
"8": "Implementation Planning",
|
||||
@@ -44,10 +44,10 @@
|
||||
"42": "1. Input Schemas",
|
||||
"43": "2. Basic CLI Usage Examples",
|
||||
"44": "2. Standard Streams & Exit Codes",
|
||||
"45": "ECPSnapshot",
|
||||
"46": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"45": "LocalEmbeddingsAdapter",
|
||||
"46": "InherenceClassifier",
|
||||
"47": "detect_language",
|
||||
"48": "test_models.py",
|
||||
"48": "classifier.py",
|
||||
"49": "content_northvolt_de.md",
|
||||
"50": "content_presal_pt.md",
|
||||
"51": "content_tangential_es.md",
|
||||
@@ -79,7 +79,7 @@
|
||||
"77": "pt/tangential.md",
|
||||
"78": "tests/__init__.py",
|
||||
"79": "text-nlp-classifier",
|
||||
"80": "get_hl_gl_ceid",
|
||||
"80": "test_extract_article_contents.py",
|
||||
"81": "Extrator de Notícias do Google News — Guia Completo de Funcionamento",
|
||||
"82": "extract_google_news.py",
|
||||
"83": "ExtractionResult",
|
||||
@@ -96,6 +96,22 @@
|
||||
"94": "Specification Quality Checklist: Google News Headlines Extractor",
|
||||
"95": "CLI Contract: Google News Headlines Extractor",
|
||||
"96": "readiness.md",
|
||||
"98": "build_parser",
|
||||
"99": "sample_rss_xml"
|
||||
"97": "🧠 TextNLPClassifierApp",
|
||||
"98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor",
|
||||
"99": "sample_rss_xml",
|
||||
"100": "models.py",
|
||||
"101": "ECPSnapshot",
|
||||
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"103": "4. Requisitos Funcionais (FR)",
|
||||
"104": "Tasks: Article Content Multi-Engine Extractor",
|
||||
"105": "ClassificationResult",
|
||||
"106": "Implementation Plan: Article Content Multi-Engine Extractor",
|
||||
"107": "2. Cenários de Validação",
|
||||
"108": "1. Technical Decisions & Tradeoffs",
|
||||
"109": "Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"110": "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"111": "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"112": "CLI Contract: Article Content Multi-Engine Extractor",
|
||||
"113": "001-multilingual-entity-classifier/spec.md",
|
||||
"114": "JSON Schema Contract: Article Content Multi-Engine Extractor"
|
||||
}
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
||||
|
||||
## Corpus Check
|
||||
- 147 files · ~65,826 words
|
||||
- 161 files · ~78,476 words
|
||||
- Verdict: corpus is large enough that graph structure adds value.
|
||||
|
||||
## Summary
|
||||
- 775 nodes · 931 edges · 99 communities (61 shown, 38 thin omitted)
|
||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
|
||||
- 1000 nodes · 1226 edges · 115 communities (77 shown, 38 thin omitted)
|
||||
- Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 39 edges (avg confidence: 0.95)
|
||||
- Token cost: 0 input · 0 output
|
||||
|
||||
## Graph Freshness
|
||||
- Built from commit: `67cc40f9`
|
||||
- Built from commit: `6e3d5761`
|
||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||
- Run `graphify update .` after code changes (no API cost).
|
||||
|
||||
@@ -20,7 +20,7 @@
|
||||
- SpecKit Utilities
|
||||
- Graphify Commands
|
||||
- speckit-analyze/SKILL.md
|
||||
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- Feature Specification Template
|
||||
- Graphify Rules
|
||||
- Implementation Planning
|
||||
@@ -56,10 +56,10 @@
|
||||
- 1. Input Schemas
|
||||
- 2. Basic CLI Usage Examples
|
||||
- 2. Standard Streams & Exit Codes
|
||||
- ECPSnapshot
|
||||
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- LocalEmbeddingsAdapter
|
||||
- InherenceClassifier
|
||||
- detect_language
|
||||
- test_models.py
|
||||
- classifier.py
|
||||
- content_northvolt_de.md
|
||||
- content_presal_pt.md
|
||||
- content_tangential_es.md
|
||||
@@ -91,7 +91,7 @@
|
||||
- pt/tangential.md
|
||||
- tests/__init__.py
|
||||
- text-nlp-classifier
|
||||
- get_hl_gl_ceid
|
||||
- test_extract_article_contents.py
|
||||
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||
- extract_google_news.py
|
||||
- ExtractionResult
|
||||
@@ -107,37 +107,52 @@
|
||||
- 1. Entidades de Domínio & DTOs
|
||||
- Specification Quality Checklist: Google News Headlines Extractor
|
||||
- CLI Contract: Google News Headlines Extractor
|
||||
- build_parser
|
||||
- 🧠 TextNLPClassifierApp
|
||||
- Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
|
||||
- sample_rss_xml
|
||||
- models.py
|
||||
- ECPSnapshot
|
||||
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- 4. Requisitos Funcionais (FR)
|
||||
- Tasks: Article Content Multi-Engine Extractor
|
||||
- ClassificationResult
|
||||
- Implementation Plan: Article Content Multi-Engine Extractor
|
||||
- 2. Cenários de Validação
|
||||
- 1. Technical Decisions & Tradeoffs
|
||||
- Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- CLI Contract: Article Content Multi-Engine Extractor
|
||||
- JSON Schema Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## God Nodes (most connected - your core abstractions)
|
||||
1. `ECPSnapshot` - 31 edges
|
||||
2. `InherenceClassifier` - 25 edges
|
||||
3. `DecisionCategory` - 18 edges
|
||||
4. `ClassificationResult` - 17 edges
|
||||
5. `main()` - 14 edges
|
||||
5. `process_batch()` - 15 edges
|
||||
6. `LocalEmbeddingsAdapter` - 14 edges
|
||||
7. `LLMFallbackAdapter` - 14 edges
|
||||
8. `detect_language()` - 14 edges
|
||||
9. `Tasks: [FEATURE NAME]` - 13 edges
|
||||
10. `SearchQuery` - 12 edges
|
||||
9. `main()` - 13 edges
|
||||
10. `ArticleCrawler` - 13 edges
|
||||
|
||||
## Surprising Connections (you probably didn't know these)
|
||||
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
|
||||
tests/test_extract_google_news.py → scripts/extract_google_news.py
|
||||
- `classifier()` --uses--> `InherenceClassifier` [INFERRED]
|
||||
tests/test_benchmark_24.py → src/classifier.py
|
||||
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `test_ecp_snapshot_missing_required()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `petrobras_ecp()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_classifier.py → src/models.py
|
||||
|
||||
## Import Cycles
|
||||
- None detected.
|
||||
|
||||
## Communities (99 total, 38 thin omitted)
|
||||
## Communities (115 total, 38 thin omitted)
|
||||
|
||||
### Community 0 - "Task Planning"
|
||||
Cohesion: 0.07
|
||||
@@ -159,9 +174,9 @@ Nodes (24): For /graphify add and --watch, For /graphify query, For the commit h
|
||||
Cohesion: 0.08
|
||||
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
|
||||
|
||||
### Community 5 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.05
|
||||
Nodes (34): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+26 more)
|
||||
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.15
|
||||
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
|
||||
|
||||
### Community 6 - "Feature Specification Template"
|
||||
Cohesion: 0.15
|
||||
@@ -260,8 +275,8 @@ Cohesion: 0.50
|
||||
Nodes (3): Boundaries, Output, Scan
|
||||
|
||||
### Community 40 - "main"
|
||||
Cohesion: 0.19
|
||||
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
|
||||
Cohesion: 0.14
|
||||
Nodes (17): ArgumentParser, CaptureFixture, build_parser(), main(), Cria e configura o parser de argumentos CLI., Ponto de entrada do script CLI., Path, Valida execução padrão do CLI com saída JSON no stdout. (+9 more)
|
||||
|
||||
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.22
|
||||
@@ -279,41 +294,41 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
|
||||
Cohesion: 0.29
|
||||
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
||||
|
||||
### Community 45 - "ECPSnapshot"
|
||||
Cohesion: 0.06
|
||||
Nodes (55): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+47 more)
|
||||
### Community 45 - "LocalEmbeddingsAdapter"
|
||||
Cohesion: 0.14
|
||||
Nodes (8): LocalEmbeddingsAdapter, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface(), test_llm_adapter_interface()
|
||||
|
||||
### Community 46 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.15
|
||||
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
|
||||
### Community 46 - "InherenceClassifier"
|
||||
Cohesion: 0.12
|
||||
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more)
|
||||
|
||||
### Community 47 - "detect_language"
|
||||
Cohesion: 0.14
|
||||
Nodes (21): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+13 more)
|
||||
|
||||
### Community 48 - "test_models.py"
|
||||
Cohesion: 0.09
|
||||
Nodes (29): emit_error(), main(), parse_args(), Namespace, ClassificationError, ErrorCode, Any, extract_evidence_snippets() (+21 more)
|
||||
### Community 48 - "classifier.py"
|
||||
Cohesion: 0.24
|
||||
Nodes (11): Core deterministic classification engine (Tier 1 core)., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+3 more)
|
||||
|
||||
### Community 80 - "get_hl_gl_ceid"
|
||||
Cohesion: 0.25
|
||||
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
|
||||
### Community 80 - "test_extract_article_contents.py"
|
||||
Cohesion: 0.06
|
||||
Nodes (56): ArticleCrawler, extract_all_engines(), ExtractedArticle, ExtractionBatchReport, InputArticle, load_search_json(), log_info(), main() (+48 more)
|
||||
|
||||
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
|
||||
Cohesion: 0.08
|
||||
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
|
||||
|
||||
### Community 82 - "extract_google_news.py"
|
||||
Cohesion: 0.20
|
||||
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
|
||||
Cohesion: 0.15
|
||||
Nodes (18): extract_google_news(), _fetch_rss_content(), get_hl_gl_ceid(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Remove pontuação e espaços extras para comparação de redundância. (+10 more)
|
||||
|
||||
### Community 83 - "ExtractionResult"
|
||||
Cohesion: 0.29
|
||||
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
|
||||
|
||||
### Community 84 - "test_extract_google_news.py"
|
||||
Cohesion: 0.21
|
||||
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
|
||||
Cohesion: 0.15
|
||||
Nodes (15): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado. (+7 more)
|
||||
|
||||
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
|
||||
Cohesion: 0.14
|
||||
@@ -355,28 +370,88 @@ Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness,
|
||||
Cohesion: 0.33
|
||||
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
|
||||
|
||||
### Community 98 - "build_parser"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
|
||||
### Community 97 - "🧠 TextNLPClassifierApp"
|
||||
Cohesion: 0.06
|
||||
Nodes (33): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. Extração Completa Automática, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Amostragem Rápida (Limit 2 Notícias), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências (+25 more)
|
||||
|
||||
### Community 98 - "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.05
|
||||
Nodes (34): 1. Requirement Completeness, 2. Requirement Clarity & Non-Ambiguity, 3. Requirement Consistency & Data Contracts, 4. Scenario & Edge Case Coverage, 5. Non-Functional & Operational Readiness, Extraction Pipeline Checklist: Article Content Multi-Engine Extractor, Notes, Content Quality (+26 more)
|
||||
|
||||
### Community 99 - "sample_rss_xml"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
|
||||
|
||||
### Community 100 - "models.py"
|
||||
Cohesion: 0.15
|
||||
Nodes (17): emit_error(), main(), parse_args(), Namespace, Enum, ClassificationError, ErrorCode, MatchedGraphEntity (+9 more)
|
||||
|
||||
### Community 101 - "ECPSnapshot"
|
||||
Cohesion: 0.22
|
||||
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more)
|
||||
|
||||
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.14
|
||||
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
||||
|
||||
### Community 103 - "4. Requisitos Funcionais (FR)"
|
||||
Cohesion: 0.10
|
||||
Nodes (20): 1.1 Objetivo do Produto, 1. Visão Geral e Contexto, 2. Personas e Casos de Uso, 3. Arquitetura e Fluxo do Sistema, 4. Requisitos Funcionais (FR), 5. Requisitos Não Funcionais (NFR), 6.1 Esquema do JSON de Entrada (`Input`), 6.2 Esquema do JSON Consolidado de Saída (`Output`) (+12 more)
|
||||
|
||||
### Community 104 - "Tasks: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.11
|
||||
Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more)
|
||||
|
||||
### Community 105 - "ClassificationResult"
|
||||
Cohesion: 0.16
|
||||
Nodes (11): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional local vector embeddings adapter (Tier 2). Disabled by default.… (+3 more)
|
||||
|
||||
### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.17
|
||||
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Article Content Multi-Engine Extractor, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Next: `/speckit-tasks`)*, Project Structure (+3 more)
|
||||
|
||||
### Community 107 - "2. Cenários de Validação"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 1. Pré-requisitos, 2. Cenários de Validação, 3. Validação Automatizada de Testes, Cenário 1: Extração com Amostragem Rápida (Limit 2), Cenário 2: Caminho Customizado de Saída, Cenário 3: Modo Silencioso (`--silent`), Cenário 4: Resiliência contra URLs Inválidas, Quickstart & Validation Guide: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 108 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Navegação e Renderização com `foxcape` em Sessão Única, Decision 2: Orquestração Tripla de Extração de Conteúdo (NLP & Web Scraping), Decision 3: Resiliência e Isolamento de Falhas por Camada, Decision 4: Herança Inteligente de Idioma para NLP, Decision 5: Gerenciamento de Memória e Descarte do Raw HTML, Decision 6: Segregação de Streams e Feedback Visual em `stderr`, Research: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 109 - "Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.25
|
||||
Nodes (8): Complexity Tracking, Constitution Check, Documentation (this feature), Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC), Project Structure, Source Code (repository root), Summary, Technical Context
|
||||
|
||||
### Community 110 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.29
|
||||
Nodes (7): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
|
||||
### Community 111 - "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
|
||||
### Community 112 - "CLI Contract: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): 1. Comando de Execução, 2. Argumentos e Flags, 3. Códigos de Saída (Exit Codes), 4. Comportamento de Streams (I/O), CLI Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 114 - "JSON Schema Contract: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## Knowledge Gaps
|
||||
- **361 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+356 more)
|
||||
- **465 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+460 more)
|
||||
These have ≤1 connection - possible missing edges or undocumented components.
|
||||
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
|
||||
## Suggested Questions
|
||||
_Questions this graph is uniquely positioned to answer:_
|
||||
|
||||
- **Why does `main()` connect `test_models.py` to `main`, `ECPSnapshot`?**
|
||||
_High betweenness centrality (0.031) - this node is a cross-community bridge._
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `test_models.py`, `detect_language`?**
|
||||
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
|
||||
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `models.py`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `InherenceClassifier`, `detect_language`, `classifier.py`?**
|
||||
_High betweenness centrality (0.005) - this node is a cross-community bridge._
|
||||
- **Why does `InherenceClassifier` connect `InherenceClassifier` to `models.py`, `ECPSnapshot`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `detect_language`, `classifier.py`?**
|
||||
_High betweenness centrality (0.003) - this node is a cross-community bridge._
|
||||
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?**
|
||||
_High betweenness centrality (0.003) - this node is a cross-community bridge._
|
||||
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
||||
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
||||
|
||||
+7433
-1423
File diff suppressed because it is too large
Load Diff
@@ -420,9 +420,9 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"requirements.txt": {
|
||||
"mtime": 1787236111.876247,
|
||||
"seen": 1787236262.0105975,
|
||||
"ast_hash": "9fb7e3f8ecb04ff30d650c772f509557",
|
||||
"mtime": 1787240110.13472,
|
||||
"seen": 1787240516.5993943,
|
||||
"ast_hash": "f5d99630fa9c93c5fbaa44814f107e70",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/fixtures/benchmark_24/de/contextual.md": {
|
||||
@@ -652,5 +652,89 @@
|
||||
"seen": 1787235020.6850297,
|
||||
"ast_hash": "627a6c953b250083ebe76ee74d605bb5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"README.md": {
|
||||
"mtime": 1787240502.4633865,
|
||||
"seen": 1787240516.5990582,
|
||||
"ast_hash": "a075f140e2dd8383bec3991f92b32b26",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"scripts/extract_article_contents.py": {
|
||||
"mtime": 1787241061.3097162,
|
||||
"seen": 1787241108.1811028,
|
||||
"ast_hash": "08cd37738da1500b6912ffa1fb2fc0cc",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_extract_article_contents.py": {
|
||||
"mtime": 1787240910.5980105,
|
||||
"seen": 1787241108.1826632,
|
||||
"ast_hash": "98cf3b30fa6ce044d943eff381eaf8a7",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/prd_extrator_artigos_nlp.md": {
|
||||
"mtime": 1787237774.0092583,
|
||||
"seen": 1787240516.5991437,
|
||||
"ast_hash": "c90f25764acefb91f1160feaf497fb33",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/checklists/extraction.md": {
|
||||
"mtime": 1787239873.797626,
|
||||
"seen": 1787240516.6008725,
|
||||
"ast_hash": "3c78fb563727a7f6f23240dcaad3c010",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/checklists/requirements.md": {
|
||||
"mtime": 1787237950.8140705,
|
||||
"seen": 1787240516.600874,
|
||||
"ast_hash": "16b2bf7a7910722a7aa6c57cdcc31f1c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/contracts/cli-contract.md": {
|
||||
"mtime": 1787239569.0560138,
|
||||
"seen": 1787240516.6008754,
|
||||
"ast_hash": "9326b1b3c124ad41636c3fae5c9de8b5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/contracts/json-schema.md": {
|
||||
"mtime": 1787239578.9314032,
|
||||
"seen": 1787240516.6008763,
|
||||
"ast_hash": "93f1da07b16c61de15dc93071434d145",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/data-model.md": {
|
||||
"mtime": 1787239558.2362826,
|
||||
"seen": 1787240516.6008773,
|
||||
"ast_hash": "03ed32de948199c9a498c486e7bdafbd",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/plan.md": {
|
||||
"mtime": 1787239601.395844,
|
||||
"seen": 1787240516.6008782,
|
||||
"ast_hash": "406c3ebad1d32984c1b98f8a9b7ae9f1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/quickstart.md": {
|
||||
"mtime": 1787239589.9896657,
|
||||
"seen": 1787240516.6008794,
|
||||
"ast_hash": "7cac454150d33dad67f683367fca749f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/research.md": {
|
||||
"mtime": 1787239545.8330145,
|
||||
"seen": 1787240516.6008804,
|
||||
"ast_hash": "beef81df32e2dc16798ece67726daa95",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/spec.md": {
|
||||
"mtime": 1787239197.9064422,
|
||||
"seen": 1787240516.6008813,
|
||||
"ast_hash": "9fd196b7ba4d9ba7adb2cad25d094764",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/tasks.md": {
|
||||
"mtime": 1787240464.6520643,
|
||||
"seen": 1787240516.6008823,
|
||||
"ast_hash": "d95c4763728496703cd89590288a2bed",
|
||||
"semantic_hash": ""
|
||||
}
|
||||
}
|
||||
+106
-56
@@ -1,16 +1,16 @@
|
||||
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
||||
|
||||
## Corpus Check
|
||||
- 148 files · ~67,357 words
|
||||
- 161 files · ~78,549 words
|
||||
- Verdict: corpus is large enough that graph structure adds value.
|
||||
|
||||
## Summary
|
||||
- 802 nodes · 957 edges · 104 communities (66 shown, 38 thin omitted)
|
||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
|
||||
- 1000 nodes · 1220 edges · 115 communities (77 shown, 38 thin omitted)
|
||||
- Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 39 edges (avg confidence: 0.95)
|
||||
- Token cost: 0 input · 0 output
|
||||
|
||||
## Graph Freshness
|
||||
- Built from commit: `67cc40f9`
|
||||
- Built from commit: `6e3d5761`
|
||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||
- Run `graphify update .` after code changes (no API cost).
|
||||
|
||||
@@ -56,10 +56,10 @@
|
||||
- 1. Input Schemas
|
||||
- 2. Basic CLI Usage Examples
|
||||
- 2. Standard Streams & Exit Codes
|
||||
- ClassificationResult
|
||||
- classifier.py
|
||||
- InherenceClassifier
|
||||
- detect_language
|
||||
- test_models.py
|
||||
- ClassificationError
|
||||
- content_northvolt_de.md
|
||||
- content_presal_pt.md
|
||||
- content_tangential_es.md
|
||||
@@ -91,7 +91,7 @@
|
||||
- pt/tangential.md
|
||||
- tests/__init__.py
|
||||
- text-nlp-classifier
|
||||
- get_hl_gl_ceid
|
||||
- test_extract_article_contents.py
|
||||
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||
- extract_google_news.py
|
||||
- ExtractionResult
|
||||
@@ -108,41 +108,51 @@
|
||||
- Specification Quality Checklist: Google News Headlines Extractor
|
||||
- CLI Contract: Google News Headlines Extractor
|
||||
- 🧠 TextNLPClassifierApp
|
||||
- build_parser
|
||||
- Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
|
||||
- sample_rss_xml
|
||||
- classifier.py
|
||||
- main
|
||||
- ECPSnapshot
|
||||
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- main
|
||||
- 4. Requisitos Funcionais (FR)
|
||||
- Tasks: Article Content Multi-Engine Extractor
|
||||
- models.py
|
||||
- Implementation Plan: Article Content Multi-Engine Extractor
|
||||
- 2. Cenários de Validação
|
||||
- 1. Technical Decisions & Tradeoffs
|
||||
- Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- CLI Contract: Article Content Multi-Engine Extractor
|
||||
- JSON Schema Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## God Nodes (most connected - your core abstractions)
|
||||
1. `ECPSnapshot` - 31 edges
|
||||
2. `InherenceClassifier` - 25 edges
|
||||
3. `DecisionCategory` - 18 edges
|
||||
3. `DecisionCategory` - 17 edges
|
||||
4. `ClassificationResult` - 17 edges
|
||||
5. `main()` - 14 edges
|
||||
5. `process_batch()` - 15 edges
|
||||
6. `LocalEmbeddingsAdapter` - 14 edges
|
||||
7. `LLMFallbackAdapter` - 14 edges
|
||||
8. `detect_language()` - 14 edges
|
||||
9. `Tasks: [FEATURE NAME]` - 13 edges
|
||||
10. `SearchQuery` - 12 edges
|
||||
9. `main()` - 13 edges
|
||||
10. `ArticleCrawler` - 13 edges
|
||||
|
||||
## Surprising Connections (you probably didn't know these)
|
||||
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `main()` --uses--> `ErrorCode` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
|
||||
tests/test_extract_google_news.py → scripts/extract_google_news.py
|
||||
- `classifier()` --uses--> `InherenceClassifier` [INFERRED]
|
||||
tests/test_benchmark_24.py → src/classifier.py
|
||||
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
|
||||
## Import Cycles
|
||||
- None detected.
|
||||
|
||||
## Communities (104 total, 38 thin omitted)
|
||||
## Communities (115 total, 38 thin omitted)
|
||||
|
||||
### Community 0 - "Task Planning"
|
||||
Cohesion: 0.07
|
||||
@@ -165,8 +175,8 @@ Cohesion: 0.08
|
||||
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
|
||||
|
||||
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.06
|
||||
Nodes (33): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+25 more)
|
||||
Cohesion: 0.15
|
||||
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
|
||||
|
||||
### Community 6 - "Feature Specification Template"
|
||||
Cohesion: 0.15
|
||||
@@ -265,8 +275,8 @@ Cohesion: 0.50
|
||||
Nodes (3): Boundaries, Output, Scan
|
||||
|
||||
### Community 40 - "main"
|
||||
Cohesion: 0.19
|
||||
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
|
||||
Cohesion: 0.14
|
||||
Nodes (17): ArgumentParser, CaptureFixture, build_parser(), main(), Cria e configura o parser de argumentos CLI., Ponto de entrada do script CLI., Path, Valida execução padrão do CLI com saída JSON no stdout. (+9 more)
|
||||
|
||||
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.22
|
||||
@@ -284,41 +294,41 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
|
||||
Cohesion: 0.29
|
||||
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
||||
|
||||
### Community 45 - "ClassificationResult"
|
||||
Cohesion: 0.10
|
||||
Nodes (18): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., LocalEmbeddingsAdapter (+10 more)
|
||||
### Community 45 - "classifier.py"
|
||||
Cohesion: 0.14
|
||||
Nodes (9): LocalEmbeddingsAdapter, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Core deterministic classification engine (Tier 1 core)., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface() (+1 more)
|
||||
|
||||
### Community 46 - "InherenceClassifier"
|
||||
Cohesion: 0.12
|
||||
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more)
|
||||
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess without --output and verify stdout is pure parseable…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC. (+19 more)
|
||||
|
||||
### Community 47 - "detect_language"
|
||||
Cohesion: 0.14
|
||||
Nodes (20): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization. (+12 more)
|
||||
Cohesion: 0.09
|
||||
Nodes (30): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+22 more)
|
||||
|
||||
### Community 48 - "test_models.py"
|
||||
Cohesion: 0.21
|
||||
Nodes (12): Classify inherence of content against an ECP snapshot., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+4 more)
|
||||
### Community 48 - "ClassificationError"
|
||||
Cohesion: 0.29
|
||||
Nodes (3): ClassificationError, Any, test_classification_error_serialization()
|
||||
|
||||
### Community 80 - "get_hl_gl_ceid"
|
||||
Cohesion: 0.25
|
||||
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
|
||||
### Community 80 - "test_extract_article_contents.py"
|
||||
Cohesion: 0.06
|
||||
Nodes (56): ArticleCrawler, extract_all_engines(), ExtractedArticle, ExtractionBatchReport, InputArticle, load_search_json(), log_info(), main() (+48 more)
|
||||
|
||||
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
|
||||
Cohesion: 0.08
|
||||
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
|
||||
|
||||
### Community 82 - "extract_google_news.py"
|
||||
Cohesion: 0.20
|
||||
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
|
||||
Cohesion: 0.15
|
||||
Nodes (18): extract_google_news(), _fetch_rss_content(), get_hl_gl_ceid(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Remove pontuação e espaços extras para comparação de redundância. (+10 more)
|
||||
|
||||
### Community 83 - "ExtractionResult"
|
||||
Cohesion: 0.29
|
||||
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
|
||||
|
||||
### Community 84 - "test_extract_google_news.py"
|
||||
Cohesion: 0.21
|
||||
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
|
||||
Cohesion: 0.15
|
||||
Nodes (15): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado. (+7 more)
|
||||
|
||||
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
|
||||
Cohesion: 0.14
|
||||
@@ -361,47 +371,87 @@ Cohesion: 0.33
|
||||
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
|
||||
|
||||
### Community 97 - "🧠 TextNLPClassifierApp"
|
||||
Cohesion: 0.07
|
||||
Nodes (26): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências, 3. Baixar Binários do Navegador Stealth (Camoufox), 3. Fórmula 1 (Inglaterra / Inglês) (+18 more)
|
||||
Cohesion: 0.06
|
||||
Nodes (33): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. Extração Completa Automática, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Amostragem Rápida (Limit 2 Notícias), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências (+25 more)
|
||||
|
||||
### Community 98 - "build_parser"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
|
||||
### Community 98 - "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.05
|
||||
Nodes (34): 1. Requirement Completeness, 2. Requirement Clarity & Non-Ambiguity, 3. Requirement Consistency & Data Contracts, 4. Scenario & Edge Case Coverage, 5. Non-Functional & Operational Readiness, Extraction Pipeline Checklist: Article Content Multi-Engine Extractor, Notes, Content Quality (+26 more)
|
||||
|
||||
### Community 99 - "sample_rss_xml"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
|
||||
|
||||
### Community 100 - "classifier.py"
|
||||
### Community 100 - "main"
|
||||
Cohesion: 0.23
|
||||
Nodes (9): emit_error(), Enum, Core deterministic classification engine (Tier 1 core)., ClassificationError, ErrorCode, MatchedGraphEntity, Data models and validation schemas for Multilingual NLP Entity Inherence…, str (+1 more)
|
||||
Nodes (13): emit_error(), main(), parse_args(), Namespace, Enum, ErrorCode, str, CLI execution tests covering flags, arguments, stdout, and error handling. (+5 more)
|
||||
|
||||
### Community 101 - "ECPSnapshot"
|
||||
Cohesion: 0.20
|
||||
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more)
|
||||
Cohesion: 0.22
|
||||
Nodes (11): parametrize, ECPSnapshot, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), Unit tests for ECP models, schema validation, and structured error handling., test_classification_result_serialization() (+3 more)
|
||||
|
||||
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.14
|
||||
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
||||
|
||||
### Community 103 - "main"
|
||||
Cohesion: 0.31
|
||||
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
|
||||
### Community 103 - "4. Requisitos Funcionais (FR)"
|
||||
Cohesion: 0.10
|
||||
Nodes (20): 1.1 Objetivo do Produto, 1. Visão Geral e Contexto, 2. Personas e Casos de Uso, 3. Arquitetura e Fluxo do Sistema, 4. Requisitos Funcionais (FR), 5. Requisitos Não Funcionais (NFR), 6.1 Esquema do JSON de Entrada (`Input`), 6.2 Esquema do JSON Consolidado de Saída (`Output`) (+12 more)
|
||||
|
||||
### Community 104 - "Tasks: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.11
|
||||
Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more)
|
||||
|
||||
### Community 105 - "models.py"
|
||||
Cohesion: 0.16
|
||||
Nodes (12): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional local vector embeddings adapter (Tier 2). Disabled by default.… (+4 more)
|
||||
|
||||
### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.17
|
||||
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Article Content Multi-Engine Extractor, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Completed)*, Project Structure (+3 more)
|
||||
|
||||
### Community 107 - "2. Cenários de Validação"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 1. Pré-requisitos, 2. Cenários de Validação, 3. Validação Automatizada de Testes, Cenário 1: Extração com Amostragem Rápida (Limit 2), Cenário 2: Caminho Customizado de Saída, Cenário 3: Modo Silencioso (`--silent`), Cenário 4: Resiliência contra URLs Inválidas, Quickstart & Validation Guide: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 108 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.22
|
||||
Nodes (8): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Navegação e Renderização com `foxcape` em Sessão Única, Decision 2: Orquestração Tripla de Extração de Conteúdo (NLP & Web Scraping), Decision 3: Resiliência e Isolamento de Falhas por Camada, Decision 4: Herança Inteligente de Idioma para NLP, Decision 5: Gerenciamento de Memória e Descarte do Raw HTML, Decision 6: Segregação de Streams e Feedback Visual em `stderr`, Research: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 109 - "Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.25
|
||||
Nodes (8): Complexity Tracking, Constitution Check, Documentation (this feature), Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC), Project Structure, Source Code (repository root), Summary, Technical Context
|
||||
|
||||
### Community 110 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.29
|
||||
Nodes (7): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
|
||||
### Community 111 - "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
|
||||
### Community 112 - "CLI Contract: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): 1. Comando de Execução, 2. Argumentos e Flags, 3. Códigos de Saída (Exit Codes), 4. Comportamento de Streams (I/O), CLI Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
### Community 114 - "JSON Schema Contract: Article Content Multi-Engine Extractor"
|
||||
Cohesion: 0.50
|
||||
Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## Knowledge Gaps
|
||||
- **381 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+376 more)
|
||||
- **465 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+460 more)
|
||||
These have ≤1 connection - possible missing edges or undocumented components.
|
||||
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
|
||||
## Suggested Questions
|
||||
_Questions this graph is uniquely positioned to answer:_
|
||||
|
||||
- **Why does `main()` connect `main` to `main`, `classifier.py`, `ECPSnapshot`, `InherenceClassifier`?**
|
||||
_High betweenness centrality (0.029) - this node is a cross-community bridge._
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `classifier.py`, `main`, `ClassificationResult`, `InherenceClassifier`, `test_models.py`?**
|
||||
_High betweenness centrality (0.020) - this node is a cross-community bridge._
|
||||
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
|
||||
_High betweenness centrality (0.020) - this node is a cross-community bridge._
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `models.py`, `classifier.py`, `InherenceClassifier`, `detect_language`?**
|
||||
_High betweenness centrality (0.006) - this node is a cross-community bridge._
|
||||
- **Why does `InherenceClassifier` connect `InherenceClassifier` to `main`, `ECPSnapshot`, `models.py`, `classifier.py`, `detect_language`?**
|
||||
_High betweenness centrality (0.003) - this node is a cross-community bridge._
|
||||
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?**
|
||||
_High betweenness centrality (0.003) - this node is a cross-community bridge._
|
||||
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
||||
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
||||
|
||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
||||
{"nodes": [{"id": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_md", "label": "json-schema.md", "file_type": "document", "node_kind": "page", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_json_schema_contract_article_content_multi_engine_extractor", "label": "JSON Schema Contract: Article Content Multi-Engine Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_1_schema_de_entrada_input_json", "label": "1. Schema de Entrada (Input JSON)", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L3"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_2_schema_de_sa\u00edda_output_json", "label": "2. Schema de Sa\u00edda (Output JSON)", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L27"}], "edges": [{"source": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_md", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_json_schema_contract_article_content_multi_engine_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_json_schema_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_1_schema_de_entrada_input_json", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L3", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_json_schema_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_json_schema_2_schema_de_sa\u00edda_output_json", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/json-schema.md", "source_location": "L27", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
||||
{"nodes": [{"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_md", "label": "requirements.md", "file_type": "document", "node_kind": "page", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "label": "Specification Quality Checklist: Article Content Multi-Engine Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_content_quality", "label": "Content Quality", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_requirement_completeness", "label": "Requirement Completeness", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L14"}, {"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_feature_readiness", "label": "Feature Readiness", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L25"}, {"id": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_notes", "label": "Notes", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L32"}], "edges": [{"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_md", "target": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_md", "target": "$graphify-root$_specs_003_article_content_extractor_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L5", "weight": 1.0, "target_file": "$graphify-root$/specs/003-article-content-extractor/spec.md"}, {"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_content_quality", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_requirement_completeness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L14", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_feature_readiness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L25", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_specification_quality_checklist_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_checklists_requirements_notes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/checklists/requirements.md", "source_location": "L32", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
||||
{"nodes": [{"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_md", "label": "cli-contract.md", "file_type": "document", "node_kind": "page", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "label": "CLI Contract: Article Content Multi-Engine Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_1_comando_de_execu\u00e7\u00e3o", "label": "1. Comando de Execu\u00e7\u00e3o", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L3"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_2_argumentos_e_flags", "label": "2. Argumentos e Flags", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L11"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_3_c\u00f3digos_de_sa\u00edda_exit_codes", "label": "3. C\u00f3digos de Sa\u00edda (Exit Codes)", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L25"}, {"id": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_4_comportamento_de_streams_i_o", "label": "4. Comportamento de Streams (I/O)", "file_type": "document", "node_kind": "heading", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L35"}], "edges": [{"source": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_md", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_1_comando_de_execu\u00e7\u00e3o", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L3", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_2_argumentos_e_flags", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L11", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_3_c\u00f3digos_de_sa\u00edda_exit_codes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L25", "weight": 1.0}, {"source": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_cli_contract_article_content_multi_engine_extractor", "target": "$graphify-root$_specs_003_article_content_extractor_contracts_cli_contract_4_comportamento_de_streams_i_o", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/003-article-content-extractor/contracts/cli-contract.md", "source_location": "L35", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
||||
{"nodes": [{"id": "pkg_text_nlp_classifier", "label": "text-nlp-classifier", "file_type": "code", "type": "package", "ecosystem": "python", "source_file": "pyproject.toml", "source_location": "L1", "version": "0.1.0"}], "edges": []}
|
||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
Vendored
+1
-1
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+7339
-2011
File diff suppressed because it is too large
Load Diff
+141
-63
@@ -294,15 +294,15 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"classify.py": {
|
||||
"mtime": 1787197110.3753626,
|
||||
"seen": 1787197159.9828603,
|
||||
"ast_hash": "d09a35a5e25d42f6ecc537d5b67bef8e",
|
||||
"mtime": 1787264412.2457643,
|
||||
"seen": 1787264519.0539427,
|
||||
"ast_hash": "e89fc4b64bd7606fc466d5338108705b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"pyproject.toml": {
|
||||
"mtime": 1787234982.768605,
|
||||
"seen": 1787235020.6850271,
|
||||
"ast_hash": "26f5f979658067bf447f375f044a8f21",
|
||||
"mtime": 1787264482.8509245,
|
||||
"seen": 1787264519.0539448,
|
||||
"ast_hash": "409c1fc0d7bbda6830ccb71c6196cb2f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/__init__.py": {
|
||||
@@ -318,45 +318,45 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/adapters/base.py": {
|
||||
"mtime": 1787196396.6199949,
|
||||
"seen": 1787196463.0953116,
|
||||
"ast_hash": "40dc6e175c8708467748c1f42a0e064a",
|
||||
"mtime": 1787264412.2437606,
|
||||
"seen": 1787264519.0542011,
|
||||
"ast_hash": "220635d5d74e73f58259069bf5207908",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/adapters/embeddings.py": {
|
||||
"mtime": 1787196402.7484262,
|
||||
"seen": 1787196463.0953145,
|
||||
"ast_hash": "0444823bb05720d4c4ca1662a00b2c01",
|
||||
"mtime": 1787264437.92723,
|
||||
"seen": 1787264519.0542023,
|
||||
"ast_hash": "7beb0ecfb7482c1fd10422da99f69abe",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/adapters/llm.py": {
|
||||
"mtime": 1787196408.180451,
|
||||
"seen": 1787196463.0953166,
|
||||
"ast_hash": "3d95d98625df3fcd343f2fecb78707c5",
|
||||
"mtime": 1787264412.2437606,
|
||||
"seen": 1787264519.0542035,
|
||||
"ast_hash": "ba2990328f5b25ac6cdfc3b57acc9d39",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/classifier.py": {
|
||||
"mtime": 1787197085.0250447,
|
||||
"seen": 1787197159.985785,
|
||||
"ast_hash": "a2e70f968109d213fdab0ae81ac819ff",
|
||||
"mtime": 1787264459.052089,
|
||||
"seen": 1787264519.0542045,
|
||||
"ast_hash": "b8fb440374ecd66bb8bf67c70c19ed1f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/language.py": {
|
||||
"mtime": 1787196384.5545745,
|
||||
"seen": 1787196463.0953217,
|
||||
"ast_hash": "985619013e58a7f55af2e955817f223c",
|
||||
"mtime": 1787264437.9312274,
|
||||
"seen": 1787264519.0542057,
|
||||
"ast_hash": "cbaf52272ebb05372e34a05cd201c270",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/models.py": {
|
||||
"mtime": 1787195855.3266282,
|
||||
"seen": 1787196463.0953236,
|
||||
"ast_hash": "ce73bfe85eb54f907e2052c80636ecaf",
|
||||
"mtime": 1787264437.9302285,
|
||||
"seen": 1787264519.054207,
|
||||
"ast_hash": "16e604c14f7d66634dc9e2ed7959dd7c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/parser.py": {
|
||||
"mtime": 1787195873.9119804,
|
||||
"seen": 1787196463.0953257,
|
||||
"ast_hash": "0a1d64ee7088b266125af071f85f3826",
|
||||
"mtime": 1787264437.9312274,
|
||||
"seen": 1787264519.054208,
|
||||
"ast_hash": "d43fca3f063534e3626f86f36a97a2cc",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/__init__.py": {
|
||||
@@ -366,39 +366,39 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_adapters.py": {
|
||||
"mtime": 1787196418.6328156,
|
||||
"seen": 1787196463.09533,
|
||||
"ast_hash": "5caf95327ea030dbcc37a99bd9052ebd",
|
||||
"mtime": 1787264437.9292288,
|
||||
"seen": 1787264519.0543096,
|
||||
"ast_hash": "7ed9ba6f62fef404bfc28a2e16ada647",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_benchmark_24.py": {
|
||||
"mtime": 1787196325.5409796,
|
||||
"seen": 1787196463.0953324,
|
||||
"ast_hash": "487d30d6de2dee529fa91b86a1d62c0a",
|
||||
"mtime": 1787264437.9302285,
|
||||
"seen": 1787264519.0543122,
|
||||
"ast_hash": "052a51561002be035de1a79455256b67",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_classifier.py": {
|
||||
"mtime": 1787196364.7195728,
|
||||
"seen": 1787196463.0953343,
|
||||
"ast_hash": "4350a5b11a08d12ccb85a3f332d71e87",
|
||||
"mtime": 1787264437.92723,
|
||||
"seen": 1787264519.0543132,
|
||||
"ast_hash": "14da0c7a7d4c08762437ca777c086dea",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_cli.py": {
|
||||
"mtime": 1787195939.567137,
|
||||
"seen": 1787196463.0953364,
|
||||
"ast_hash": "90b0131d43cfc64a76e499c7a65d95de",
|
||||
"mtime": 1787264437.9292288,
|
||||
"seen": 1787264519.0543146,
|
||||
"ast_hash": "b6f57994937363f8a1b0f2f66ea53d3a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_language.py": {
|
||||
"mtime": 1787195891.3329668,
|
||||
"seen": 1787196463.0953388,
|
||||
"ast_hash": "66191b43fcd7aef47c07cd20d5a8f03c",
|
||||
"mtime": 1787264437.9292288,
|
||||
"seen": 1787264519.0543184,
|
||||
"ast_hash": "f1a09702410864434274f1b347769742",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_models.py": {
|
||||
"mtime": 1787195886.5009987,
|
||||
"seen": 1787196463.095341,
|
||||
"ast_hash": "1fd725c8ca0f52e52f56c7223e9574d0",
|
||||
"mtime": 1787264437.9302285,
|
||||
"seen": 1787264519.0543194,
|
||||
"ast_hash": "7ec72b612bf758be8c9a3617ce18b132",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"examples/content_northvolt_de.md": {
|
||||
@@ -420,9 +420,9 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"requirements.txt": {
|
||||
"mtime": 1787236111.876247,
|
||||
"seen": 1787236262.0105975,
|
||||
"ast_hash": "9fb7e3f8ecb04ff30d650c772f509557",
|
||||
"mtime": 1787240110.13472,
|
||||
"seen": 1787240516.5993943,
|
||||
"ast_hash": "f5d99630fa9c93c5fbaa44814f107e70",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/fixtures/benchmark_24/de/contextual.md": {
|
||||
@@ -570,27 +570,27 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_adversarial.py": {
|
||||
"mtime": 1787197141.2340336,
|
||||
"seen": 1787197159.9881344,
|
||||
"ast_hash": "5e8daf517ee32c261617d96264ef0473",
|
||||
"mtime": 1787264437.9242287,
|
||||
"seen": 1787264519.0543108,
|
||||
"ast_hash": "d353633ce31a624a8dcda790efec4d5e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"scripts/extract_google_news.py": {
|
||||
"mtime": 1787236453.5231855,
|
||||
"seen": 1787236509.1147037,
|
||||
"ast_hash": "4216de490a87378d21dab707a3166679",
|
||||
"mtime": 1787264437.92723,
|
||||
"seen": 1787264519.0540311,
|
||||
"ast_hash": "97d8a193901d5e2edf6359fad256b7ef",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_extract_google_news.py": {
|
||||
"mtime": 1787236301.513448,
|
||||
"seen": 1787236353.5582018,
|
||||
"ast_hash": "1fc9a108a8abfecf1d37c8642a885e9e",
|
||||
"mtime": 1787264437.9312274,
|
||||
"seen": 1787264519.054317,
|
||||
"ast_hash": "7561ad8fedcf1bc6cc8ed1e86a345532",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/googlenews_extractor_guia_completo.md": {
|
||||
"mtime": 1787231605.0830677,
|
||||
"seen": 1787234772.8112028,
|
||||
"ast_hash": "59db8e652000ba6087a00289a0ce0959",
|
||||
"mtime": 1787264437.9282286,
|
||||
"seen": 1787264519.0578,
|
||||
"ast_hash": "d7c3c64e2009dded496a08c7c91f63b9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/checklists/readiness.md": {
|
||||
@@ -654,9 +654,87 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"README.md": {
|
||||
"mtime": 1787237002.3288481,
|
||||
"seen": 1787237013.7487457,
|
||||
"ast_hash": "1b624e85d8ce1521683cbea2664dd2d2",
|
||||
"mtime": 1787264334.4498444,
|
||||
"seen": 1787264519.0577974,
|
||||
"ast_hash": "196db110cfde06179fb750fc75afb6fe",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"scripts/extract_article_contents.py": {
|
||||
"mtime": 1787264465.3415215,
|
||||
"seen": 1787264519.0540295,
|
||||
"ast_hash": "2bea6b2a048526cb49d95fb88a001600",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_extract_article_contents.py": {
|
||||
"mtime": 1787264437.9292288,
|
||||
"seen": 1787264519.0543158,
|
||||
"ast_hash": "56aee72480f05c714865791a935d539c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/prd_extrator_artigos_nlp.md": {
|
||||
"mtime": 1787237774.0092583,
|
||||
"seen": 1787240516.5991437,
|
||||
"ast_hash": "c90f25764acefb91f1160feaf497fb33",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/checklists/extraction.md": {
|
||||
"mtime": 1787239873.797626,
|
||||
"seen": 1787240516.6008725,
|
||||
"ast_hash": "3c78fb563727a7f6f23240dcaad3c010",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/checklists/requirements.md": {
|
||||
"mtime": 1787237950.8140705,
|
||||
"seen": 1787240516.600874,
|
||||
"ast_hash": "16b2bf7a7910722a7aa6c57cdcc31f1c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/contracts/cli-contract.md": {
|
||||
"mtime": 1787239569.0560138,
|
||||
"seen": 1787240516.6008754,
|
||||
"ast_hash": "9326b1b3c124ad41636c3fae5c9de8b5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/contracts/json-schema.md": {
|
||||
"mtime": 1787239578.9314032,
|
||||
"seen": 1787240516.6008763,
|
||||
"ast_hash": "93f1da07b16c61de15dc93071434d145",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/data-model.md": {
|
||||
"mtime": 1787239558.2362826,
|
||||
"seen": 1787240516.6008773,
|
||||
"ast_hash": "03ed32de948199c9a498c486e7bdafbd",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/plan.md": {
|
||||
"mtime": 1787264323.707266,
|
||||
"seen": 1787264519.0605621,
|
||||
"ast_hash": "b2232491829a303a49e749c75f3adca2",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/quickstart.md": {
|
||||
"mtime": 1787239589.9896657,
|
||||
"seen": 1787240516.6008794,
|
||||
"ast_hash": "7cac454150d33dad67f683367fca749f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/research.md": {
|
||||
"mtime": 1787239545.8330145,
|
||||
"seen": 1787240516.6008804,
|
||||
"ast_hash": "beef81df32e2dc16798ece67726daa95",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/spec.md": {
|
||||
"mtime": 1787264319.2883234,
|
||||
"seen": 1787264519.060745,
|
||||
"ast_hash": "187b1307e28a02f0ac44895b5f9cca55",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/003-article-content-extractor/tasks.md": {
|
||||
"mtime": 1787240464.6520643,
|
||||
"seen": 1787240516.6008823,
|
||||
"ast_hash": "d95c4763728496703cd89590288a2bed",
|
||||
"semantic_hash": ""
|
||||
}
|
||||
}
|
||||
@@ -24,3 +24,17 @@ python_functions = ["test_*"]
|
||||
|
||||
[tool.pyright]
|
||||
extraPaths = [".", "src"]
|
||||
|
||||
[tool.ruff]
|
||||
line-length = 100
|
||||
target-version = "py310"
|
||||
|
||||
[tool.ruff.lint]
|
||||
select = ["E", "F", "I", "W"]
|
||||
ignore = ["E501"]
|
||||
|
||||
[tool.mypy]
|
||||
python_version = "3.12"
|
||||
ignore_missing_imports = true
|
||||
check_untyped_defs = true
|
||||
|
||||
|
||||
@@ -3,6 +3,10 @@ foxcape>=0.1.1
|
||||
beautifulsoup4>=4.12.0
|
||||
googlenewsdecoder>=0.1.7
|
||||
selectolax>=0.3.27
|
||||
trafilatura>=1.8.0
|
||||
newspaper4k>=0.9.3.1
|
||||
readability-lxml>=0.8.1
|
||||
lxml>=4.9.0
|
||||
# Optional Tier 2 / Tier 3 dependencies (not required for POC core execution)
|
||||
# sentence-transformers>=2.2.0
|
||||
# httpx>=0.24.0
|
||||
|
||||
@@ -0,0 +1,770 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Extrator e Parser de Artigos Multimotor (Foxcape + Trafilatura + Newspaper4k + Readability).
|
||||
|
||||
Lê listagens JSON de notícias (ex: out/river_plate.json), acessa e renderiza as páginas
|
||||
em modo stealth headless utilizando Foxcape reutilizando a mesma sessão de navegador,
|
||||
executa a extração em paralelo/sequência com 3 motores de conteúdo (Trafilatura,
|
||||
Newspaper4k e Readability) e salva o resultado enriquecido e higienizado em JSON.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any, Literal
|
||||
|
||||
import trafilatura
|
||||
from bs4 import BeautifulSoup
|
||||
from foxcape import Foxcape, FoxcapeConfig
|
||||
from newspaper import Article
|
||||
from readability import Document
|
||||
|
||||
# ==============================================================================
|
||||
# Modelos de Dados e Dataclasses
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class InputArticle:
|
||||
"""Metadados originais da notícia contida no JSON de entrada."""
|
||||
|
||||
titulo: str
|
||||
url: str
|
||||
subtitulo: str | None = None
|
||||
quando_publicado: str | None = None
|
||||
pagina: int = 1
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"titulo": self.titulo,
|
||||
"subtitulo": self.subtitulo,
|
||||
"quando_publicado": self.quando_publicado,
|
||||
"url": self.url,
|
||||
"pagina": self.pagina,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TrafilaturaData:
|
||||
"""Dados completos extraídos pelo motor Trafilatura."""
|
||||
|
||||
title: str | None = None
|
||||
author: str | None = None
|
||||
date: str | None = None
|
||||
description: str | None = None
|
||||
sitename: str | None = None
|
||||
hostname: str | None = None
|
||||
language: str | None = None
|
||||
categories: list[str] = field(default_factory=list)
|
||||
tags: list[str] = field(default_factory=list)
|
||||
canonical_url: str | None = None
|
||||
image: str | None = None
|
||||
pagetype: str | None = None
|
||||
fingerprint: str | None = None
|
||||
license: str | None = None
|
||||
comments: str | None = None
|
||||
text: str = ""
|
||||
markdown: str | None = None
|
||||
raw_json: dict[str, Any] | None = None
|
||||
error: str | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"title": self.title,
|
||||
"author": self.author,
|
||||
"date": self.date,
|
||||
"description": self.description,
|
||||
"sitename": self.sitename,
|
||||
"hostname": self.hostname,
|
||||
"language": self.language,
|
||||
"categories": self.categories,
|
||||
"tags": self.tags,
|
||||
"canonical_url": self.canonical_url,
|
||||
"image": self.image,
|
||||
"pagetype": self.pagetype,
|
||||
"fingerprint": self.fingerprint,
|
||||
"license": self.license,
|
||||
"comments": self.comments,
|
||||
"text": self.text,
|
||||
"markdown": self.markdown,
|
||||
"raw_json": self.raw_json,
|
||||
"error": self.error,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class NewspaperData:
|
||||
"""Dados completos extraídos e enriquecidos com NLP pelo motor Newspaper4k."""
|
||||
|
||||
title: str | None = None
|
||||
authors: list[str] = field(default_factory=list)
|
||||
publish_date: str | None = None
|
||||
text: str = ""
|
||||
summary: str | None = None
|
||||
keywords: list[str] = field(default_factory=list)
|
||||
keyword_scores: dict[str, float] = field(default_factory=dict)
|
||||
top_image: str | None = None
|
||||
images: list[str] = field(default_factory=list)
|
||||
movies: list[str] = field(default_factory=list)
|
||||
tags: list[str] = field(default_factory=list)
|
||||
canonical_link: str | None = None
|
||||
article_html: str | None = None
|
||||
meta_description: str | None = None
|
||||
meta_keywords: list[str] = field(default_factory=list)
|
||||
meta_favicon: str | None = None
|
||||
meta_site_name: str | None = None
|
||||
meta_lang: str | None = None
|
||||
meta_data: dict[str, Any] = field(default_factory=dict)
|
||||
error: str | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"title": self.title,
|
||||
"authors": self.authors,
|
||||
"publish_date": self.publish_date,
|
||||
"text": self.text,
|
||||
"summary": self.summary,
|
||||
"keywords": self.keywords,
|
||||
"keyword_scores": self.keyword_scores,
|
||||
"top_image": self.top_image,
|
||||
"images": self.images,
|
||||
"movies": self.movies,
|
||||
"tags": self.tags,
|
||||
"canonical_link": self.canonical_link,
|
||||
"article_html": self.article_html,
|
||||
"meta_description": self.meta_description,
|
||||
"meta_keywords": self.meta_keywords,
|
||||
"meta_favicon": self.meta_favicon,
|
||||
"meta_site_name": self.meta_site_name,
|
||||
"meta_lang": self.meta_lang,
|
||||
"meta_data": self.meta_data,
|
||||
"error": self.error,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ReadabilityData:
|
||||
"""Dados completos higienizados pelo algoritmo Readability."""
|
||||
|
||||
title: str | None = None
|
||||
short_title: str | None = None
|
||||
author: str | None = None
|
||||
cleaned_html: str | None = None
|
||||
cleaned_text: str | None = None
|
||||
error: str | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"title": self.title,
|
||||
"short_title": self.short_title,
|
||||
"author": self.author,
|
||||
"cleaned_html": self.cleaned_html,
|
||||
"cleaned_text": self.cleaned_text,
|
||||
"error": self.error,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExtractedArticle:
|
||||
"""Resultado consolidado da extração de um artigo."""
|
||||
|
||||
input_meta: InputArticle
|
||||
extraction_status: Literal["success", "failed"]
|
||||
error_message: str | None
|
||||
crawled_url: str
|
||||
page_title: str | None
|
||||
http_status: int | None
|
||||
trafilatura: TrafilaturaData | None = None
|
||||
newspaper4k: NewspaperData | None = None
|
||||
readability: ReadabilityData | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"input_meta": self.input_meta.to_dict(),
|
||||
"extraction_status": self.extraction_status,
|
||||
"error_message": self.error_message,
|
||||
"crawled_url": self.crawled_url,
|
||||
"page_title": self.page_title,
|
||||
"http_status": self.http_status,
|
||||
"trafilatura": self.trafilatura.to_dict() if self.trafilatura else None,
|
||||
"newspaper4k": self.newspaper4k.to_dict() if self.newspaper4k else None,
|
||||
"readability": self.readability.to_dict() if self.readability else None,
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExtractionBatchReport:
|
||||
"""Relatório consolidado de saída do processamento de um lote."""
|
||||
|
||||
source_file: str
|
||||
processed_at: str
|
||||
total_articles: int
|
||||
successful_articles: int
|
||||
failed_articles: int
|
||||
articles: list[ExtractedArticle] = field(default_factory=list)
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"source_file": self.source_file,
|
||||
"processed_at": self.processed_at,
|
||||
"total_articles": self.total_articles,
|
||||
"successful_articles": self.successful_articles,
|
||||
"failed_articles": self.failed_articles,
|
||||
"articles": [a.to_dict() for a in self.articles],
|
||||
}
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Parsers / Extratores Especializados
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
class TrafilaturaExtractor:
|
||||
"""Motor de extração baseado na biblioteca Trafilatura."""
|
||||
|
||||
@staticmethod
|
||||
def extract(html: str, url: str | None = None) -> TrafilaturaData:
|
||||
try:
|
||||
# Extração bare document completa
|
||||
doc = trafilatura.bare_extraction(
|
||||
html,
|
||||
url=url,
|
||||
include_comments=True,
|
||||
include_tables=True,
|
||||
include_images=True,
|
||||
include_links=True,
|
||||
include_formatting=True,
|
||||
with_metadata=True,
|
||||
)
|
||||
|
||||
# Extração em markdown
|
||||
markdown_text = trafilatura.extract(
|
||||
html,
|
||||
output_format="markdown",
|
||||
include_comments=True,
|
||||
include_tables=True,
|
||||
include_images=True,
|
||||
include_links=True,
|
||||
include_formatting=True,
|
||||
url=url,
|
||||
)
|
||||
|
||||
# Extração em JSON nativo
|
||||
json_output_str = trafilatura.extract(
|
||||
html,
|
||||
output_format="json",
|
||||
include_comments=True,
|
||||
include_tables=True,
|
||||
include_images=True,
|
||||
include_links=True,
|
||||
url=url,
|
||||
)
|
||||
raw_json = json.loads(json_output_str) if json_output_str else None
|
||||
|
||||
if doc:
|
||||
if isinstance(doc, dict):
|
||||
categories = list(doc.get("categories", [])) if doc.get("categories") else []
|
||||
tags = list(doc.get("tags", [])) if doc.get("tags") else []
|
||||
return TrafilaturaData(
|
||||
title=doc.get("title"),
|
||||
author=doc.get("author"),
|
||||
date=doc.get("date"),
|
||||
description=doc.get("description"),
|
||||
sitename=doc.get("sitename"),
|
||||
hostname=doc.get("hostname"),
|
||||
language=doc.get("language"),
|
||||
categories=categories,
|
||||
tags=tags,
|
||||
canonical_url=doc.get("url") or url,
|
||||
image=doc.get("image"),
|
||||
pagetype=doc.get("pagetype"),
|
||||
fingerprint=doc.get("fingerprint"),
|
||||
license=doc.get("license"),
|
||||
comments=doc.get("comments"),
|
||||
text=(doc.get("text") or "").strip(),
|
||||
markdown=(markdown_text or "").strip() if markdown_text else None,
|
||||
raw_json=raw_json,
|
||||
error=None,
|
||||
)
|
||||
else:
|
||||
categories = list(doc.categories) if doc.categories else []
|
||||
tags = list(doc.tags) if doc.tags else []
|
||||
return TrafilaturaData(
|
||||
title=doc.title,
|
||||
author=doc.author,
|
||||
date=doc.date,
|
||||
description=doc.description,
|
||||
sitename=doc.sitename,
|
||||
hostname=doc.hostname,
|
||||
language=doc.language,
|
||||
categories=categories,
|
||||
tags=tags,
|
||||
canonical_url=doc.url or url,
|
||||
image=doc.image,
|
||||
pagetype=doc.pagetype,
|
||||
fingerprint=doc.fingerprint,
|
||||
license=doc.license,
|
||||
comments=doc.comments,
|
||||
text=(doc.text or "").strip(),
|
||||
markdown=(markdown_text or "").strip() if markdown_text else None,
|
||||
raw_json=raw_json,
|
||||
error=None,
|
||||
)
|
||||
else:
|
||||
raw_text = trafilatura.extract(html, output_format="txt", url=url) or ""
|
||||
return TrafilaturaData(
|
||||
title=raw_json.get("title") if raw_json else None,
|
||||
text=raw_text.strip(),
|
||||
markdown=markdown_text.strip() if markdown_text else None,
|
||||
raw_json=raw_json,
|
||||
error=None,
|
||||
)
|
||||
except Exception as e:
|
||||
return TrafilaturaData(error=str(e))
|
||||
|
||||
|
||||
class NewspaperExtractor:
|
||||
"""Motor de extração baseado no Newspaper4k com NLP."""
|
||||
|
||||
@staticmethod
|
||||
def extract(html: str, url: str = "", language: str = "en") -> NewspaperData:
|
||||
try:
|
||||
lang_code = language.split("-")[0].lower() if language else "en"
|
||||
|
||||
article = Article(url=url, language=lang_code)
|
||||
article.download(input_html=html)
|
||||
article.parse()
|
||||
|
||||
# Executar NLP para summary e keywords com fallback gracioso
|
||||
try:
|
||||
article.nlp()
|
||||
summary = article.summary
|
||||
keywords = list(article.keywords) if article.keywords else []
|
||||
keyword_scores = getattr(article, "keyword_scores", {}) or {}
|
||||
except Exception:
|
||||
summary = None
|
||||
keywords = []
|
||||
keyword_scores = {}
|
||||
|
||||
publish_date_str = (
|
||||
article.publish_date.isoformat()
|
||||
if article.publish_date and hasattr(article.publish_date, "isoformat")
|
||||
else str(article.publish_date)
|
||||
if article.publish_date
|
||||
else None
|
||||
)
|
||||
|
||||
# Metadados e tags adicionais
|
||||
meta_keywords = list(article.meta_keywords) if article.meta_keywords else []
|
||||
tags = list(article.tags) if getattr(article, "tags", None) else []
|
||||
movies = list(article.movies) if getattr(article, "movies", None) else []
|
||||
images = list(article.images) if article.images else []
|
||||
|
||||
return NewspaperData(
|
||||
title=article.title or None,
|
||||
authors=list(article.authors) if article.authors else [],
|
||||
publish_date=publish_date_str,
|
||||
text=article.text or "",
|
||||
summary=summary,
|
||||
keywords=keywords,
|
||||
keyword_scores=dict(keyword_scores),
|
||||
top_image=article.top_image or getattr(article, "meta_img", None) or None,
|
||||
images=images,
|
||||
movies=movies,
|
||||
tags=tags,
|
||||
canonical_link=getattr(article, "canonical_link", None) or None,
|
||||
article_html=getattr(article, "article_html", None) or None,
|
||||
meta_description=getattr(article, "meta_description", None) or None,
|
||||
meta_keywords=meta_keywords,
|
||||
meta_favicon=getattr(article, "meta_favicon", None) or None,
|
||||
meta_site_name=getattr(article, "meta_site_name", None) or None,
|
||||
meta_lang=getattr(article, "meta_lang", None) or None,
|
||||
meta_data=dict(article.meta_data) if article.meta_data else {},
|
||||
error=None,
|
||||
)
|
||||
except Exception as e:
|
||||
return NewspaperData(error=str(e))
|
||||
|
||||
|
||||
class ReadabilityExtractor:
|
||||
"""Motor de extração baseado no algoritmo Readability (readability-lxml)."""
|
||||
|
||||
@staticmethod
|
||||
def extract(html: str) -> ReadabilityData:
|
||||
try:
|
||||
doc = Document(html)
|
||||
title = doc.title()
|
||||
short_title = doc.short_title()
|
||||
cleaned_html = doc.summary()
|
||||
author = None
|
||||
try:
|
||||
author = doc.author()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# Extração de texto limpo a partir do HTML higienizado
|
||||
soup = BeautifulSoup(cleaned_html, "html.parser")
|
||||
cleaned_text = soup.get_text(separator="\n\n", strip=True)
|
||||
|
||||
return ReadabilityData(
|
||||
title=title or None,
|
||||
short_title=short_title or None,
|
||||
author=author or None,
|
||||
cleaned_html=cleaned_html or None,
|
||||
cleaned_text=cleaned_text or None,
|
||||
error=None,
|
||||
)
|
||||
except Exception as e:
|
||||
return ReadabilityData(error=str(e))
|
||||
|
||||
|
||||
def extract_all_engines(
|
||||
html: str, url: str = "", language: str = "en"
|
||||
) -> tuple[TrafilaturaData, NewspaperData, ReadabilityData]:
|
||||
"""Executa a extração simultânea pelos três motores de conteúdo com isolamento defensivo."""
|
||||
try:
|
||||
traf_data = TrafilaturaExtractor.extract(html, url=url)
|
||||
except Exception as exc:
|
||||
traf_data = TrafilaturaData(
|
||||
title=None,
|
||||
author=None,
|
||||
date=None,
|
||||
description=None,
|
||||
categories=[],
|
||||
tags=[],
|
||||
canonical_url=None,
|
||||
text="",
|
||||
raw_json=None,
|
||||
error=str(exc),
|
||||
)
|
||||
|
||||
try:
|
||||
newspaper_data = NewspaperExtractor.extract(html, url=url, language=language)
|
||||
except Exception as exc:
|
||||
newspaper_data = NewspaperData(
|
||||
title=None,
|
||||
authors=[],
|
||||
publish_date=None,
|
||||
text="",
|
||||
summary=None,
|
||||
keywords=[],
|
||||
top_image=None,
|
||||
images=[],
|
||||
meta_data={},
|
||||
error=str(exc),
|
||||
)
|
||||
|
||||
try:
|
||||
readability_data = ReadabilityExtractor.extract(html)
|
||||
except Exception as exc:
|
||||
readability_data = ReadabilityData(
|
||||
title=None,
|
||||
short_title=None,
|
||||
cleaned_html=None,
|
||||
cleaned_text=None,
|
||||
error=str(exc),
|
||||
)
|
||||
|
||||
return traf_data, newspaper_data, readability_data
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Motor de Navegação Stealth Headless com Foxcape
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
class ArticleCrawler:
|
||||
"""Gerenciador de ciclo de vida e requisições via Foxcape Headless."""
|
||||
|
||||
def __init__(self, timeout_sec: int = 30) -> None:
|
||||
self.timeout_sec = timeout_sec
|
||||
self.timeout_ms = timeout_sec * 1000
|
||||
self._scraper: Foxcape | None = None
|
||||
|
||||
def start(self) -> None:
|
||||
if self._scraper is None:
|
||||
config = FoxcapeConfig(headless=True, humanize=False)
|
||||
self._scraper = Foxcape(config=config)
|
||||
self._scraper.start()
|
||||
|
||||
def close(self) -> None:
|
||||
if self._scraper is not None:
|
||||
try:
|
||||
self._scraper.close()
|
||||
except Exception:
|
||||
pass
|
||||
self._scraper = None
|
||||
|
||||
def __enter__(self) -> ArticleCrawler:
|
||||
self.start()
|
||||
return self
|
||||
|
||||
def __exit__(self, exc_type: Any, exc_val: Any, exc_tb: Any) -> None:
|
||||
self.close()
|
||||
|
||||
def crawl(self, url: str) -> tuple[str, str | None, int | None]:
|
||||
"""
|
||||
Navega até a URL, aguarda o carregamento do DOM e retorna (html, page_title, http_status).
|
||||
"""
|
||||
if self._scraper is None:
|
||||
self.start()
|
||||
|
||||
assert self._scraper is not None
|
||||
result = self._scraper.get(
|
||||
url,
|
||||
wait_until="domcontentloaded",
|
||||
timeout_ms=self.timeout_ms,
|
||||
human_delay=False,
|
||||
)
|
||||
return result.html, result.title, result.status_code
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Helpers de I/O e Orquestrador de Lote
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def log_info(message: str, silent: bool = False) -> None:
|
||||
"""Escreve mensagem informativa no stderr."""
|
||||
if not silent:
|
||||
sys.stderr.write(f"[INFO] {message}\n")
|
||||
sys.stderr.flush()
|
||||
|
||||
|
||||
def load_search_json(file_path: Path) -> tuple[str | None, str, list[InputArticle]]:
|
||||
"""Carrega o arquivo JSON gerado pelo extrator de notícias."""
|
||||
if not file_path.exists():
|
||||
raise FileNotFoundError(f"Arquivo de entrada não encontrado: {file_path}")
|
||||
|
||||
with file_path.open("r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
|
||||
query = data.get("query")
|
||||
language = data.get("language", "en")
|
||||
raw_items = data.get("items", [])
|
||||
|
||||
articles = []
|
||||
for item in raw_items:
|
||||
if isinstance(item, dict) and "url" in item and "titulo" in item:
|
||||
articles.append(
|
||||
InputArticle(
|
||||
titulo=item["titulo"],
|
||||
url=item["url"],
|
||||
subtitulo=item.get("subtitulo"),
|
||||
quando_publicado=item.get("quando_publicado"),
|
||||
pagina=item.get("pagina", 1),
|
||||
)
|
||||
)
|
||||
|
||||
return query, language, articles
|
||||
|
||||
|
||||
def save_extracted_json(report: ExtractionBatchReport, output_path: Path) -> None:
|
||||
"""Salva o relatório consolidado em formato JSON com UTF-8."""
|
||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
with output_path.open("w", encoding="utf-8") as f:
|
||||
json.dump(report.to_dict(), f, ensure_ascii=False, indent=2)
|
||||
|
||||
|
||||
def process_batch(
|
||||
input_path: Path,
|
||||
output_path: Path | None = None,
|
||||
limit: int | None = None,
|
||||
language_override: str | None = None,
|
||||
timeout: int = 30,
|
||||
silent: bool = False,
|
||||
) -> ExtractionBatchReport:
|
||||
"""
|
||||
Executa o pipeline completo de extração em lote para o arquivo de entrada.
|
||||
"""
|
||||
query, search_lang, input_articles = load_search_json(input_path)
|
||||
effective_lang = language_override or search_lang or "en"
|
||||
|
||||
if limit is not None and limit > 0:
|
||||
input_articles = input_articles[:limit]
|
||||
|
||||
total = len(input_articles)
|
||||
log_info(
|
||||
f"🚀 Iniciando extração de {total} artigo(s) a partir de '{input_path}' (Idioma NLP: '{effective_lang}')...",
|
||||
silent=silent,
|
||||
)
|
||||
|
||||
# Determinar caminho de saída padrão se não especificado
|
||||
if output_path is None:
|
||||
output_path = input_path.parent / f"{input_path.stem}_extracted.json"
|
||||
|
||||
extracted_list: list[ExtractedArticle] = []
|
||||
successful_count = 0
|
||||
failed_count = 0
|
||||
start_time = time.time()
|
||||
|
||||
with ArticleCrawler(timeout_sec=timeout) as crawler:
|
||||
for idx, article in enumerate(input_articles, start=1):
|
||||
url = article.url
|
||||
log_info(f"🌐 [{idx}/{total}] Navegando com Foxcape: {url}", silent=silent)
|
||||
|
||||
try:
|
||||
html, page_title, http_status = crawler.crawl(url)
|
||||
|
||||
log_info(
|
||||
f"⚙️ [{idx}/{total}] Processando extratores (Trafilatura, Newspaper4k, Readability)...",
|
||||
silent=silent,
|
||||
)
|
||||
traf_data, newspaper_data, readability_data = extract_all_engines(
|
||||
html=html, url=url, language=effective_lang
|
||||
)
|
||||
|
||||
extracted_article = ExtractedArticle(
|
||||
input_meta=article,
|
||||
extraction_status="success",
|
||||
error_message=None,
|
||||
crawled_url=url,
|
||||
page_title=page_title,
|
||||
http_status=http_status,
|
||||
trafilatura=traf_data,
|
||||
newspaper4k=newspaper_data,
|
||||
readability=readability_data,
|
||||
)
|
||||
successful_count += 1
|
||||
log_info(
|
||||
f'✅ [{idx}/{total}] Sucesso (Título: "{article.titulo[:50]}...")',
|
||||
silent=silent,
|
||||
)
|
||||
|
||||
except Exception as exc:
|
||||
failed_count += 1
|
||||
error_msg = str(exc)
|
||||
log_info(
|
||||
f"⚠️ [{idx}/{total}] Falha ao processar URL '{url}': {error_msg}",
|
||||
silent=silent,
|
||||
)
|
||||
extracted_article = ExtractedArticle(
|
||||
input_meta=article,
|
||||
extraction_status="failed",
|
||||
error_message=error_msg,
|
||||
crawled_url=url,
|
||||
page_title=None,
|
||||
http_status=None,
|
||||
trafilatura=None,
|
||||
newspaper4k=None,
|
||||
readability=None,
|
||||
)
|
||||
|
||||
extracted_list.append(extracted_article)
|
||||
|
||||
elapsed = time.time() - start_time
|
||||
now_iso = datetime.now(timezone.utc).isoformat()
|
||||
|
||||
report = ExtractionBatchReport(
|
||||
source_file=str(input_path),
|
||||
processed_at=now_iso,
|
||||
total_articles=total,
|
||||
successful_articles=successful_count,
|
||||
failed_articles=failed_count,
|
||||
articles=extracted_list,
|
||||
)
|
||||
|
||||
save_extracted_json(report, output_path)
|
||||
log_info(f"💾 Relatório final gravado com sucesso em: '{output_path}'", silent=silent)
|
||||
log_info(
|
||||
f"📊 Resumo: {total} total | {successful_count} sucessos | {failed_count} falhas | Tempo: {elapsed:.2f}s",
|
||||
silent=silent,
|
||||
)
|
||||
|
||||
return report
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Interface CLI
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def parse_arguments(args: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Extrator e Parser de Artigos Multimotor (Foxcape + Trafilatura + Newspaper4k + Readability)",
|
||||
formatter_class=argparse.RawTextHelpFormatter,
|
||||
)
|
||||
parser.add_argument(
|
||||
"-i",
|
||||
"--input",
|
||||
required=True,
|
||||
type=str,
|
||||
help="Caminho para o arquivo JSON de busca de notícias (ex: out/river_plate.json)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-o",
|
||||
"--output",
|
||||
required=False,
|
||||
type=str,
|
||||
default=None,
|
||||
help="Caminho do arquivo JSON de destino (padrão: <input_stem>_extracted.json)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-l",
|
||||
"--limit",
|
||||
required=False,
|
||||
type=int,
|
||||
default=None,
|
||||
help="Limita a quantidade máxima de artigos a serem processados",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--lang",
|
||||
"--language",
|
||||
dest="language",
|
||||
required=False,
|
||||
type=str,
|
||||
default=None,
|
||||
help="Sobrescreve o código de idioma para o NLP do Newspaper4k (ex: pt, es, en)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-t",
|
||||
"--timeout",
|
||||
required=False,
|
||||
type=int,
|
||||
default=30,
|
||||
help="Timeout em segundos para carregamento do DOM de cada página no Foxcape (padrão: 30)",
|
||||
)
|
||||
parser.add_argument(
|
||||
"-s",
|
||||
"--silent",
|
||||
action="store_true",
|
||||
help="Suprime mensagens de log e progresso no stderr",
|
||||
)
|
||||
return parser.parse_args(args)
|
||||
|
||||
|
||||
def main(args: list[str] | None = None) -> int:
|
||||
try:
|
||||
parsed = parse_arguments(args)
|
||||
input_path = Path(parsed.input)
|
||||
output_path = Path(parsed.output) if parsed.output else None
|
||||
|
||||
if not input_path.exists():
|
||||
sys.stderr.write(f"Erro: Arquivo de entrada '{input_path}' não existe.\n")
|
||||
return 1
|
||||
|
||||
process_batch(
|
||||
input_path=input_path,
|
||||
output_path=output_path,
|
||||
limit=parsed.limit,
|
||||
language_override=parsed.language,
|
||||
timeout=parsed.timeout,
|
||||
silent=parsed.silent,
|
||||
)
|
||||
return 0
|
||||
except KeyboardInterrupt:
|
||||
sys.stderr.write("\nExecução cancelada pelo usuário.\n")
|
||||
return 130
|
||||
except Exception as exc:
|
||||
sys.stderr.write(f"Erro fatal durante a execução: {exc}\n")
|
||||
return 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -42,9 +42,7 @@ class SearchQuery:
|
||||
raise ValueError("A palavra-chave não pode ser vazia.")
|
||||
|
||||
if not self.language or len(self.language.strip()) < 2:
|
||||
raise ValueError(
|
||||
"O idioma deve conter pelo menos 2 caracteres (ex: 'pt', 'en', 'es')."
|
||||
)
|
||||
raise ValueError("O idioma deve conter pelo menos 2 caracteres (ex: 'pt', 'en', 'es').")
|
||||
|
||||
if self.max_pages < 1 or self.max_pages > 10:
|
||||
raise ValueError("O número máximo de páginas deve estar entre 1 e 10.")
|
||||
@@ -434,9 +432,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||
verbose = not getattr(args, "silent", False)
|
||||
|
||||
try:
|
||||
result = extract_google_news(
|
||||
query, resolve_urls=args.resolve_urls, verbose=verbose
|
||||
)
|
||||
result = extract_google_news(query, resolve_urls=args.resolve_urls, verbose=verbose)
|
||||
json_output = json.dumps(
|
||||
result.to_dict(),
|
||||
ensure_ascii=False,
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
|
||||
|
||||
**Purpose**: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
|
||||
**Created**: 2026-08-20
|
||||
**Feature**: [spec.md](../spec.md) | [data-model.md](../data-model.md) | [contracts/](../contracts/)
|
||||
|
||||
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
|
||||
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
|
||||
|
||||
---
|
||||
|
||||
## 1. Requirement Completeness
|
||||
|
||||
- [x] CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
|
||||
- [x] CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
|
||||
- [x] CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
|
||||
- [x] CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
|
||||
- [x] CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
|
||||
- [x] CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]
|
||||
|
||||
---
|
||||
|
||||
## 2. Requirement Clarity & Non-Ambiguity
|
||||
|
||||
- [x] CHK007 Is the default naming convention for output JSON (`<input_stem>_extracted.json`) unambiguously specified when `--output` is omitted? [Clarity, Spec §FR-007, Clarifications]
|
||||
- [x] CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
|
||||
- [x] CHK009 Is the mechanism for inheriting the `language` parameter with fallback to `"en"` clearly defined? [Clarity, Spec §FR-005, Clarifications]
|
||||
- [x] CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 3. Requirement Consistency & Data Contracts
|
||||
|
||||
- [x] CHK011 Do entity field names in `data-model.md` align consistently with the schemas in `contracts/json-schema.md`? [Consistency, DataModel §1.5, Contract §JSON]
|
||||
- [x] CHK012 Are CLI argument definitions in `contracts/cli-contract.md` aligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI]
|
||||
- [x] CHK013 Is stream segregation (`stderr` for progress logs, `stdout` for JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 4. Scenario & Edge Case Coverage
|
||||
|
||||
- [x] CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
|
||||
- [x] CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
|
||||
- [x] CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
|
||||
- [x] CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
|
||||
- [x] CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 5. Non-Functional & Operational Readiness
|
||||
|
||||
- [x] CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
|
||||
- [x] CHK020 Are execution sampling requirements via `--limit` testable within 30 seconds? [Measurability, Spec §SC-004]
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All 20 items reviewed and validated against `docs/prd_extrator_artigos_nlp.md` and spec artifacts.
|
||||
- Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
|
||||
- Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.
|
||||
@@ -0,0 +1,34 @@
|
||||
# Specification Quality Checklist: Article Content Multi-Engine Extractor
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-08-20
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs) in user-facing outcomes
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- All items passed specification validation. Ready for planning phase (`/speckit-plan`).
|
||||
@@ -0,0 +1,47 @@
|
||||
# CLI Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## 1. Comando de Execução
|
||||
|
||||
```bash
|
||||
python scripts/extract_article_contents.py [OPTIONS]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Argumentos e Flags
|
||||
|
||||
| Flag Curta | Flag Longa | Tipo | Obrigatório | Padrão | Descrição |
|
||||
|---|---|---|---|---|---|
|
||||
| `-i` | `--input` | String (Path) | **Sim** | — | Caminho para o arquivo JSON de entrada (ex: `out/river_plate.json`). |
|
||||
| `-o` | `--output` | String (Path) | Não | `<input_stem>_extracted.json` | Caminho do arquivo JSON de destino. Se omitido, salva no mesmo diretório com sufixo `_extracted.json`. |
|
||||
| `-l` | `--limit` | Inteiro | Não | Todos | Limita a quantidade máxima de artigos a serem processados (útil para amostragem/testes). |
|
||||
| `--lang` | `--language` | String | Não | Do JSON / `"en"` | Sobrescreve o código de idioma para o módulo de NLP do Newspaper4k (ex: `pt`, `es`, `en`). |
|
||||
| `-t` | `--timeout` | Inteiro | Não | `30` | Timeout em segundos para o carregamento do DOM de cada página no Foxcape. |
|
||||
| `-s` | `--silent` | Flag booleana | Não | `False` | Suprime mensagens visuais de progresso e logs em `stderr`. |
|
||||
| `-h` | `--help` | Flag booleana | Não | `False` | Exibe manual de ajuda com todos os parâmetros disponíveis. |
|
||||
|
||||
---
|
||||
|
||||
## 3. Códigos de Saída (Exit Codes)
|
||||
|
||||
| Código | Significado | Condição |
|
||||
|---|---|---|
|
||||
| `0` | **Sucesso** | Execução concluída e arquivo JSON de saída gravado com êxito (mesmo que artigos individuais tenham falhado). |
|
||||
| `1` | **Erro de Argumento** | Arquivo de entrada inexistente, formato inválido ou parâmetros numéricos fora dos limites. |
|
||||
| `2` | **Erro de Inicialização** | Falha ao inicializar o motor Foxcape/navegador headless ou falta de dependências essenciais. |
|
||||
|
||||
---
|
||||
|
||||
## 4. Comportamento de Streams (I/O)
|
||||
|
||||
- **`stderr`**: Recebe mensagens de status informativas em tempo real:
|
||||
```text
|
||||
[INFO] 🚀 Iniciando extração de 10 artigos a partir de 'out/river_plate.json'
|
||||
[INFO] 🌐 [1/10] Foxcape navegando: https://www.tycsports.com/...
|
||||
[INFO] ⚙️ [1/10] Extraindo dados (Trafilatura, Newspaper4k, Readability)...
|
||||
[INFO] ✅ [1/10] Sucesso (Título: "Los puntajes de River...")
|
||||
...
|
||||
[INFO] 💾 Relatório final gravado com sucesso em: 'out/river_plate_extracted.json'
|
||||
[INFO] 📊 Resumo: 10 total | 10 sucessos | 0 falhas | Tempo: 14.2s
|
||||
```
|
||||
- **`stdout`**: Mantém-se silencioso se `--output` for fornecido (ou padrão), ou emite o JSON final caso o usuário redirecione a saída explicitamente.
|
||||
@@ -0,0 +1,88 @@
|
||||
# JSON Schema Contract: Article Content Multi-Engine Extractor
|
||||
|
||||
## 1. Schema de Entrada (Input JSON)
|
||||
|
||||
O arquivo de entrada deve conter a seguinte estrutura JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"query": "string (opcional)",
|
||||
"language": "string (opcional, ex: 'pt', 'es', 'en')",
|
||||
"locale": "string (opcional, ex: 'BR', 'AR', 'US')",
|
||||
"total_itens": "integer (opcional)",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "string (obrigatório)",
|
||||
"url": "string (obrigatório, URL HTTP/HTTPS)",
|
||||
"subtitulo": "string (opcional)",
|
||||
"quando_publicado": "string (opcional)",
|
||||
"pagina": "integer (opcional, default: 1)"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Schema de Saída (Output JSON)
|
||||
|
||||
O arquivo gerado conterá a estrutura consolidada:
|
||||
|
||||
```json
|
||||
{
|
||||
"source_file": "out/river_plate.json",
|
||||
"processed_at": "2026-08-20T15:30:00.000000+00:00",
|
||||
"total_articles": 1,
|
||||
"successful_articles": 1,
|
||||
"failed_articles": 0,
|
||||
"articles": [
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"url": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"subtitulo": "River empató sem gols...",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 03:27:26 GMT",
|
||||
"pagina": 1
|
||||
},
|
||||
"extraction_status": "success",
|
||||
"error_message": null,
|
||||
"crawled_url": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"page_title": "Los puntajes de River vs. Independiente Santa Fe - TyC Sports",
|
||||
"http_status": 200,
|
||||
"trafilatura": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"author": "Ernesto Provitilo",
|
||||
"date": "2026-08-20",
|
||||
"description": "El análisis uno por uno...",
|
||||
"categories": ["River Plate", "Copa Sudamericana"],
|
||||
"tags": ["River", "Santa Fe"],
|
||||
"canonical_url": "https://www.tycsports.com/river-plate/los-puntajes-id755914.html",
|
||||
"text": "Franco Armani (6): Seguro en las pocas llegadas del rival...",
|
||||
"raw_json": { ... },
|
||||
"error": null
|
||||
},
|
||||
"newspaper4k": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"authors": ["Ernesto Provitilo"],
|
||||
"publish_date": "2026-08-20T03:27:26",
|
||||
"text": "Franco Armani (6): Seguro en las pocas llegadas del rival...",
|
||||
"summary": "Resumo gerado por NLP com as sentenças principais...",
|
||||
"keywords": ["river", "santa fe", "puntajes", "armani"],
|
||||
"top_image": "https://media.tycsports.com/adjuntos/800/2026/08/20/armani.jpg",
|
||||
"images": [
|
||||
"https://media.tycsports.com/adjuntos/800/2026/08/20/armani.jpg"
|
||||
],
|
||||
"meta_data": { ... },
|
||||
"error": null
|
||||
},
|
||||
"readability": {
|
||||
"title": "Los puntajes de River vs. Independiente Santa Fe",
|
||||
"short_title": "Los puntajes de River",
|
||||
"cleaned_html": "<div><p>Franco Armani (6): Seguro en las pocas llegadas...</p></div>",
|
||||
"cleaned_text": "Franco Armani (6): Seguro en las pocas llegadas...",
|
||||
"error": null
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,94 @@
|
||||
# Data Model: Article Content Multi-Engine Extractor
|
||||
|
||||
## 1. Entities & Value Objects
|
||||
|
||||
### 1.1 InputArticle (Value Object)
|
||||
Representa uma notícia contida no arquivo JSON de entrada.
|
||||
|
||||
| Campo | Tipo | Obrigatório | Descrição |
|
||||
|---|---|---|---|
|
||||
| `titulo` | `str` | Sim | Título original capturado na busca. |
|
||||
| `url` | `str` | Sim | URL final/resolvida da matéria. |
|
||||
| `pagina` | `int` | Não (default: 1) | Página em que o artigo foi encontrado. |
|
||||
| `subtitulo` | `str | None` | Não | Subtítulo ou resumo do feed RSS. |
|
||||
| `quando_publicado` | `str | None` | Não | String de data/hora original da listagem. |
|
||||
|
||||
---
|
||||
|
||||
### 1.2 TrafilaturaData (Value Object)
|
||||
Dados estruturados extraídos pelo motor Trafilatura.
|
||||
|
||||
| Campo | Tipo | Descrição |
|
||||
|---|---|---|
|
||||
| `title` | `str | None` | Título do artigo extraído pelo Trafilatura. |
|
||||
| `author` | `str | None` | Autor(es) identificados. |
|
||||
| `date` | `str | None` | Data de publicação (ISO YYYY-MM-DD se identificada). |
|
||||
| `description` | `str | None` | Descrição editorial / lead. |
|
||||
| `categories` | `list[str]` | Categorias editoriais extraídas. |
|
||||
| `tags` | `list[str]` | Tags associadas ao artigo. |
|
||||
| `canonical_url` | `str | None` | URL canônica declarada no HTML. |
|
||||
| `text` | `str` | Texto principal limpo e higienizado. |
|
||||
| `raw_json` | `dict[str, Any]` | Payload completo retornado pelo `trafilatura.extract(..., output_format='json')`. |
|
||||
| `error` | `str | None` | Mensagem de erro caso o motor falhe. |
|
||||
|
||||
---
|
||||
|
||||
### 1.3 NewspaperData (Value Object)
|
||||
Dados estruturados e enriquecidos com NLP pelo motor Newspaper4k.
|
||||
|
||||
| Campo | Tipo | Descrição |
|
||||
|---|---|---|
|
||||
| `title` | `str | None` | Título identificado pelo Newspaper. |
|
||||
| `authors` | `list[str]` | Lista de autores extraídos. |
|
||||
| `publish_date` | `str | None` | Data de publicação formatada em ISO string. |
|
||||
| `text` | `str` | Texto integral limpo da matéria. |
|
||||
| `summary` | `str | None` | Resumo automático gerado pelo módulo de NLP. |
|
||||
| `keywords` | `list[str]` | Palavras-chave relevantes identificadas por NLP. |
|
||||
| `top_image` | `str | None` | URL da imagem de destaque principal. |
|
||||
| `images` | `list[str]` | Lista de URLs de imagens presentes no artigo. |
|
||||
| `meta_data` | `dict[str, Any]` | Dicionário com metadados brutos OpenGraph e Schema. |
|
||||
| `error` | `str | None` | Mensagem de erro caso o motor falhe. |
|
||||
|
||||
---
|
||||
|
||||
### 1.4 ReadabilityData (Value Object)
|
||||
Dados higienizados pelo algoritmo Readability (`readability-lxml`).
|
||||
|
||||
| Campo | Tipo | Descrição |
|
||||
|---|---|---|
|
||||
| `title` | `str | None` | Título limpo da página. |
|
||||
| `short_title` | `str | None` | Título curto/resumido. |
|
||||
| `cleaned_html` | `str | None` | Bloco HTML do corpo do artigo higienizado sem anúncios/scripts. |
|
||||
| `cleaned_text` | `str | None` | Texto puro derivado do corpo higienizado. |
|
||||
| `error` | `str | None` | Mensagem de erro caso o motor falhe. |
|
||||
|
||||
---
|
||||
|
||||
### 1.5 ExtractedArticle (Entity)
|
||||
Resultado consolidado da extração de uma notícia específica.
|
||||
|
||||
| Campo | Tipo | Descrição |
|
||||
|---|---|---|
|
||||
| `input_meta` | `InputArticle` | Metadados da notícia original. |
|
||||
| `extraction_status` | `Literal["success", "failed"]` | Status global do processamento do artigo. |
|
||||
| `error_message` | `str | None` | Mensagem de erro se o carregamento da página falhou. |
|
||||
| `crawled_url` | `str` | URL efetivamente navegada no navegador. |
|
||||
| `page_title` | `str | None` | Título retornado pelo DOM (`document.title`). |
|
||||
| `http_status` | `int | None` | Código HTTP retornado pelo servidor (se disponível). |
|
||||
| `trafilatura` | `TrafilaturaData | None` | Resultado do motor Trafilatura. |
|
||||
| `newspaper4k` | `NewspaperData | None` | Resultado do motor Newspaper4k. |
|
||||
| `readability` | `ReadabilityData | None` | Resultado do motor Readability. |
|
||||
|
||||
---
|
||||
|
||||
### 1.6 ExtractionBatchReport (Aggregate Root)
|
||||
Relatório consolidado de saída do processamento de um lote.
|
||||
|
||||
| Campo | Tipo | Descrição |
|
||||
|---|---|---|
|
||||
| `source_file` | `str` | Caminho do arquivo JSON de entrada processado. |
|
||||
| `processed_at` | `str` | Timestamp ISO 8601 UTC do momento da execução. |
|
||||
| `total_articles` | `int` | Total de artigos processados do arquivo de entrada. |
|
||||
| `successful_articles` | `int` | Quantidade de artigos extraídos com sucesso. |
|
||||
| `failed_articles` | `int` | Quantidade de artigos que falharam na extração. |
|
||||
| `articles` | `list[ExtractedArticle]` | Lista de artigos enriquecidos com suas extrações. |
|
||||
@@ -0,0 +1,85 @@
|
||||
# Implementation Plan: Article Content Multi-Engine Extractor
|
||||
|
||||
**Branch**: `003-article-content-extractor` | **Date**: 2026-08-20 | **Spec**: [spec.md](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/003-article-content-extractor/spec.md)
|
||||
|
||||
**Input**: Feature specification from `specs/003-article-content-extractor/spec.md`
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
Construção do extrator de conteúdo de artigos em lote (`scripts/extract_article_contents.py`), que consome listagens de notícias em JSON, acessa e renderiza as páginas de forma stealth via `foxcape` em modo headless reutilizando sessão de navegador, e executa uma tríplice extração de conteúdo com **Trafilatura**, **Newspaper4k** e **Readability**, salvando o resultado consolidado e higienizado em JSON na pasta `out/`.
|
||||
|
||||
---
|
||||
|
||||
## Technical Context
|
||||
|
||||
**Language/Version**: Python 3.10+
|
||||
**Primary Dependencies**: `foxcape` (Camoufox / stealth scraping), `trafilatura` (artigo e metadados), `newspaper4k` (artigo e NLP), `readability-lxml` (miolo e legibilidade), `beautifulsoup4`, `lxml`
|
||||
**Storage**: Arquivos JSON locais no diretório `out/`
|
||||
**Testing**: `pytest` com testes unitários e de integração mockando/testando o pipeline
|
||||
**Target Platform**: Windows / Linux / macOS (Terminal CLI)
|
||||
**Project Type**: CLI tool & modular extraction engine
|
||||
**Performance Goals**: Processamento em lote mantendo sessão de navegador ativa (estimativa de 1 a 2 segundos por notícia com DOMContentLoaded)
|
||||
**Constraints**: Operação 100% headless, descarte de HTML bruto da memória após extração para conter volume, logs em `stderr`
|
||||
**Scale/Scope**: Lotes de 1 a 100+ notícias por execução
|
||||
|
||||
---
|
||||
|
||||
## Constitution Check
|
||||
|
||||
*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.*
|
||||
|
||||
| Princípio | Avaliação | Status |
|
||||
|---|---|---|
|
||||
| **I. Library / Modular Design** | Classes de extração desacopladas por motor (`TrafilaturaExtractor`, `NewspaperExtractor`, `ReadabilityExtractor`). | ✅ Aprovado |
|
||||
| **II. CLI Interface** | CLI via `scripts/extract_article_contents.py` com flags descritivas, `stderr` para status e `stdout` para JSON. | ✅ Aprovado |
|
||||
| **III. Test-First / Automated Tests** | Testes automatizados cobrindo parsing de JSON, orquestração dos 3 motores e resiliência a falhas de rede. | ✅ Aprovado |
|
||||
| **IV. Simplicity & YAGNI** | Uso direto dos módulos especializados existentes sem sobre-engenharia desnecessária. | ✅ Aprovado |
|
||||
|
||||
---
|
||||
|
||||
## Project Structure
|
||||
|
||||
### Documentation (this feature)
|
||||
|
||||
```text
|
||||
specs/003-article-content-extractor/
|
||||
├── plan.md # Este plano de implementação
|
||||
├── research.md # Decisões técnicas e tradeoffs
|
||||
├── data-model.md # Entidades e modelos de dados
|
||||
├── quickstart.md # Guia de validação e execução
|
||||
├── contracts/
|
||||
│ ├── cli-contract.md # Contrato de linha de comando
|
||||
│ └── json-schema.md # Esquemas JSON de entrada e saída
|
||||
└── checklists/
|
||||
└── requirements.md # Checklist de validação da especificação
|
||||
```
|
||||
|
||||
### Source Code Layout
|
||||
|
||||
```text
|
||||
scripts/
|
||||
├── extract_google_news.py # Extrator RSS do Google News existente
|
||||
└── extract_article_contents.py # [NEW] Extrator e Parser Multimotor de Artigos
|
||||
|
||||
tests/
|
||||
├── test_extract_google_news.py # Testes do extrator Google News existente
|
||||
└── test_extract_article_contents.py # [NEW] Testes unitários e de integração do novo extrator
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Phases
|
||||
|
||||
### Phase 0: Outline & Research *(Completed)*
|
||||
- Decisões de arquitetura consolidadas em [research.md](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/003-article-content-extractor/research.md).
|
||||
- Definição do uso de sessão única de navegador `Foxcape(config=FoxcapeConfig(headless=True))` para aceleração em lote.
|
||||
- Definição da estratégia de try/catch em 2 níveis para resiliência máxima.
|
||||
|
||||
### Phase 1: Design & Contracts *(Completed)*
|
||||
- Entidades e contratos definidos em [data-model.md](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/003-article-content-extractor/data-model.md) e [contracts/](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/003-article-content-extractor/contracts/).
|
||||
- Guia de execução rápida e validação em [quickstart.md](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/specs/003-article-content-extractor/quickstart.md).
|
||||
|
||||
### Phase 2: Tasks & Execution *(Completed)*
|
||||
- Decomposição das tarefas de implementação em `tasks.md` e execução 100% concluída.
|
||||
@@ -0,0 +1,77 @@
|
||||
# Quickstart & Validation Guide: Article Content Multi-Engine Extractor
|
||||
|
||||
Este guia descreve os passos para executar, testar e validar o extrator multimotor de artigos de notícias.
|
||||
|
||||
---
|
||||
|
||||
## 1. Pré-requisitos
|
||||
|
||||
Certifique-se de que as dependências necessárias estão instaladas no ambiente Python:
|
||||
|
||||
```bash
|
||||
pip install foxcape trafilatura newspaper4k readability-lxml beautifulsoup4 lxml pytest
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Cenários de Validação
|
||||
|
||||
### Cenário 1: Extração com Amostragem Rápida (Limit 2)
|
||||
Testa o fluxo completo ponta a ponta com apenas 2 notícias para validação imediata:
|
||||
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json --limit 2
|
||||
```
|
||||
|
||||
**Resultado esperado:**
|
||||
- Logs informativos no `stderr` indicando o progresso `[1/2]` e `[2/2]`.
|
||||
- Arquivo `out/river_plate_extracted.json` gerado automaticamente.
|
||||
- O JSON contém 2 artigos com os nós `trafilatura`, `newspaper4k` e `readability` populados.
|
||||
|
||||
---
|
||||
|
||||
### Cenário 2: Caminho Customizado de Saída
|
||||
Testa a especificação explícita do arquivo de saída:
|
||||
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json -o out/custom_test.json --limit 1
|
||||
```
|
||||
|
||||
**Resultado esperado:**
|
||||
- Arquivo `out/custom_test.json` criado com 1 artigo extraído com sucesso.
|
||||
|
||||
---
|
||||
|
||||
### Cenário 3: Modo Silencioso (`--silent`)
|
||||
Testa a supressão de logs para integração em automações/pipes:
|
||||
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json --limit 1 --silent
|
||||
```
|
||||
|
||||
**Resultado esperado:**
|
||||
- Nenhuma saída de log impressa no terminal.
|
||||
- Código de saída 0 retornado.
|
||||
|
||||
---
|
||||
|
||||
### Cenário 4: Resiliência contra URLs Inválidas
|
||||
Testa como o sistema lida com falhas pontuais de conexão ou páginas offline sem quebrar o lote:
|
||||
|
||||
```bash
|
||||
# Executa contra fixture de teste contendo URLs inexistentes
|
||||
pytest tests/test_extract_article_contents.py -k "test_resilience_on_failed_url"
|
||||
```
|
||||
|
||||
**Resultado esperado:**
|
||||
- O teste passa confirmando que o artigo com erro recebeu `extraction_status: "failed"` e os demais concluíram com sucesso.
|
||||
|
||||
---
|
||||
|
||||
## 3. Validação Automatizada de Testes
|
||||
|
||||
Executar a suíte de testes unitários e de integração:
|
||||
|
||||
```bash
|
||||
pytest tests/test_extract_article_contents.py -v
|
||||
```
|
||||
@@ -0,0 +1,50 @@
|
||||
# Research: Article Content Multi-Engine Extractor
|
||||
|
||||
## 1. Technical Decisions & Tradeoffs
|
||||
|
||||
### Decision 1: Motor de Navegação e Renderização com `foxcape` em Sessão Única
|
||||
- **Decision**: Utilizar `foxcape` com `FoxcapeConfig(headless=True, humanize=False)` reutilizando uma única instância de navegador através de context manager (`with Foxcape(...) as scraper:`) para todo o lote.
|
||||
- **Rationale**:
|
||||
- Abrir e fechar o navegador (Camoufox) para cada URL aumentaria o tempo total de processamento em 3 a 5 segundos por artigo.
|
||||
- Reutilizar a sessão mantém a conexão quente, acelera o carregamento do DOM (`wait_until="domcontentloaded"`) e reduz significativamente o consumo de CPU/RAM.
|
||||
- O modo stealth e as evasões de fingerprinting do Foxcape contornam bloqueios Cloudflare, TLS e proteções comuns em portais de notícias.
|
||||
- **Alternatives Considered**:
|
||||
- `requests` / `httpx`: Muito rápidos, porém não executam JavaScript nem resolvem páginas que necessitam de renderização DOM dinâmica (SPAs).
|
||||
- `playwright` padrão: Suscetível a detecção anti-bot e requer configuração manual de stealth plugins.
|
||||
|
||||
---
|
||||
|
||||
### Decision 2: Orquestração Tripla de Extração de Conteúdo (NLP & Web Scraping)
|
||||
- **Decision**: Executar 3 motores de extração complementares e consolidados:
|
||||
1. **Trafilatura**: Padrão ouro em extração de texto limpo, metadados editoriais (`author`, `date`, `categories`, `tags`, `canonical_url`) e estrutura JSON nativa.
|
||||
2. **Newspaper4k**: Processamento avançado de artigo (`Article`), extração de autores, data de publicação, imagens (`top_image`, `images`), e processamento NLP nativo (`nlp()`) gerando resumo automático e *keywords* no idioma do artigo.
|
||||
3. **Readability (`readability-lxml`)**: Heurística clássica de legibilidade (Arc90) para isolar o nó HTML principal sem anúncios ou elementos supérfluos, além de extrair título limpo.
|
||||
- **Rationale**: Cada motor possui pontos fortes distintos. A combinação dos três em uma única passagem oferece a visão mais rica e confiável possível sobre o artigo.
|
||||
- **Alternatives Considered**:
|
||||
- Usar apenas um dos extratores: Perderia a complementaridade (ex.: Trafilatura tem melhor parsing de texto, mas Newspaper4k oferece NLP de keywords/resumo, e Readability oferece o HTML limpo do corpo).
|
||||
|
||||
---
|
||||
|
||||
### Decision 3: Resiliência e Isolamento de Falhas por Camada
|
||||
- **Decision**: Implementar try/catch defensivo em 2 níveis:
|
||||
1. **Nível de Rede/Navegador**: Se o Foxcape falhar em carregar uma URL (timeout, erro 404, bloqueio), registra `extraction_status: "failed"` com a mensagem de erro e avança para a próxima URL.
|
||||
2. **Nível de Extrator**: Cada extrator (`trafilatura`, `newspaper4k`, `readability`) roda em bloco isolado. Se um falhar, os outros dois concluem normalmente e o campo do extrator com falha registra `{"error": "<motivo>"}`.
|
||||
- **Rationale**: Garante taxa de sucesso máxima para lotes grandes sem interrupção abrupta do processamento.
|
||||
|
||||
---
|
||||
|
||||
### Decision 4: Herança Inteligente de Idioma para NLP
|
||||
- **Decision**: O Newspaper4k recebe o idioma informado no cabeçalho do JSON de busca (`"language": "es"`, `"pt"`, etc.), com fallback padrão para `"en"`, e permite sobrescrita pelo usuário via linha de comando (`-l, --language`).
|
||||
- **Rationale**: As rotinas de NLP do Newspaper4k (extração de palavras-chave e resumo) dependem de dicionários e stopwords específicos do idioma.
|
||||
|
||||
---
|
||||
|
||||
### Decision 5: Gerenciamento de Memória e Descarte do Raw HTML
|
||||
- **Decision**: Descartar a string HTML bruta da memória após a passagem pelos 3 extratores, persistindo no JSON final somente as entidades limpas e estruturadas.
|
||||
- **Rationale**: O HTML bruto de 50 artigos pode ocupar mais de 50MB, tornando o JSON volumoso e lento para análise downstream.
|
||||
|
||||
---
|
||||
|
||||
### Decision 6: Segregação de Streams e Feedback Visual em `stderr`
|
||||
- **Decision**: Emitir logs informativos e de progresso item a item (com numeração `[1/50]`, status e tempos) exclusivamente para `sys.stderr`, mantendo o `sys.stdout` intacto.
|
||||
- **Rationale**: Permite acompanhar a execução no terminal em tempo real sem comprometer a interoperabilidade com pipes UNIX (`jq`, redirecionamentos).
|
||||
@@ -0,0 +1,114 @@
|
||||
# Feature Specification: Article Content Multi-Engine Extractor
|
||||
|
||||
**Feature Branch**: `003-article-content-extractor`
|
||||
**Created**: 2026-08-20
|
||||
**Status**: Completed
|
||||
**Input**: User description: "Extrator e Parser de Artigos Multimotor a partir de listagens JSON usando Foxcape headless e tripla extração com Trafilatura, Newspaper4k e Readability (conforme docs/prd_extrator_artigos_nlp.md)"
|
||||
|
||||
---
|
||||
|
||||
## Clarifications
|
||||
|
||||
### Session 2026-08-20
|
||||
- Q: Como o script deve tratar o armazenamento do HTML bruto (*raw HTML*) baixado pelo Foxcape no arquivo JSON de saída? → A: Não incluir o HTML bruto no JSON final (descartar após extrações e persistir apenas os dados estruturados e limpos dos 3 motores para manter o arquivo leve e performático).
|
||||
- Q: Como o idioma para o processamento de NLP do Newspaper4k deve ser definido durante a extração? → A: Automático via JSON de entrada (herda o campo `"language"` do cabeçalho da busca com fallback para `"en"`), permitindo sobrescrita opcional via flag CLI (`--language` / `-l`).
|
||||
- Q: Qual deve ser o padrão de nomenclatura e localização do arquivo JSON gerado quando o operador não fornecer a flag `--output`? → A: Salvar no mesmo diretório adicionando o sufixo `_extracted.json` ao nome base do arquivo de entrada (ex.: `out/river_plate.json` → `out/river_plate_extracted.json`).
|
||||
|
||||
---
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Extração Completa e Consolidada de Artigos em Lote (Priority: P1) 🌟 MVP
|
||||
|
||||
Como analista ou operador de dados, quero fornecer um arquivo JSON de listagem de notícias e obter como resultado um novo arquivo JSON enriquecido contendo o texto completo limpo, metadados e sumários estruturados de cada notícia processada por múltiplos motores de extração, sem que bloqueios anti-bot impeçam a coleta.
|
||||
|
||||
**Why this priority**: É o objetivo central do produto. Transforma referências e manchetes superficiais em conteúdo aprofundado, higienizado e categorizado para análise ou ingestão posterior.
|
||||
|
||||
**Independent Test**: Executar a extração apontando para um arquivo JSON com notícias válidas (ex.: `out/river_plate.json`) e verificar a geração de um arquivo de saída estruturado em `out/` contendo para cada artigo os blocos preenchidos de extração textual e metadados.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
1. **Given** um arquivo JSON de entrada contendo artigos com URLs válidas, **When** o processo de extração for disparado, **Then** o sistema acessa furtivamente cada URL em modo headless, aguarda o carregamento do DOM, obtém o HTML renderizado e processa simultaneamente a extração por três motores distintos, consolidando os resultados em um único arquivo JSON sem armazenar o HTML bruto.
|
||||
2. **Given** um arquivo de entrada vazio ou sem itens válidos, **When** o processo for executado, **Then** o sistema gera um arquivo de saída indicando 0 artigos processados e finaliza com status de sucesso.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Resiliência e Isolamento de Falhas por Artigo e Motor (Priority: P2)
|
||||
|
||||
Como engenheiro de dados executando rotinas em lote, quero que falhas em URLs individuais (como páginas inexistentes, timeouts ou instabilidade temporária do servidor) ou inconsistências em um dos motores de extração não interrompam o processamento das demais notícias do lote.
|
||||
|
||||
**Why this priority**: Garante que execuções longas com dezenas de notícias não sejam perdidas por falha pontual de um único portal externo.
|
||||
|
||||
**Independent Test**: Executar a extração contra um arquivo contendo uma URL inválida misturada com URLs válidas, confirmando que as válidas foram processadas com sucesso e a inválida foi registrada com status de erro sem abortar o pipeline.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
1. **Given** uma notícia com URL inacessível (ex.: erro 404 ou timeout de conexão), **When** a rotina processa a lista, **Then** o sistema registra o item com status de falha e mensagem explicativa, continuando o processamento do próximo item.
|
||||
2. **Given** um HTML que cause erro em um dos três motores de extração, **When** a etapa de análise é executada, **Then** os outros dois motores continuam sua extração normalmente e o erro do motor específico é encapsulado no registro daquele motor.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Controle de Execução via Linha de Comando e Feedback Visual (Priority: P3)
|
||||
|
||||
Como operador de terminal, quero parametrizar a execução via CLI (definindo arquivo de entrada, caminho de saída opcional com padrão `_extracted.json`, limite de itens, idioma e nível de verbosidade) e acompanhar o progresso visualmente no terminal em tempo real sem comprometer a saída padrão de dados.
|
||||
|
||||
**Why this priority**: Oferece usabilidade, capacidade de testes parciais rápidos (amostragem) e compatibilidade com pipes e automações.
|
||||
|
||||
**Independent Test**: Executar o comando passando a flag `--limit 2` e verificar que apenas 2 notícias foram processadas, com mensagens de progresso emitidas no canal de diagnóstico (`stderr`).
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
1. **Given** a execução via linha de comando com parâmetros `--input out/river_plate.json` (sem `--output`) e `--limit 2`, **When** o processo inicia, **Then** o terminal exibe logs com status e percentual de avanço no canal de erro/diagnóstico, processa estritamente 2 itens e grava o arquivo automaticamente como `out/river_plate_extracted.json`.
|
||||
2. **Given** o uso da flag `--silent`, **When** o script é executado, **Then** nenhuma mensagem de log é impressa no terminal.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- **Página com paywall severo ou bloqueio de bot**: O sistema deve capturar o HTML retornado, registrar eventuais limitações na extração e prosseguir sem quebrar a execução.
|
||||
- **Páginas com renderização pesada via JavaScript (SPA)**: O sistema deve aguardar o evento de carregamento do DOM antes de coletar o HTML para garantir que o conteúdo dinâmico esteja presente.
|
||||
- **Ausência de texto no corpo da notícia (apenas vídeo/galeria de fotos)**: Os extratores devem retornar campos de texto vazios de forma graciosa sem gerar exceções não tratadas.
|
||||
- **Caracteres especiais e encodings variados (UTF-8, Latin-1, etc.)**: Os textos extraídos devem ser normalizados para UTF-8 válido no JSON final.
|
||||
|
||||
---
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
- **FR-001**: O sistema DEVE receber como entrada um arquivo JSON contendo uma lista estruturada de notícias e validar a presença das URLs a serem processadas.
|
||||
- **FR-002**: O sistema DEVE navegar até cada URL utilizando navegação furtiva automatizada em modo headless, aguardando o carregamento completo do DOM.
|
||||
- **FR-003**: O sistema DEVE manter uma única sessão de navegador ativa reutilizada ao longo do lote para otimizar velocidade e consumo de memória.
|
||||
- **FR-004**: O sistema DEVE processar o HTML renderizado através do motor Trafilatura, extraindo texto limpo, título, autor, data, categorias/tags, descrição e metadados estruturados.
|
||||
- **FR-005**: O sistema DEVE processar o HTML renderizado através do motor Newspaper4k, utilizando o idioma herdado do JSON de entrada (com fallback para `"en"` ou sobrescrito por CLI) para extrair corpo do artigo, autores, data de publicação, resumo por NLP, palavras-chave por NLP, imagens e metadados OpenGraph.
|
||||
- **FR-006**: O sistema DEVE processar o HTML renderizado através do motor Readability, extraindo o conteúdo limpo principal (HTML sanitizado e texto puro) e título.
|
||||
- **FR-007**: O sistema DEVE consolidar os resultados dos três motores em um documento JSON único por execução, gravando-o por padrão como `<input_stem>_extracted.json` no mesmo diretório (ou no caminho fornecido via `--output`), sem persistir o HTML bruto baixado.
|
||||
- **FR-008**: O sistema DEVE registrar o status de extração (`success` ou `failed`) e mensagens de erro individuais para cada notícia processada.
|
||||
- **FR-009**: O sistema DEVE fornecer interface de linha de comando (CLI) com suporte a flags de arquivo de entrada (`-i, --input`), saída (`-o, --output`), limite de itens (`--limit`), idioma opcional (`-l, --language`), timeout (`-t, --timeout`) e modo silencioso (`-s, --silent`).
|
||||
- **FR-010**: O sistema DEVE enviar logs de progresso e status em tempo real exclusivamente para o fluxo de erro padrão (`stderr`), preservando o fluxo de saída padrão (`stdout`).
|
||||
|
||||
---
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **InputArticle**: Representa a notícia recebida no arquivo de entrada, contendo título original, URL resolvida, subtítulo, data de publicação da listagem e número da página.
|
||||
- **ExtractedArticleResult**: Representa o resultado consolidado da extração de um artigo, agregando os metadados de entrada, status de coleta, URL final navegada, código de resposta HTTP e os payloads detalhados de cada um dos três extratores (`trafilatura`, `newspaper4k`, `readability`), omitindo o HTML bruto.
|
||||
- **BatchExtractionReport**: Representa o relatório global do lote, contendo metadados de auditoria (arquivo de origem, data/hora de processamento, total de itens, sucessos e falhas) e a lista de `ExtractedArticleResult`.
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: O sistema processa com sucesso pelo menos 90% das notícias válidas fornecidas em lote sem intervenção manual.
|
||||
- **SC-002**: Para páginas padrão de notícias com acesso público, todos os três motores de extração preenchem seus respectivos campos de texto limpo e título.
|
||||
- **SC-003**: A falha no carregamento ou na extração de 1 artigo isolado tem taxa de propagação de erro de 0% sobre os demais itens da fila.
|
||||
- **SC-004**: Operadores conseguem executar amostragens parciais de testes em menos de 30 segundos utilizando a flag de limite de itens.
|
||||
- **SC-005**: O arquivo JSON final gerado é 100% compatível com validadores JSON padrão (UTF-8 formatado).
|
||||
|
||||
---
|
||||
|
||||
## Assumptions
|
||||
|
||||
- Os arquivos JSON de entrada seguirão a estrutura produzida pelo extrator de notícias do Google News deste repositório (com chave `items` e propriedade `url` em cada item).
|
||||
- O ambiente de execução possui conectividade à internet para acessar os portais de notícias.
|
||||
- Recursos de hardware suficientes para execução de um processo de navegador headless (Firefox/Camoufox) em segundo plano.
|
||||
- As dependências de NLP e extração (`foxcape`, `trafilatura`, `newspaper4k`, `readability-lxml`) estarão devidamente instaladas no ambiente Python.
|
||||
@@ -0,0 +1,132 @@
|
||||
# Tasks: Article Content Multi-Engine Extractor
|
||||
|
||||
**Feature**: `003-article-content-extractor`
|
||||
**Spec**: [`specs/003-article-content-extractor/spec.md`](file:///c:/Users/aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\spec.md)
|
||||
**Plan**: [`specs/003-article-content-extractor/plan.md`](file:///c:/Users/aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\plan.md)
|
||||
**Status**: Completed
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Setup & Dependencies
|
||||
|
||||
**Purpose**: Garantir as dependências do ecossistema e a estrutura inicial do projeto.
|
||||
|
||||
- [X] T001 Atualizar dependências em `requirements.txt` incluindo `trafilatura`, `newspaper4k` e `readability-lxml`
|
||||
- [X] T002 [P] Validar importação e disponibilidade das bibliotecas `foxcape`, `trafilatura`, `newspaper` e `readability` no ambiente Python
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Foundational (Estruturas e Modelos Base)
|
||||
|
||||
**Purpose**: Estruturas de dados, contratos de erro e classes base necessárias para todas as histórias de usuário.
|
||||
|
||||
- [X] T003 Definir modelos de dados e dataclasses (`InputArticle`, `TrafilaturaData`, `NewspaperData`, `ReadabilityData`, `ExtractedArticle`, `ExtractionBatchReport`) em `scripts/extract_article_contents.py`
|
||||
- [X] T004 Implementar funções utilitárias de I/O para leitura segura de JSON de busca e gravação com UTF-8 em `scripts/extract_article_contents.py`
|
||||
- [X] T005 [P] Criar suíte de testes base e fixtures de mock de HTML em `tests/test_extract_article_contents.py`
|
||||
|
||||
**Checkpoint**: Estruturas base e fixtures prontas para início do desenvolvimento das histórias de usuário.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 1 - Extração Completa e Consolidada de Artigos em Lote (Priority: P1) 🌟 MVP
|
||||
|
||||
**Goal**: Implementar a navegação headless stealth via Foxcape e os 3 motores de extração (Trafilatura, Newspaper4k, Readability) gerando o JSON consolidado.
|
||||
|
||||
**Independent Test**: Executar contra um HTML de teste ou URL mockada e verificar a extração de texto, títulos, metadados, autores, imagens e sumários NLP em um JSON sem raw HTML.
|
||||
|
||||
### Testes da User Story 1 (TDD)
|
||||
- [X] T006 [P] [US1] Criar testes unitários para o parser `TrafilaturaExtractor` em `tests/test_extract_article_contents.py`
|
||||
- [X] T007 [P] [US1] Criar testes unitários para o parser `NewspaperExtractor` (NLP, autores, imagens, resumo) em `tests/test_extract_article_contents.py`
|
||||
- [X] T008 [P] [US1] Criar testes unitários para o parser `ReadabilityExtractor` (HTML limpo, títulos) em `tests/test_extract_article_contents.py`
|
||||
- [X] T009 [US1] Criar teste de integração para o pipeline completo de extração multimotor em `tests/test_extract_article_contents.py`
|
||||
|
||||
### Implementação da User Story 1
|
||||
- [X] T010 [P] [US1] Implementar classe `TrafilaturaExtractor` para extração máxima de metadados, categorias, tags e texto limpo em `scripts/extract_article_contents.py`
|
||||
- [X] T011 [P] [US1] Implementar classe `NewspaperExtractor` com suporte a herança de idioma e métodos NLP (`parse`, `nlp`) em `scripts/extract_article_contents.py`
|
||||
- [X] T012 [P] [US1] Implementar classe `ReadabilityExtractor` para higienização e extração do miolo textual em `scripts/extract_article_contents.py`
|
||||
- [X] T013 [US1] Implementar gerenciador de sessão persistente do `Foxcape` (`with Foxcape(...)`) com espera de DOM (`domcontentloaded`) em `scripts/extract_article_contents.py`
|
||||
- [X] T014 [US1] Implementar orquestrador de lote e consolidação de resultados (descartando HTML bruto da memória) em `scripts/extract_article_contents.py`
|
||||
|
||||
**Checkpoint**: MVP funcional — O sistema já é capaz de ler uma lista de URLs, navegar com Foxcape, extrair pelos 3 motores e salvar o JSON consolidado.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 2 - Resiliência e Isolamento de Falhas por Artigo e Motor (Priority: P2)
|
||||
|
||||
**Goal**: Garantir tolerância a falhas para que timeouts, erros 404, bloqueios ou quebras em um único motor não abortem o lote.
|
||||
|
||||
**Independent Test**: Executar contra uma lista contendo URLs válidas e inválidas, confirmando que a inválida recebe status `failed` e as válidas continuam normalmente.
|
||||
|
||||
### Testes da User Story 2 (TDD)
|
||||
- [X] T015 [P] [US2] Criar teste para isolamento de erro em falha de navegação (timeout / 404) em `tests/test_extract_article_contents.py`
|
||||
- [X] T016 [P] [US2] Criar teste para isolamento de erro quando um único motor falha em `tests/test_extract_article_contents.py`
|
||||
|
||||
### Implementação da User Story 2
|
||||
- [X] T017 [US2] Implementar tratamento de exceções de rede e status HTTP individual por artigo em `scripts/extract_article_contents.py`
|
||||
- [X] T018 [US2] Implementar tratamento de exceções defensivo e encapsulamento de erro por motor de extração em `scripts/extract_article_contents.py`
|
||||
|
||||
**Checkpoint**: Sistema 100% resiliente contra instabilidades de portais e erros pontuais de parsing.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: User Story 3 - Controle de Execução via CLI e Feedback Visual (Priority: P3)
|
||||
|
||||
**Goal**: Interface de linha de comando completa com flags descritivas, suporte a limites (`--limit`), modo silencioso (`--silent`) e logs em `stderr`.
|
||||
|
||||
**Independent Test**: Executar `python scripts/extract_article_contents.py -i out/river_plate.json --limit 2` e verificar os logs em `stderr` e a criação de `out/river_plate_extracted.json`.
|
||||
|
||||
### Testes da User Story 3 (TDD)
|
||||
- [X] T019 [P] [US3] Criar testes de CLI para parsing de argumentos (`--input`, `--output`, `--limit`, `--language`, `--silent`, `--timeout`) em `tests/test_extract_article_contents.py`
|
||||
- [X] T020 [P] [US3] Criar teste de validação de segregação de streams (`stderr` vs `stdout`) e exit codes em `tests/test_extract_article_contents.py`
|
||||
|
||||
### Implementação da User Story 3
|
||||
- [X] T021 [US3] Implementar parser CLI (`argparse`) com todas as opções e convenção padrão de saída (`<stem>_extracted.json`) em `scripts/extract_article_contents.py`
|
||||
- [X] T022 [US3] Implementar sistema de logging visual em tempo real com emojis e status direcionado exclusivamente para `sys.stderr` em `scripts/extract_article_contents.py`
|
||||
- [X] T023 [US3] Implementar controle de códigos de saída (0 para sucesso, 1 para argumento inválido, 2 para erro de inicialização) em `scripts/extract_article_contents.py`
|
||||
|
||||
**Checkpoint**: Todas as histórias de usuário (US1, US2, US3) implementadas e integradas.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6: Polish & Validação Final
|
||||
|
||||
**Purpose**: Verificação ponta a ponta, documentação e conformidade.
|
||||
|
||||
- [X] T024 [P] Executar suíte completa de testes automatizados com `pytest`
|
||||
- [X] T025 Executar validação real de ponta a ponta contra `out/river_plate.json` gerando `out/river_plate_extracted.json`
|
||||
- [X] T026 [P] Atualizar documentação de uso no `README.md`
|
||||
- [X] T027 Executar `graphify update .` para manter o grafo de conhecimento do repositório sincronizado
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
P1[Phase 1: Setup & Dependencies\n T001, T002] --> P2[Phase 2: Foundational\n T003, T004, T005]
|
||||
P2 --> P3[Phase 3: User Story 1 MVP\n T006-T014]
|
||||
P3 --> P4[Phase 4: User Story 2 Resiliência\n T015-T018]
|
||||
P4 --> P5[Phase 5: User Story 3 CLI & Logs\n T019-T023]
|
||||
P5 --> P6[Phase 6: Polish & Validação\n T024-T027]
|
||||
```
|
||||
|
||||
### Oportunidades de Execução Paralela
|
||||
- **Phase 1**: `T002` pode rodar em paralelo após `T001`.
|
||||
- **Phase 3 (Testes & Parsers)**: `T006`, `T007`, `T008` (testes unitários) e `T010`, `T011`, `T012` (implementações dos 3 parsers) podem ser desenvolvidos em paralelo por atuarem em classes isoladas.
|
||||
- **Phase 4 & 5 (Testes)**: `T015`, `T016`, `T019`, `T020` podem ser implementados em paralelo antes das respectivas integrações no CLI.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### MVP First (User Story 1 Only)
|
||||
1. Completar Fase 1 (Setup) e Fase 2 (Foundational).
|
||||
2. Implementar Fase 3 (User Story 1).
|
||||
3. **Validar MVP**: Testar extração de 1 artigo local com sucesso.
|
||||
|
||||
### Entrega Incremental
|
||||
1. Setup + Foundational → Base pronta.
|
||||
2. User Story 1 → Tripla extração funcional (MVP).
|
||||
3. User Story 2 → Resiliência total contra falhas externas.
|
||||
4. User Story 3 → Interface CLI rica e ergonômica.
|
||||
5. Polish → Testes 100% passando e validação real em lote.
|
||||
@@ -3,8 +3,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
from typing import Any, Dict, List, Optional
|
||||
from src.models import ECPSnapshot, ClassificationResult
|
||||
|
||||
from src.models import ClassificationResult, ECPSnapshot
|
||||
|
||||
|
||||
class BaseNLPAdapter(ABC):
|
||||
@@ -13,12 +13,10 @@ class BaseNLPAdapter(ABC):
|
||||
@abstractmethod
|
||||
def is_available(self) -> bool:
|
||||
"""Return True if the underlying provider or model is installed and configured."""
|
||||
pass
|
||||
|
||||
@abstractmethod
|
||||
def evaluate_similarity(self, text: str, terms: List[str]) -> float:
|
||||
def evaluate_similarity(self, text: str, terms: list[str]) -> float:
|
||||
"""Compute semantic similarity score between text and a set of candidate terms."""
|
||||
pass
|
||||
|
||||
@abstractmethod
|
||||
def disambiguate(
|
||||
@@ -26,6 +24,5 @@ class BaseNLPAdapter(ABC):
|
||||
ecp: ECPSnapshot,
|
||||
content_md: str,
|
||||
initial_result: ClassificationResult,
|
||||
) -> Optional[ClassificationResult]:
|
||||
) -> ClassificationResult | None:
|
||||
"""Optionally refine an ambiguous classification result."""
|
||||
pass
|
||||
|
||||
@@ -6,15 +6,16 @@ without requiring sentence-transformers to be pre-installed in the core environm
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import List, Optional
|
||||
from src.models import ECPSnapshot, ClassificationResult
|
||||
from src.adapters.base import BaseNLPAdapter
|
||||
from src.models import ClassificationResult, ECPSnapshot
|
||||
|
||||
|
||||
class LocalEmbeddingsAdapter(BaseNLPAdapter):
|
||||
"""Optional adapter for local multilingual semantic vector embeddings."""
|
||||
|
||||
def __init__(self, model_name: str = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2") -> None:
|
||||
def __init__(
|
||||
self, model_name: str = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
|
||||
) -> None:
|
||||
self.model_name = model_name
|
||||
self._model = None
|
||||
self._initialized = False
|
||||
@@ -22,11 +23,12 @@ class LocalEmbeddingsAdapter(BaseNLPAdapter):
|
||||
def is_available(self) -> bool:
|
||||
try:
|
||||
import sentence_transformers # noqa: F401
|
||||
|
||||
return True
|
||||
except ImportError:
|
||||
return False
|
||||
|
||||
def evaluate_similarity(self, text: str, terms: List[str]) -> float:
|
||||
def evaluate_similarity(self, text: str, terms: list[str]) -> float:
|
||||
if not self.is_available() or not terms:
|
||||
return 0.0
|
||||
# Placeholder stub for local embedding computation
|
||||
@@ -37,6 +39,6 @@ class LocalEmbeddingsAdapter(BaseNLPAdapter):
|
||||
ecp: ECPSnapshot,
|
||||
content_md: str,
|
||||
initial_result: ClassificationResult,
|
||||
) -> Optional[ClassificationResult]:
|
||||
) -> ClassificationResult | None:
|
||||
# Embeddings adapter does not alter decisions in POC unless explicitly wired
|
||||
return None
|
||||
|
||||
+5
-5
@@ -7,22 +7,22 @@ without requiring OpenAI/Anthropic/Gemini API keys for core POC execution.
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
from typing import List, Optional
|
||||
from src.models import ECPSnapshot, ClassificationResult
|
||||
|
||||
from src.adapters.base import BaseNLPAdapter
|
||||
from src.models import ClassificationResult, ECPSnapshot
|
||||
|
||||
|
||||
class LLMFallbackAdapter(BaseNLPAdapter):
|
||||
"""Optional adapter for LLM fallback boundary disambiguation."""
|
||||
|
||||
def __init__(self, model_name: str = "gpt-4o-mini", api_key: Optional[str] = None) -> None:
|
||||
def __init__(self, model_name: str = "gpt-4o-mini", api_key: str | None = None) -> None:
|
||||
self.model_name = model_name
|
||||
self.api_key = api_key or os.environ.get("OPENAI_API_KEY")
|
||||
|
||||
def is_available(self) -> bool:
|
||||
return bool(self.api_key)
|
||||
|
||||
def evaluate_similarity(self, text: str, terms: List[str]) -> float:
|
||||
def evaluate_similarity(self, text: str, terms: list[str]) -> float:
|
||||
return 0.0
|
||||
|
||||
def disambiguate(
|
||||
@@ -30,7 +30,7 @@ class LLMFallbackAdapter(BaseNLPAdapter):
|
||||
ecp: ECPSnapshot,
|
||||
content_md: str,
|
||||
initial_result: ClassificationResult,
|
||||
) -> Optional[ClassificationResult]:
|
||||
) -> ClassificationResult | None:
|
||||
# If API key is not configured or case is already clear, skip
|
||||
if not self.is_available():
|
||||
return None
|
||||
|
||||
+54
-35
@@ -3,18 +3,15 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from typing import Any, Dict, List, Optional, Tuple
|
||||
from typing import Any
|
||||
|
||||
from src.models import (
|
||||
ECPSnapshot,
|
||||
RelatedEntity,
|
||||
ClassificationResult,
|
||||
ClassificationError,
|
||||
DecisionCategory,
|
||||
ErrorCode,
|
||||
)
|
||||
from src.language import detect_language, normalize_text
|
||||
from src.parser import strip_markdown, extract_evidence_snippets
|
||||
from src.models import (
|
||||
ClassificationResult,
|
||||
DecisionCategory,
|
||||
ECPSnapshot,
|
||||
)
|
||||
from src.parser import extract_evidence_snippets, strip_markdown
|
||||
|
||||
|
||||
def match_phrase_in_text(phrase: str, normalized_text: str) -> bool:
|
||||
@@ -52,10 +49,12 @@ class InherenceClassifier:
|
||||
|
||||
if enable_embeddings:
|
||||
from src.adapters.embeddings import LocalEmbeddingsAdapter
|
||||
|
||||
self._embeddings_adapter = LocalEmbeddingsAdapter()
|
||||
|
||||
if enable_llm:
|
||||
from src.adapters.llm import LLMFallbackAdapter
|
||||
|
||||
self._llm_adapter = LLMFallbackAdapter()
|
||||
|
||||
def classify(self, ecp: ECPSnapshot, content_md: str) -> ClassificationResult:
|
||||
@@ -70,7 +69,7 @@ class InherenceClassifier:
|
||||
|
||||
# 2. Match Target Entity & Aliases (deduplicate normalized terms)
|
||||
all_target_terms = [ecp.target_name] + [a for a in ecp.aliases if a != ecp.target_name]
|
||||
matched_target_terms: List[str] = []
|
||||
matched_target_terms: list[str] = []
|
||||
target_mention_count = 0
|
||||
seen_norm_terms: set[str] = set()
|
||||
|
||||
@@ -85,27 +84,35 @@ class InherenceClassifier:
|
||||
target_mention_count += count
|
||||
|
||||
# 3. Match Context Anchors
|
||||
matched_anchors: List[str] = []
|
||||
matched_anchors: list[str] = []
|
||||
for anchor in ecp.anchors:
|
||||
if match_phrase_in_text(anchor, norm_text):
|
||||
matched_anchors.append(anchor)
|
||||
|
||||
# 4. Match Negative Anchors (homonym disambiguators)
|
||||
matched_negative_anchors: List[str] = []
|
||||
matched_negative_anchors: list[str] = []
|
||||
for neg in ecp.negative_anchors:
|
||||
if match_phrase_in_text(neg, norm_text):
|
||||
matched_negative_anchors.append(neg)
|
||||
|
||||
# 5. Match Related Graph Entities
|
||||
matched_graph_entities: List[Dict[str, Any]] = []
|
||||
matched_graph_entities: list[dict[str, Any]] = []
|
||||
highest_graph_weight = 0.0
|
||||
for rel in ecp.related_entities:
|
||||
rel_name = rel.name if hasattr(rel, "name") else rel.get("name", "")
|
||||
rel_id = rel.entity_id if hasattr(rel, "entity_id") else rel.get("entity_id", "")
|
||||
rel_type = rel.relation_type if hasattr(rel, "relation_type") else rel.get("relation_type", "")
|
||||
rel_weight = float(rel.weight if hasattr(rel, "weight") else rel.get("weight", 1.0))
|
||||
rel_scope = str(rel.scope if hasattr(rel, "scope") else rel.get("scope", "general"))
|
||||
rel_aliases = rel.aliases if hasattr(rel, "aliases") else rel.get("aliases", [])
|
||||
if isinstance(rel, dict):
|
||||
rel_name = rel.get("name", "")
|
||||
rel_id = rel.get("entity_id", "")
|
||||
rel_type = rel.get("relation_type", "")
|
||||
rel_weight = float(rel.get("weight", 1.0))
|
||||
rel_scope = str(rel.get("scope", "general"))
|
||||
rel_aliases = rel.get("aliases", [])
|
||||
else:
|
||||
rel_name = rel.name
|
||||
rel_id = rel.entity_id
|
||||
rel_type = rel.relation_type
|
||||
rel_weight = float(rel.weight)
|
||||
rel_scope = str(rel.scope)
|
||||
rel_aliases = rel.aliases
|
||||
|
||||
rel_terms = [rel_name] + list(rel_aliases)
|
||||
rel_matched = False
|
||||
@@ -114,28 +121,33 @@ class InherenceClassifier:
|
||||
rel_matched = True
|
||||
break
|
||||
if rel_matched:
|
||||
matched_graph_entities.append({
|
||||
"entity_id": rel_id,
|
||||
"name": rel_name,
|
||||
"relation_type": rel_type,
|
||||
"weight": rel_weight,
|
||||
"scope": rel_scope,
|
||||
})
|
||||
if rel_weight > highest_graph_weight:
|
||||
highest_graph_weight = rel_weight
|
||||
matched_graph_entities.append(
|
||||
{
|
||||
"entity_id": rel_id,
|
||||
"name": rel_name,
|
||||
"relation_type": rel_type,
|
||||
"weight": rel_weight,
|
||||
"scope": rel_scope,
|
||||
}
|
||||
)
|
||||
highest_graph_weight = max(highest_graph_weight, rel_weight)
|
||||
|
||||
# 6. Evaluate Decision Rules
|
||||
warnings: List[str] = []
|
||||
warnings: list[str] = []
|
||||
has_direct_match = len(matched_target_terms) > 0
|
||||
has_negative_match = len(matched_negative_anchors) > 0
|
||||
has_graph_match = len(matched_graph_entities) > 0
|
||||
has_anchor_match = len(matched_anchors) > 0
|
||||
|
||||
# Term collection for evidence extraction
|
||||
evidence_terms = matched_target_terms + [g["name"] for g in matched_graph_entities] + matched_anchors
|
||||
evidence_terms = (
|
||||
matched_target_terms + [g["name"] for g in matched_graph_entities] + matched_anchors
|
||||
)
|
||||
|
||||
# Decision 1: Dominant Negative Anchors (overrides passing mentions)
|
||||
if has_negative_match and (not has_anchor_match or len(matched_negative_anchors) >= len(matched_anchors)):
|
||||
if has_negative_match and (
|
||||
not has_anchor_match or len(matched_negative_anchors) >= len(matched_anchors)
|
||||
):
|
||||
decision = DecisionCategory.NOT_RELATED
|
||||
is_inherent = False
|
||||
confidence = 0.90
|
||||
@@ -147,7 +159,9 @@ class InherenceClassifier:
|
||||
# Strong direct match with supporting context
|
||||
decision = DecisionCategory.DIRECT_INHERENT
|
||||
is_inherent = True
|
||||
confidence = min(0.98, 0.85 + (0.04 * len(matched_anchors)) + (0.02 * target_mention_count))
|
||||
confidence = min(
|
||||
0.98, 0.85 + (0.04 * len(matched_anchors)) + (0.02 * target_mention_count)
|
||||
)
|
||||
rationale = (
|
||||
f"Direct match of target entity '{ecp.target_name}' with strong contextual anchor density "
|
||||
f"({len(matched_anchors)} anchor(s) matched)."
|
||||
@@ -168,7 +182,10 @@ class InherenceClassifier:
|
||||
# Contextual inherence via connected graph entity with domain anchor alignment
|
||||
decision = DecisionCategory.CONTEXTUAL_INHERENT
|
||||
is_inherent = True
|
||||
confidence = round(min(0.95, 0.70 + (highest_graph_weight * 0.20) + (0.03 * len(matched_anchors))), 4)
|
||||
confidence = round(
|
||||
min(0.95, 0.70 + (highest_graph_weight * 0.20) + (0.03 * len(matched_anchors))),
|
||||
4,
|
||||
)
|
||||
top_rel = matched_graph_entities[0]
|
||||
rationale = (
|
||||
f"Matched connected entity '{top_rel['name']}' ({top_rel['relation_type']}) "
|
||||
@@ -191,7 +208,9 @@ class InherenceClassifier:
|
||||
decision = DecisionCategory.NOT_RELATED
|
||||
is_inherent = False
|
||||
confidence = 0.85
|
||||
rationale = "General domain topics mentioned, but target entity or related entities are absent."
|
||||
rationale = (
|
||||
"General domain topics mentioned, but target entity or related entities are absent."
|
||||
)
|
||||
|
||||
else:
|
||||
# Completely unrelated
|
||||
|
||||
+492
-55
@@ -4,66 +4,458 @@ from __future__ import annotations
|
||||
|
||||
import re
|
||||
import unicodedata
|
||||
from typing import Dict, List, Set, Tuple
|
||||
|
||||
# Supported ISO 639-1 language codes
|
||||
SUPPORTED_LANGUAGES: Set[str] = {"pt", "en", "es", "de", "it", "fr"}
|
||||
SUPPORTED_LANGUAGES: set[str] = {"pt", "en", "es", "de", "it", "fr"}
|
||||
|
||||
# Characteristic function words / stopwords for deterministic language identification
|
||||
LANGUAGE_STOPWORDS: Dict[str, Set[str]] = {
|
||||
LANGUAGE_STOPWORDS: dict[str, set[str]] = {
|
||||
"pt": {
|
||||
"de", "a", "o", "que", "e", "do", "da", "em", "um", "para", "é", "com", "não",
|
||||
"uma", "os", "no", "se", "na", "por", "mais", "as", "dos", "como", "mas", "foi",
|
||||
"ao", "ele", "das", "tem", "à", "seu", "sua", "ou", "ser", "quando", "muito",
|
||||
"nos", "já", "está", "eu", "também", "só", "pelo", "pela", "até", "isso", "ela",
|
||||
"entre", "depois", "sem", "mesmo", "aos", "ter", "seus", "quem", "nas", "me",
|
||||
"esse", "eles", "estão", "você", "tinha", "foram", "essa", "num", "nem", "suas",
|
||||
"anunciou", "produção", "empresa", "mercado", "setor", "governo", "ano"
|
||||
"de",
|
||||
"a",
|
||||
"o",
|
||||
"que",
|
||||
"e",
|
||||
"do",
|
||||
"da",
|
||||
"em",
|
||||
"um",
|
||||
"para",
|
||||
"é",
|
||||
"com",
|
||||
"não",
|
||||
"uma",
|
||||
"os",
|
||||
"no",
|
||||
"se",
|
||||
"na",
|
||||
"por",
|
||||
"mais",
|
||||
"as",
|
||||
"dos",
|
||||
"como",
|
||||
"mas",
|
||||
"foi",
|
||||
"ao",
|
||||
"ele",
|
||||
"das",
|
||||
"tem",
|
||||
"à",
|
||||
"seu",
|
||||
"sua",
|
||||
"ou",
|
||||
"ser",
|
||||
"quando",
|
||||
"muito",
|
||||
"nos",
|
||||
"já",
|
||||
"está",
|
||||
"eu",
|
||||
"também",
|
||||
"só",
|
||||
"pelo",
|
||||
"pela",
|
||||
"até",
|
||||
"isso",
|
||||
"ela",
|
||||
"entre",
|
||||
"depois",
|
||||
"sem",
|
||||
"mesmo",
|
||||
"aos",
|
||||
"ter",
|
||||
"seus",
|
||||
"quem",
|
||||
"nas",
|
||||
"me",
|
||||
"esse",
|
||||
"eles",
|
||||
"estão",
|
||||
"você",
|
||||
"tinha",
|
||||
"foram",
|
||||
"essa",
|
||||
"num",
|
||||
"nem",
|
||||
"suas",
|
||||
"anunciou",
|
||||
"produção",
|
||||
"empresa",
|
||||
"mercado",
|
||||
"setor",
|
||||
"governo",
|
||||
"ano",
|
||||
},
|
||||
"en": {
|
||||
"the", "be", "to", "of", "and", "a", "in", "that", "have", "i", "it", "for",
|
||||
"not", "on", "with", "he", "as", "you", "do", "at", "this", "but", "his", "by",
|
||||
"from", "they", "we", "say", "her", "she", "or", "an", "will", "my", "one",
|
||||
"all", "would", "there", "their", "what", "so", "up", "out", "if", "about",
|
||||
"who", "get", "which", "go", "me", "when", "make", "can", "like", "time", "no",
|
||||
"just", "him", "know", "take", "people", "into", "year", "your", "good", "some",
|
||||
"could", "them", "see", "other", "than", "then", "now", "look", "only", "come"
|
||||
"the",
|
||||
"be",
|
||||
"to",
|
||||
"of",
|
||||
"and",
|
||||
"a",
|
||||
"in",
|
||||
"that",
|
||||
"have",
|
||||
"i",
|
||||
"it",
|
||||
"for",
|
||||
"not",
|
||||
"on",
|
||||
"with",
|
||||
"he",
|
||||
"as",
|
||||
"you",
|
||||
"do",
|
||||
"at",
|
||||
"this",
|
||||
"but",
|
||||
"his",
|
||||
"by",
|
||||
"from",
|
||||
"they",
|
||||
"we",
|
||||
"say",
|
||||
"her",
|
||||
"she",
|
||||
"or",
|
||||
"an",
|
||||
"will",
|
||||
"my",
|
||||
"one",
|
||||
"all",
|
||||
"would",
|
||||
"there",
|
||||
"their",
|
||||
"what",
|
||||
"so",
|
||||
"up",
|
||||
"out",
|
||||
"if",
|
||||
"about",
|
||||
"who",
|
||||
"get",
|
||||
"which",
|
||||
"go",
|
||||
"me",
|
||||
"when",
|
||||
"make",
|
||||
"can",
|
||||
"like",
|
||||
"time",
|
||||
"no",
|
||||
"just",
|
||||
"him",
|
||||
"know",
|
||||
"take",
|
||||
"people",
|
||||
"into",
|
||||
"year",
|
||||
"your",
|
||||
"good",
|
||||
"some",
|
||||
"could",
|
||||
"them",
|
||||
"see",
|
||||
"other",
|
||||
"than",
|
||||
"then",
|
||||
"now",
|
||||
"look",
|
||||
"only",
|
||||
"come",
|
||||
},
|
||||
"es": {
|
||||
"de", "la", "que", "el", "en", "y", "a", "los", "del", "se", "las", "por", "un",
|
||||
"para", "con", "no", "una", "su", "al", "lo", "como", "más", "pero", "sus", "le",
|
||||
"ya", "o", "este", "sí", "porque", "esta", "entre", "cuando", "muy", "sin", "sobre",
|
||||
"también", "me", "hasta", "hay", "donde", "quien", "desde", "todo", "nos", "durante",
|
||||
"todos", "uno", "les", "ni", "contra", "otros", "ese", "eso", "ante", "ellos",
|
||||
"e", "esto", "mí", "antes", "algunos", "qué", "unos", "yo", "otro", "otras",
|
||||
"anunció", "producción", "empresa", "mercado", "sector", "año", "gobierno"
|
||||
"de",
|
||||
"la",
|
||||
"que",
|
||||
"el",
|
||||
"en",
|
||||
"y",
|
||||
"a",
|
||||
"los",
|
||||
"del",
|
||||
"se",
|
||||
"las",
|
||||
"por",
|
||||
"un",
|
||||
"para",
|
||||
"con",
|
||||
"no",
|
||||
"una",
|
||||
"su",
|
||||
"al",
|
||||
"lo",
|
||||
"como",
|
||||
"más",
|
||||
"pero",
|
||||
"sus",
|
||||
"le",
|
||||
"ya",
|
||||
"o",
|
||||
"este",
|
||||
"sí",
|
||||
"porque",
|
||||
"esta",
|
||||
"entre",
|
||||
"cuando",
|
||||
"muy",
|
||||
"sin",
|
||||
"sobre",
|
||||
"también",
|
||||
"me",
|
||||
"hasta",
|
||||
"hay",
|
||||
"donde",
|
||||
"quien",
|
||||
"desde",
|
||||
"todo",
|
||||
"nos",
|
||||
"durante",
|
||||
"todos",
|
||||
"uno",
|
||||
"les",
|
||||
"ni",
|
||||
"contra",
|
||||
"otros",
|
||||
"ese",
|
||||
"eso",
|
||||
"ante",
|
||||
"ellos",
|
||||
"e",
|
||||
"esto",
|
||||
"mí",
|
||||
"antes",
|
||||
"algunos",
|
||||
"qué",
|
||||
"unos",
|
||||
"yo",
|
||||
"otro",
|
||||
"otras",
|
||||
"anunció",
|
||||
"producción",
|
||||
"empresa",
|
||||
"mercado",
|
||||
"sector",
|
||||
"año",
|
||||
"gobierno",
|
||||
},
|
||||
"de": {
|
||||
"der", "die", "und", "in", "den", "von", "zu", "das", "mit", "sich", "des", "auf",
|
||||
"für", "ist", "im", "dem", "nicht", "ein", "eine", "als", "auch", "es", "an",
|
||||
"werden", "aus", "er", "hat", "dass", "sie", "nach", "wird", "bei", "einer", "um",
|
||||
"am", "sind", "noch", "wie", "einem", "über", "einen", "so", "zum", "war", "haben",
|
||||
"nur", "oder", "aber", "vor", "zur", "bis", "mehr", "durch", "man", "sein", "wurde",
|
||||
"sei", "prozent", "hatte", "kann", "gegen", "vom", "können", "schon", "wenn", "habe",
|
||||
"seine", "ihre", "unter", "wir", "sollen", "neue", "neuen", "batteriezellen", "unternehmen"
|
||||
"der",
|
||||
"die",
|
||||
"und",
|
||||
"in",
|
||||
"den",
|
||||
"von",
|
||||
"zu",
|
||||
"das",
|
||||
"mit",
|
||||
"sich",
|
||||
"des",
|
||||
"auf",
|
||||
"für",
|
||||
"ist",
|
||||
"im",
|
||||
"dem",
|
||||
"nicht",
|
||||
"ein",
|
||||
"eine",
|
||||
"als",
|
||||
"auch",
|
||||
"es",
|
||||
"an",
|
||||
"werden",
|
||||
"aus",
|
||||
"er",
|
||||
"hat",
|
||||
"dass",
|
||||
"sie",
|
||||
"nach",
|
||||
"wird",
|
||||
"bei",
|
||||
"einer",
|
||||
"um",
|
||||
"am",
|
||||
"sind",
|
||||
"noch",
|
||||
"wie",
|
||||
"einem",
|
||||
"über",
|
||||
"einen",
|
||||
"so",
|
||||
"zum",
|
||||
"war",
|
||||
"haben",
|
||||
"nur",
|
||||
"oder",
|
||||
"aber",
|
||||
"vor",
|
||||
"zur",
|
||||
"bis",
|
||||
"mehr",
|
||||
"durch",
|
||||
"man",
|
||||
"sein",
|
||||
"wurde",
|
||||
"sei",
|
||||
"prozent",
|
||||
"hatte",
|
||||
"kann",
|
||||
"gegen",
|
||||
"vom",
|
||||
"können",
|
||||
"schon",
|
||||
"wenn",
|
||||
"habe",
|
||||
"seine",
|
||||
"ihre",
|
||||
"unter",
|
||||
"wir",
|
||||
"sollen",
|
||||
"neue",
|
||||
"neuen",
|
||||
"batteriezellen",
|
||||
"unternehmen",
|
||||
},
|
||||
"it": {
|
||||
"di", "e", "il", "che", "la", "a", "in", "per", "un", "del", "non", "i", "si", "da",
|
||||
"le", "con", "sono", "della", "dei", "degli", "una", "al", "ma", "più", "delle",
|
||||
"questo", "nel", "alla", "anche", "ha", "gli", "come", "dall", "dalla", "ed",
|
||||
"se", "ci", "lo", "su", "loro", "dopo", "qualche", "nella", "uno", "mio", "tuo",
|
||||
"suo", "nostro", "vostro", "loro", "stato", "stata", "tra", "fra", "mentre", "prima",
|
||||
"quando", "molto", "tutto", "tutti", "tutte", "tutta", "senza", "ancora", "solo",
|
||||
"azienda", "mercato", "settore", "anno", "governo", "produzione", "motori"
|
||||
"di",
|
||||
"e",
|
||||
"il",
|
||||
"che",
|
||||
"la",
|
||||
"a",
|
||||
"in",
|
||||
"per",
|
||||
"un",
|
||||
"del",
|
||||
"non",
|
||||
"i",
|
||||
"si",
|
||||
"da",
|
||||
"le",
|
||||
"con",
|
||||
"sono",
|
||||
"della",
|
||||
"dei",
|
||||
"degli",
|
||||
"una",
|
||||
"al",
|
||||
"ma",
|
||||
"più",
|
||||
"delle",
|
||||
"questo",
|
||||
"nel",
|
||||
"alla",
|
||||
"anche",
|
||||
"ha",
|
||||
"gli",
|
||||
"come",
|
||||
"dall",
|
||||
"dalla",
|
||||
"ed",
|
||||
"se",
|
||||
"ci",
|
||||
"lo",
|
||||
"su",
|
||||
"loro",
|
||||
"dopo",
|
||||
"qualche",
|
||||
"nella",
|
||||
"uno",
|
||||
"mio",
|
||||
"tuo",
|
||||
"suo",
|
||||
"nostro",
|
||||
"vostro",
|
||||
"stato",
|
||||
"stata",
|
||||
"tra",
|
||||
"fra",
|
||||
"mentre",
|
||||
"prima",
|
||||
"quando",
|
||||
"molto",
|
||||
"tutto",
|
||||
"tutti",
|
||||
"tutte",
|
||||
"tutta",
|
||||
"senza",
|
||||
"ancora",
|
||||
"solo",
|
||||
"azienda",
|
||||
"mercato",
|
||||
"settore",
|
||||
"anno",
|
||||
"governo",
|
||||
"produzione",
|
||||
"motori",
|
||||
},
|
||||
"fr": {
|
||||
"de", "la", "le", "et", "les", "des", "en", "un", "du", "une", "que", "est", "pour",
|
||||
"qui", "dans", "a", "par", "sur", "pas", "plus", "au", "avec", "ce", "il", "sont",
|
||||
"se", "ne", "son", "sa", "ses", "aux", "ou", "comme", "mais", "nous", "vous", "ils",
|
||||
"leur", "y", "tout", "faire", "été", "aussi", "ces", "ont", "si", "fait", "même",
|
||||
"très", "après", "sans", "sous", "entre", "deux", "bien", "chez", "autre", "autres",
|
||||
"entreprise", "marché", "secteur", "année", "gouvernement", "production", "véhicules"
|
||||
}
|
||||
"de",
|
||||
"la",
|
||||
"le",
|
||||
"et",
|
||||
"les",
|
||||
"des",
|
||||
"en",
|
||||
"un",
|
||||
"du",
|
||||
"une",
|
||||
"que",
|
||||
"est",
|
||||
"pour",
|
||||
"qui",
|
||||
"dans",
|
||||
"a",
|
||||
"par",
|
||||
"sur",
|
||||
"pas",
|
||||
"plus",
|
||||
"au",
|
||||
"avec",
|
||||
"ce",
|
||||
"il",
|
||||
"sont",
|
||||
"se",
|
||||
"ne",
|
||||
"son",
|
||||
"sa",
|
||||
"ses",
|
||||
"aux",
|
||||
"ou",
|
||||
"comme",
|
||||
"mais",
|
||||
"nous",
|
||||
"vous",
|
||||
"ils",
|
||||
"leur",
|
||||
"y",
|
||||
"tout",
|
||||
"faire",
|
||||
"été",
|
||||
"aussi",
|
||||
"ces",
|
||||
"ont",
|
||||
"si",
|
||||
"fait",
|
||||
"même",
|
||||
"très",
|
||||
"après",
|
||||
"sans",
|
||||
"sous",
|
||||
"entre",
|
||||
"deux",
|
||||
"bien",
|
||||
"chez",
|
||||
"autre",
|
||||
"autres",
|
||||
"entreprise",
|
||||
"marché",
|
||||
"secteur",
|
||||
"année",
|
||||
"gouvernement",
|
||||
"production",
|
||||
"véhicules",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@@ -77,12 +469,12 @@ def normalize_text(text: str) -> str:
|
||||
return "".join(c for c in nfd if unicodedata.category(c) != "Mn")
|
||||
|
||||
|
||||
def extract_words(text: str) -> List[str]:
|
||||
def extract_words(text: str) -> list[str]:
|
||||
"""Tokenize text into lowercase alphanumeric words."""
|
||||
return re.findall(r"\b\w+\b", text.lower())
|
||||
|
||||
|
||||
def detect_language(text: str) -> Tuple[str, float]:
|
||||
def detect_language(text: str) -> tuple[str, float]:
|
||||
"""
|
||||
Detect the ISO-639-1 language code of text among supported languages (pt, en, es, de, it, fr).
|
||||
Returns (detected_language, confidence_score).
|
||||
@@ -94,11 +486,10 @@ def detect_language(text: str) -> Tuple[str, float]:
|
||||
if not words:
|
||||
return "unknown", 0.0
|
||||
|
||||
total_words = len(words)
|
||||
word_set = set(words)
|
||||
|
||||
# Score languages based on matched stopword counts
|
||||
scores: Dict[str, int] = {}
|
||||
scores: dict[str, int] = {}
|
||||
for lang, stopwords in LANGUAGE_STOPWORDS.items():
|
||||
matched = word_set.intersection(stopwords)
|
||||
scores[lang] = len(matched)
|
||||
@@ -110,13 +501,55 @@ def detect_language(text: str) -> Tuple[str, float]:
|
||||
|
||||
# Specific disambiguation rules for closely related languages (PT vs ES)
|
||||
pt_exclusive = {
|
||||
"não", "do", "da", "no", "na", "nos", "nas", "em", "um", "uma", "você", "são", "é", "dos",
|
||||
"das", "foi", "está", "estão", "com", "pelo", "pela", "pelos", "pelas", "notícia",
|
||||
"extração", "mês", "ano", "produção", "bateu"
|
||||
"não",
|
||||
"do",
|
||||
"da",
|
||||
"no",
|
||||
"na",
|
||||
"nos",
|
||||
"nas",
|
||||
"em",
|
||||
"um",
|
||||
"uma",
|
||||
"você",
|
||||
"são",
|
||||
"é",
|
||||
"dos",
|
||||
"das",
|
||||
"foi",
|
||||
"está",
|
||||
"estão",
|
||||
"com",
|
||||
"pelo",
|
||||
"pela",
|
||||
"pelos",
|
||||
"pelas",
|
||||
"notícia",
|
||||
"extração",
|
||||
"mês",
|
||||
"ano",
|
||||
"produção",
|
||||
"bateu",
|
||||
}
|
||||
es_exclusive = {
|
||||
"el", "la", "y", "del", "al", "los", "las", "su", "sus", "con", "más", "pero",
|
||||
"durante", "noticia", "extracción", "mes", "año", "producción"
|
||||
"el",
|
||||
"la",
|
||||
"y",
|
||||
"del",
|
||||
"al",
|
||||
"los",
|
||||
"las",
|
||||
"su",
|
||||
"sus",
|
||||
"con",
|
||||
"más",
|
||||
"pero",
|
||||
"durante",
|
||||
"noticia",
|
||||
"extracción",
|
||||
"mes",
|
||||
"año",
|
||||
"producción",
|
||||
}
|
||||
|
||||
# Normalize words to match accents cleanly
|
||||
@@ -133,7 +566,11 @@ def detect_language(text: str) -> Tuple[str, float]:
|
||||
elif "e" in word_set and "y" not in word_set:
|
||||
pt_score += 1
|
||||
|
||||
if top_lang in ("pt", "es") or (top_lang == "pt" and es_score > pt_score) or (top_lang == "es" and pt_score > es_score):
|
||||
if (
|
||||
top_lang in ("pt", "es")
|
||||
or (top_lang == "pt" and es_score > pt_score)
|
||||
or (top_lang == "es" and pt_score > es_score)
|
||||
):
|
||||
if es_score > pt_score:
|
||||
top_lang = "es"
|
||||
top_matches = max(top_matches, es_score)
|
||||
|
||||
+23
-19
@@ -3,9 +3,9 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass, field, asdict
|
||||
from dataclasses import dataclass, field
|
||||
from enum import Enum
|
||||
from typing import Any, Dict, List, Optional
|
||||
from typing import Any
|
||||
|
||||
|
||||
class DecisionCategory(str, Enum):
|
||||
@@ -29,12 +29,12 @@ class RelatedEntity:
|
||||
name: str
|
||||
relation_type: str
|
||||
weight: float
|
||||
aliases: List[str] = field(default_factory=list)
|
||||
aliases: list[str] = field(default_factory=list)
|
||||
scope: str = "general"
|
||||
confidence: float = 1.0
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, data: Dict[str, Any]) -> RelatedEntity:
|
||||
def from_dict(cls, data: dict[str, Any]) -> RelatedEntity:
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError("Related entity must be a JSON object")
|
||||
|
||||
@@ -58,15 +58,15 @@ class RelatedEntity:
|
||||
class ECPSnapshot:
|
||||
target_entity_id: str
|
||||
target_name: str
|
||||
aliases: List[str]
|
||||
aliases: list[str]
|
||||
domain: str
|
||||
anchors: List[str]
|
||||
negative_anchors: List[str] = field(default_factory=list)
|
||||
anchors: list[str]
|
||||
negative_anchors: list[str] = field(default_factory=list)
|
||||
graph_version: str = "1.0.0"
|
||||
related_entities: List[RelatedEntity] = field(default_factory=list)
|
||||
related_entities: list[RelatedEntity] = field(default_factory=list)
|
||||
|
||||
@classmethod
|
||||
def from_dict(cls, data: Dict[str, Any]) -> ECPSnapshot:
|
||||
def from_dict(cls, data: dict[str, Any]) -> ECPSnapshot:
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError("ECP Snapshot payload must be a JSON object")
|
||||
|
||||
@@ -124,16 +124,18 @@ class ClassificationResult:
|
||||
is_inherent: bool
|
||||
confidence: float
|
||||
detected_language: str
|
||||
matched_anchors: List[str] = field(default_factory=list)
|
||||
negative_matches: List[str] = field(default_factory=list)
|
||||
graph_matches: List[Dict[str, Any]] = field(default_factory=list)
|
||||
evidence: List[str] = field(default_factory=list)
|
||||
matched_anchors: list[str] = field(default_factory=list)
|
||||
negative_matches: list[str] = field(default_factory=list)
|
||||
graph_matches: list[dict[str, Any]] = field(default_factory=list)
|
||||
evidence: list[str] = field(default_factory=list)
|
||||
rationale: str = ""
|
||||
warnings: List[str] = field(default_factory=list)
|
||||
warnings: list[str] = field(default_factory=list)
|
||||
|
||||
def to_dict(self) -> Dict[str, Any]:
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"decision": self.decision.value if isinstance(self.decision, DecisionCategory) else str(self.decision),
|
||||
"decision": self.decision.value
|
||||
if isinstance(self.decision, DecisionCategory)
|
||||
else str(self.decision),
|
||||
"is_inherent": bool(self.is_inherent),
|
||||
"confidence": round(float(self.confidence), 4),
|
||||
"detected_language": str(self.detected_language),
|
||||
@@ -153,11 +155,13 @@ class ClassificationResult:
|
||||
class ClassificationError:
|
||||
error_code: ErrorCode
|
||||
message: str
|
||||
details: Dict[str, Any] = field(default_factory=dict)
|
||||
details: dict[str, Any] = field(default_factory=dict)
|
||||
|
||||
def to_dict(self) -> Dict[str, Any]:
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"error_code": self.error_code.value if isinstance(self.error_code, ErrorCode) else str(self.error_code),
|
||||
"error_code": self.error_code.value
|
||||
if isinstance(self.error_code, ErrorCode)
|
||||
else str(self.error_code),
|
||||
"message": str(self.message),
|
||||
"details": dict(self.details),
|
||||
}
|
||||
|
||||
+5
-4
@@ -3,7 +3,6 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from typing import List, Tuple
|
||||
|
||||
|
||||
def strip_markdown(markdown_text: str) -> str:
|
||||
@@ -40,7 +39,7 @@ def strip_markdown(markdown_text: str) -> str:
|
||||
return text
|
||||
|
||||
|
||||
def extract_sentences(text: str) -> List[str]:
|
||||
def extract_sentences(text: str) -> list[str]:
|
||||
"""Split text into individual sentences."""
|
||||
# Split by period, exclamation, question mark followed by space or newline
|
||||
raw_sentences = re.split(r"(?<=[.!?])\s+", text.strip())
|
||||
@@ -48,7 +47,9 @@ def extract_sentences(text: str) -> List[str]:
|
||||
return sentences
|
||||
|
||||
|
||||
def extract_evidence_snippets(markdown_text: str, match_terms: List[str], max_snippets: int = 3) -> List[str]:
|
||||
def extract_evidence_snippets(
|
||||
markdown_text: str, match_terms: list[str], max_snippets: int = 3
|
||||
) -> list[str]:
|
||||
"""
|
||||
Extract relevant sentence excerpts from Markdown text that contain any of the given match terms.
|
||||
Preserves original phrasing and formats as clean evidence.
|
||||
@@ -62,7 +63,7 @@ def extract_evidence_snippets(markdown_text: str, match_terms: List[str], max_sn
|
||||
sentences = [plain_text]
|
||||
|
||||
lower_terms = [t.lower() for t in match_terms if t]
|
||||
evidence: List[str] = []
|
||||
evidence: list[str] = []
|
||||
|
||||
for sentence in sentences:
|
||||
lower_sent = sentence.lower()
|
||||
|
||||
@@ -27,7 +27,7 @@ def test_classifier_with_adapter_flags():
|
||||
target_name="TestCorp",
|
||||
aliases=["TestCorp"],
|
||||
domain="Tech",
|
||||
anchors=["software"]
|
||||
anchors=["software"],
|
||||
)
|
||||
res = classifier.classify(ecp, "TestCorp builds enterprise cloud software.")
|
||||
assert res.decision.value == "DIRECT_INHERENT"
|
||||
|
||||
+57
-29
@@ -7,9 +7,8 @@ and CLI execution behavior via subprocess (exit codes, stream purity, JSON parsi
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
import pytest
|
||||
from src.models import ECPSnapshot, RelatedEntity, DecisionCategory
|
||||
|
||||
from src.models import DecisionCategory, ECPSnapshot, RelatedEntity
|
||||
|
||||
|
||||
def test_adversarial_sao_paulo_city_vs_fc():
|
||||
@@ -20,9 +19,13 @@ def test_adversarial_sao_paulo_city_vs_fc():
|
||||
aliases=["São Paulo", "SPFC", "Tricolor Paulista"],
|
||||
domain="Futebol e Esportes",
|
||||
anchors=["Morumbi", "futebol", "campeonato", "Copa Libertadores", "elenco", "estádio"],
|
||||
negative_anchors=["prefeitura de são paulo", "governo do estado de são paulo", "trânsito na capital paulista"],
|
||||
negative_anchors=[
|
||||
"prefeitura de são paulo",
|
||||
"governo do estado de são paulo",
|
||||
"trânsito na capital paulista",
|
||||
],
|
||||
graph_version="1.0.0",
|
||||
related_entities=[]
|
||||
related_entities=[],
|
||||
)
|
||||
content = (
|
||||
"# Obras Viárias na Capital\n\n"
|
||||
@@ -30,6 +33,7 @@ def test_adversarial_sao_paulo_city_vs_fc():
|
||||
"para desafogar o fluxo de veículos na região central durante os horários de pico."
|
||||
)
|
||||
from src.classifier import InherenceClassifier
|
||||
|
||||
classifier = InherenceClassifier()
|
||||
result = classifier.classify(ecp, content)
|
||||
assert result.decision in (DecisionCategory.NOT_RELATED, DecisionCategory.TANGENTIAL)
|
||||
@@ -47,13 +51,14 @@ def test_adversarial_apple_fruit_recipe():
|
||||
anchors=["iPhone", "MacBook", "iOS", "silicon", "hardware"],
|
||||
negative_anchors=["apple pie", "orchard harvest", "doce de maçã"],
|
||||
graph_version="1.0.0",
|
||||
related_entities=[]
|
||||
related_entities=[],
|
||||
)
|
||||
content = (
|
||||
"# Receita Caseira\n\n"
|
||||
"Comprei maçãs frescas no mercado para preparar um doce de maçã com canela e açúcar mascavo."
|
||||
)
|
||||
from src.classifier import InherenceClassifier
|
||||
|
||||
classifier = InherenceClassifier()
|
||||
result = classifier.classify(ecp, content)
|
||||
assert result.decision in (DecisionCategory.NOT_RELATED, DecisionCategory.TANGENTIAL)
|
||||
@@ -76,9 +81,9 @@ def test_adversarial_related_entity_without_scope_context():
|
||||
relation_type="SUPPLIER_OF",
|
||||
weight=0.95,
|
||||
scope="battery_technology",
|
||||
confidence=0.99
|
||||
confidence=0.99,
|
||||
)
|
||||
]
|
||||
],
|
||||
)
|
||||
# Content mentions Northvolt in an unrelated/passing architectural context without domain anchors
|
||||
content = (
|
||||
@@ -87,6 +92,7 @@ def test_adversarial_related_entity_without_scope_context():
|
||||
"mit moderner Holzfassade und Blick auf den See."
|
||||
)
|
||||
from src.classifier import InherenceClassifier
|
||||
|
||||
classifier = InherenceClassifier()
|
||||
result = classifier.classify(ecp, content)
|
||||
# Must be TANGENTIAL or NOT_RELATED, NEVER CONTEXTUAL_INHERENT
|
||||
@@ -98,18 +104,27 @@ def test_adversarial_related_entity_without_scope_context():
|
||||
def test_adversarial_subprocess_cli_success_stdout(tmp_path):
|
||||
"""Run CLI via subprocess without --output and verify stdout is pure parseable JSON."""
|
||||
ecp_file = tmp_path / "ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_petrobras",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_petrobras",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "content.md"
|
||||
content_file.write_text("# Notícia\n\nA Petrobras bateu recorde de extração de petróleo no pré-sal este mês.", encoding="utf-8")
|
||||
content_file.write_text(
|
||||
"# Notícia\n\nA Petrobras bateu recorde de extração de petróleo no pré-sal este mês.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
import os
|
||||
|
||||
env = dict(os.environ, PYTHONIOENCODING="utf-8", PYTHONUTF8="1")
|
||||
|
||||
res = subprocess.run(
|
||||
@@ -133,16 +148,22 @@ def test_adversarial_subprocess_cli_success_stdout(tmp_path):
|
||||
def test_adversarial_subprocess_cli_empty_content(tmp_path):
|
||||
"""Run CLI via subprocess with empty content and verify error code and exit code."""
|
||||
import os
|
||||
|
||||
env = dict(os.environ, PYTHONIOENCODING="utf-8", PYTHONUTF8="1")
|
||||
|
||||
ecp_file = tmp_path / "ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Test",
|
||||
"aliases": ["Test"],
|
||||
"domain": "Tech",
|
||||
"anchors": ["tech"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Test",
|
||||
"aliases": ["Test"],
|
||||
"domain": "Tech",
|
||||
"anchors": ["tech"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "empty.md"
|
||||
content_file.write_text(" \n\n ", encoding="utf-8")
|
||||
@@ -164,15 +185,21 @@ def test_adversarial_subprocess_cli_empty_content(tmp_path):
|
||||
def test_adversarial_subprocess_cli_missing_required_field(tmp_path):
|
||||
"""Run CLI via subprocess with missing target_name and verify error payload."""
|
||||
import os
|
||||
|
||||
env = dict(os.environ, PYTHONIOENCODING="utf-8", PYTHONUTF8="1")
|
||||
|
||||
ecp_file = tmp_path / "ecp_bad.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"aliases": ["Test"],
|
||||
"domain": "Tech",
|
||||
"anchors": ["tech"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_1",
|
||||
"aliases": ["Test"],
|
||||
"domain": "Tech",
|
||||
"anchors": ["tech"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "content.md"
|
||||
content_file.write_text("Conteúdo de teste válido.", encoding="utf-8")
|
||||
@@ -193,6 +220,7 @@ def test_adversarial_subprocess_cli_missing_required_field(tmp_path):
|
||||
def test_adversarial_subprocess_cli_corrupted_json(tmp_path):
|
||||
"""Run CLI via subprocess with corrupted JSON and verify error payload."""
|
||||
import os
|
||||
|
||||
env = dict(os.environ, PYTHONIOENCODING="utf-8", PYTHONUTF8="1")
|
||||
|
||||
ecp_file = tmp_path / "ecp_corrupted.json"
|
||||
|
||||
+19
-21
@@ -6,19 +6,17 @@ Target Success Criterion: Precision >= 90% over the 24 cases.
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from src.models import ECPSnapshot, DecisionCategory
|
||||
|
||||
from src.classifier import InherenceClassifier
|
||||
from src.models import ECPSnapshot
|
||||
|
||||
FIXTURES_DIR = Path(__file__).parent / "fixtures" / "benchmark_24"
|
||||
LANGUAGES = ["pt", "en", "es", "de", "it", "fr"]
|
||||
DECISION_TYPES = ["direct", "contextual", "tangential", "not_related"]
|
||||
|
||||
BENCHMARK_CASES = [
|
||||
(lang, dec_type)
|
||||
for lang in LANGUAGES
|
||||
for dec_type in DECISION_TYPES
|
||||
]
|
||||
BENCHMARK_CASES = [(lang, dec_type) for lang in LANGUAGES for dec_type in DECISION_TYPES]
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
@@ -44,28 +42,28 @@ def test_benchmark_case(classifier, lang: str, dec_type: str):
|
||||
result = classifier.classify(ecp, content)
|
||||
|
||||
# 1. Decision category validation
|
||||
assert (
|
||||
result.decision.value == expected["expected_decision"]
|
||||
), f"[{lang.upper()} - {dec_type}] Expected {expected['expected_decision']}, got {result.decision.value}. Rationale: {result.rationale}"
|
||||
assert result.decision.value == expected["expected_decision"], (
|
||||
f"[{lang.upper()} - {dec_type}] Expected {expected['expected_decision']}, got {result.decision.value}. Rationale: {result.rationale}"
|
||||
)
|
||||
|
||||
# 2. Derived is_inherent boolean validation
|
||||
assert (
|
||||
result.is_inherent == expected["expected_is_inherent"]
|
||||
), f"[{lang.upper()} - {dec_type}] Expected is_inherent={expected['expected_is_inherent']}, got {result.is_inherent}"
|
||||
assert result.is_inherent == expected["expected_is_inherent"], (
|
||||
f"[{lang.upper()} - {dec_type}] Expected is_inherent={expected['expected_is_inherent']}, got {result.is_inherent}"
|
||||
)
|
||||
|
||||
# 3. Language detection validation
|
||||
assert (
|
||||
result.detected_language == expected["expected_language"]
|
||||
), f"[{lang.upper()} - {dec_type}] Expected language '{expected['expected_language']}', got '{result.detected_language}'"
|
||||
assert result.detected_language == expected["expected_language"], (
|
||||
f"[{lang.upper()} - {dec_type}] Expected language '{expected['expected_language']}', got '{result.detected_language}'"
|
||||
)
|
||||
|
||||
# 4. Confidence threshold validation
|
||||
min_conf = expected.get("min_confidence", 0.0)
|
||||
assert (
|
||||
result.confidence >= min_conf
|
||||
), f"[{lang.upper()} - {dec_type}] Expected confidence >= {min_conf}, got {result.confidence}"
|
||||
assert result.confidence >= min_conf, (
|
||||
f"[{lang.upper()} - {dec_type}] Expected confidence >= {min_conf}, got {result.confidence}"
|
||||
)
|
||||
|
||||
# 5. Evidence presence for inherent content
|
||||
if result.is_inherent:
|
||||
assert (
|
||||
len(result.evidence) > 0
|
||||
), f"[{lang.upper()} - {dec_type}] Inherent decision must have non-empty evidence snippets"
|
||||
assert len(result.evidence) > 0, (
|
||||
f"[{lang.upper()} - {dec_type}] Inherent decision must have non-empty evidence snippets"
|
||||
)
|
||||
|
||||
@@ -1,8 +1,9 @@
|
||||
"""Unit tests for deterministic classification decision logic."""
|
||||
|
||||
import pytest
|
||||
from src.models import ECPSnapshot, RelatedEntity, DecisionCategory
|
||||
|
||||
from src.classifier import InherenceClassifier
|
||||
from src.models import DecisionCategory, ECPSnapshot, RelatedEntity
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
@@ -23,9 +24,9 @@ def petrobras_ecp():
|
||||
weight=0.85,
|
||||
aliases=[],
|
||||
scope="logistics",
|
||||
confidence=1.0
|
||||
confidence=1.0,
|
||||
)
|
||||
]
|
||||
],
|
||||
)
|
||||
|
||||
|
||||
|
||||
+52
-31
@@ -1,23 +1,30 @@
|
||||
"""CLI execution tests covering flags, arguments, stdout, and error handling."""
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
import pytest
|
||||
|
||||
from classify import main
|
||||
|
||||
|
||||
def test_cli_success_stdout(tmp_path, capsys):
|
||||
ecp_file = tmp_path / "ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "content.md"
|
||||
content_file.write_text("# Notícia\n\nA Petrobras bateu recorde de extração de petróleo no pré-sal este mês.", encoding="utf-8")
|
||||
content_file.write_text(
|
||||
"# Notícia\n\nA Petrobras bateu recorde de extração de petróleo no pré-sal este mês.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
exit_code = main(["--ecp", str(ecp_file), "--content", str(content_file)])
|
||||
assert exit_code == 0
|
||||
@@ -31,20 +38,30 @@ def test_cli_success_stdout(tmp_path, capsys):
|
||||
|
||||
def test_cli_output_file(tmp_path):
|
||||
ecp_file = tmp_path / "ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo", "pré-sal"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "content.md"
|
||||
content_file.write_text("Petrobras anunciou investimentos bilionários em novas refinarias de petróleo.", encoding="utf-8")
|
||||
content_file.write_text(
|
||||
"Petrobras anunciou investimentos bilionários em novas refinarias de petróleo.",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
output_file = tmp_path / "out" / "result.json"
|
||||
|
||||
exit_code = main(["--ecp", str(ecp_file), "--content", str(content_file), "--output", str(output_file)])
|
||||
exit_code = main(
|
||||
["--ecp", str(ecp_file), "--content", str(content_file), "--output", str(output_file)]
|
||||
)
|
||||
assert exit_code == 0
|
||||
assert output_file.exists()
|
||||
|
||||
@@ -67,13 +84,18 @@ def test_cli_missing_ecp_file(tmp_path, capsys):
|
||||
|
||||
def test_cli_empty_content_file(tmp_path, capsys):
|
||||
ecp_file = tmp_path / "ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"target_entity_id": "ent_1",
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo"],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "empty.md"
|
||||
content_file.write_text(" \n\n ", encoding="utf-8")
|
||||
@@ -88,11 +110,10 @@ def test_cli_empty_content_file(tmp_path, capsys):
|
||||
|
||||
def test_cli_missing_required_ecp_field(tmp_path, capsys):
|
||||
ecp_file = tmp_path / "bad_ecp.json"
|
||||
ecp_file.write_text(json.dumps({
|
||||
"target_entity_id": "ent_1",
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["petróleo"]
|
||||
}), encoding="utf-8")
|
||||
ecp_file.write_text(
|
||||
json.dumps({"target_entity_id": "ent_1", "domain": "Oil & Gas", "anchors": ["petróleo"]}),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
content_file = tmp_path / "content.md"
|
||||
content_file.write_text("Algum conteúdo válido para testar o erro.", encoding="utf-8")
|
||||
|
||||
@@ -0,0 +1,372 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Testes automatizados para o Extrator e Parser Multimotor de Artigos.
|
||||
Cobre modelos de dados, parsers (Trafilatura, Newspaper4k, Readability),
|
||||
isolamento de falhas, orquestração de lote e interface CLI.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from scripts.extract_article_contents import (
|
||||
ArticleCrawler,
|
||||
ExtractedArticle,
|
||||
ExtractionBatchReport,
|
||||
InputArticle,
|
||||
NewspaperData,
|
||||
NewspaperExtractor,
|
||||
ReadabilityData,
|
||||
ReadabilityExtractor,
|
||||
TrafilaturaData,
|
||||
TrafilaturaExtractor,
|
||||
extract_all_engines,
|
||||
load_search_json,
|
||||
main,
|
||||
process_batch,
|
||||
save_extracted_json,
|
||||
)
|
||||
|
||||
SAMPLE_HTML = """
|
||||
<!DOCTYPE html>
|
||||
<html lang="es">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<title>River Plate igualó sin goles ante Independiente Santa Fe - Olé</title>
|
||||
<meta name="description" content="El equipo de Núñez empató 0-0 en Bogotá por los octavos de final.">
|
||||
<meta name="author" content="Juan Pérez">
|
||||
<meta property="og:title" content="River Plate igualó sin goles ante Independiente Santa Fe">
|
||||
<meta property="og:image" content="https://media.ole.com.ar/river.jpg">
|
||||
</head>
|
||||
<body>
|
||||
<header><nav><a href="/">Inicio</a></nav></header>
|
||||
<article>
|
||||
<h1>River Plate igualó sin goles ante Independiente Santa Fe</h1>
|
||||
<p class="byline">Por Juan Pérez - 20 de Agosto de 2026</p>
|
||||
<p class="lead">El equipo de Núñez empató 0-0 en Bogotá por la Copa Sudamericana.</p>
|
||||
<p>Franco Armani fue la gran figura del encuentro con tres atajadas espectaculares en el primer tiempo.</p>
|
||||
<p>El partido de vuelta se disputará en el estadio Monumental la próxima semana ante una multitud.</p>
|
||||
</article>
|
||||
<footer><p>Copyright 2026 Olé</p></footer>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_input_json(tmp_path: Path) -> Path:
|
||||
data = {
|
||||
"query": "River Plate",
|
||||
"language": "es",
|
||||
"locale": "AR",
|
||||
"total_itens": 2,
|
||||
"items": [
|
||||
{
|
||||
"titulo": "River Plate igualó sin goles ante Santa Fe",
|
||||
"subtitulo": "Empate en Bogotá",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 03:27:26 GMT",
|
||||
"url": "https://www.ole.com.ar/river-0-0-santa-fe.html",
|
||||
"pagina": 1,
|
||||
},
|
||||
{
|
||||
"titulo": "Armani fue la figura de River",
|
||||
"subtitulo": "Gran actuación del arquero",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 04:00:00 GMT",
|
||||
"url": "https://www.tycsports.com/armani-figura.html",
|
||||
"pagina": 1,
|
||||
},
|
||||
],
|
||||
}
|
||||
input_file = tmp_path / "river_plate.json"
|
||||
input_file.write_text(json.dumps(data, ensure_ascii=False), encoding="utf-8")
|
||||
return input_file
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 1. Testes de Modelos e I/O de JSON
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_input_article_creation():
|
||||
article = InputArticle(
|
||||
titulo="Notícia Teste",
|
||||
url="https://example.com/noticia",
|
||||
subtitulo="Subtítulo",
|
||||
quando_publicado="Thu, 20 Aug 2026",
|
||||
pagina=1,
|
||||
)
|
||||
assert article.titulo == "Notícia Teste"
|
||||
assert article.url == "https://example.com/noticia"
|
||||
assert article.pagina == 1
|
||||
d = article.to_dict()
|
||||
assert d["titulo"] == "Notícia Teste"
|
||||
assert d["url"] == "https://example.com/noticia"
|
||||
|
||||
|
||||
def test_load_search_json_valid(sample_input_json: Path):
|
||||
query, lang, items = load_search_json(sample_input_json)
|
||||
assert query == "River Plate"
|
||||
assert lang == "es"
|
||||
assert len(items) == 2
|
||||
assert items[0].url == "https://www.ole.com.ar/river-0-0-santa-fe.html"
|
||||
|
||||
|
||||
def test_load_search_json_invalid_file(tmp_path: Path):
|
||||
non_existent = tmp_path / "missing.json"
|
||||
with pytest.raises(FileNotFoundError):
|
||||
load_search_json(non_existent)
|
||||
|
||||
|
||||
def test_save_extracted_json(tmp_path: Path):
|
||||
report = ExtractionBatchReport(
|
||||
source_file="test.json",
|
||||
processed_at="2026-08-20T12:00:00Z",
|
||||
total_articles=1,
|
||||
successful_articles=1,
|
||||
failed_articles=0,
|
||||
articles=[
|
||||
ExtractedArticle(
|
||||
input_meta=InputArticle(titulo="Teste", url="https://example.com/noticia"),
|
||||
extraction_status="success",
|
||||
error_message=None,
|
||||
crawled_url="https://example.com/noticia",
|
||||
page_title="Página Teste",
|
||||
http_status=200,
|
||||
trafilatura=TrafilaturaData(
|
||||
title="Teste",
|
||||
author="Autor",
|
||||
date="2026-08-20",
|
||||
description="Desc",
|
||||
categories=[],
|
||||
tags=[],
|
||||
canonical_url=None,
|
||||
text="Texto do teste.",
|
||||
raw_json=None,
|
||||
error=None,
|
||||
),
|
||||
newspaper4k=NewspaperData(
|
||||
title="Teste",
|
||||
authors=["Autor"],
|
||||
publish_date="2026-08-20",
|
||||
text="Texto do teste.",
|
||||
summary="Resumo",
|
||||
keywords=["teste"],
|
||||
top_image=None,
|
||||
images=[],
|
||||
meta_data={},
|
||||
error=None,
|
||||
),
|
||||
readability=ReadabilityData(
|
||||
title="Teste",
|
||||
short_title="Teste",
|
||||
cleaned_html="<p>Texto do teste.</p>",
|
||||
cleaned_text="Texto do teste.",
|
||||
error=None,
|
||||
),
|
||||
)
|
||||
],
|
||||
)
|
||||
out_file = tmp_path / "out" / "result.json"
|
||||
save_extracted_json(report, out_file)
|
||||
assert out_file.exists()
|
||||
content = json.loads(out_file.read_text(encoding="utf-8"))
|
||||
assert content["total_articles"] == 1
|
||||
assert content["articles"][0]["trafilatura"]["title"] == "Teste"
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 2. Testes Unitários dos Parsers (Trafilatura, Newspaper4k, Readability)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_trafilatura_extractor():
|
||||
extractor = TrafilaturaExtractor()
|
||||
res = extractor.extract(SAMPLE_HTML, url="https://www.ole.com.ar/noticia")
|
||||
assert isinstance(res, TrafilaturaData)
|
||||
assert res.error is None
|
||||
assert "Armani" in res.text or "River Plate" in res.text
|
||||
assert res.title is not None
|
||||
|
||||
|
||||
def test_newspaper_extractor():
|
||||
extractor = NewspaperExtractor()
|
||||
res = extractor.extract(SAMPLE_HTML, url="https://www.ole.com.ar/noticia", language="es")
|
||||
assert isinstance(res, NewspaperData)
|
||||
assert res.error is None
|
||||
assert "Armani" in res.text or "River" in res.text
|
||||
assert isinstance(res.keywords, list)
|
||||
assert len(res.keywords) > 0
|
||||
|
||||
|
||||
def test_readability_extractor():
|
||||
extractor = ReadabilityExtractor()
|
||||
res = extractor.extract(SAMPLE_HTML)
|
||||
assert isinstance(res, ReadabilityData)
|
||||
assert res.error is None
|
||||
assert res.title is not None
|
||||
assert res.cleaned_html is not None
|
||||
assert "Armani" in (res.cleaned_text or "") or "River" in (res.cleaned_text or "")
|
||||
|
||||
|
||||
def test_extract_all_engines():
|
||||
traf, news, read = extract_all_engines(
|
||||
SAMPLE_HTML, url="https://www.ole.com.ar/noticia", language="es"
|
||||
)
|
||||
assert traf.error is None
|
||||
assert news.error is None
|
||||
assert read.error is None
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 3. Testes de Isolamento de Falhas (Resiliência)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_extractor_error_isolation_on_faulty_engine():
|
||||
with patch(
|
||||
"scripts.extract_article_contents.TrafilaturaExtractor.extract",
|
||||
side_effect=RuntimeError("Trafilatura crash"),
|
||||
):
|
||||
traf, news, read = extract_all_engines(
|
||||
SAMPLE_HTML, url="https://example.com", language="es"
|
||||
)
|
||||
assert traf.error == "Trafilatura crash"
|
||||
assert news.error is None
|
||||
assert read.error is None
|
||||
|
||||
|
||||
def test_crawler_error_isolation(sample_input_json: Path, tmp_path: Path):
|
||||
out_file = tmp_path / "river_plate_extracted.json"
|
||||
|
||||
def mock_crawl(url, timeout_sec=30):
|
||||
if "ole.com.ar" in url:
|
||||
return SAMPLE_HTML, "River Plate Olé", 200
|
||||
raise ConnectionError("Connection refused by tycsports.com")
|
||||
|
||||
with patch.object(ArticleCrawler, "crawl", side_effect=mock_crawl):
|
||||
with patch.object(ArticleCrawler, "start"), patch.object(ArticleCrawler, "close"):
|
||||
report = process_batch(
|
||||
input_path=sample_input_json,
|
||||
output_path=out_file,
|
||||
silent=True,
|
||||
)
|
||||
|
||||
assert report.total_articles == 2
|
||||
assert report.successful_articles == 1
|
||||
assert report.failed_articles == 1
|
||||
assert report.articles[0].extraction_status == "success"
|
||||
assert report.articles[1].extraction_status == "failed"
|
||||
assert "Connection refused" in (report.articles[1].error_message or "")
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 4. Testes de CLI e Limitação (--limit, --silent, --language)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_process_batch_with_limit(sample_input_json: Path, tmp_path: Path):
|
||||
out_file = tmp_path / "limit_extracted.json"
|
||||
|
||||
with (
|
||||
patch.object(ArticleCrawler, "crawl", return_value=(SAMPLE_HTML, "Page Title", 200)),
|
||||
patch.object(ArticleCrawler, "start"),
|
||||
patch.object(ArticleCrawler, "close"),
|
||||
):
|
||||
report = process_batch(
|
||||
input_path=sample_input_json,
|
||||
output_path=out_file,
|
||||
limit=1,
|
||||
silent=True,
|
||||
)
|
||||
|
||||
assert report.total_articles == 1
|
||||
assert len(report.articles) == 1
|
||||
assert out_file.exists()
|
||||
|
||||
|
||||
def test_cli_main_success(sample_input_json: Path, tmp_path: Path, capsys):
|
||||
out_file = tmp_path / "cli_out.json"
|
||||
|
||||
with (
|
||||
patch.object(ArticleCrawler, "crawl", return_value=(SAMPLE_HTML, "Page Title", 200)),
|
||||
patch.object(ArticleCrawler, "start"),
|
||||
patch.object(ArticleCrawler, "close"),
|
||||
):
|
||||
exit_code = main(["-i", str(sample_input_json), "-o", str(out_file), "--limit", "1", "-s"])
|
||||
|
||||
assert exit_code == 0
|
||||
assert out_file.exists()
|
||||
|
||||
|
||||
def test_cli_main_missing_input_file(tmp_path: Path, capsys):
|
||||
missing_file = tmp_path / "does_not_exist.json"
|
||||
exit_code = main(["-i", str(missing_file), "-s"])
|
||||
assert exit_code == 1
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 5. Testes End-to-End (E2E) ao Vivo (Live Network)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_e2e_live_article_extraction(tmp_path: Path):
|
||||
"""Valida E2E a extração real ao vivo com Foxcape e os 3 motores em lote."""
|
||||
live_input_file = Path("out/river_plate.json")
|
||||
if not live_input_file.exists():
|
||||
pytest.skip("Arquivo out/river_plate.json não encontrado para teste E2E ao vivo.")
|
||||
|
||||
out_file = tmp_path / "e2e_live_extracted.json"
|
||||
|
||||
# Executa o batch real com limite de 1 notícia
|
||||
report = process_batch(
|
||||
input_path=live_input_file,
|
||||
output_path=out_file,
|
||||
limit=1,
|
||||
silent=True,
|
||||
)
|
||||
|
||||
assert report.total_articles == 1
|
||||
assert report.successful_articles == 1
|
||||
assert report.failed_articles == 0
|
||||
assert len(report.articles) == 1
|
||||
|
||||
art = report.articles[0]
|
||||
assert art.extraction_status == "success"
|
||||
assert art.crawled_url.startswith("http")
|
||||
|
||||
# Valida que todos os 3 motores extraíram dados reais
|
||||
assert art.trafilatura is not None and art.trafilatura.error is None
|
||||
assert len(art.trafilatura.text) > 50
|
||||
|
||||
assert art.newspaper4k is not None and art.newspaper4k.error is None
|
||||
assert len(art.newspaper4k.text) > 50
|
||||
assert isinstance(art.newspaper4k.keywords, list)
|
||||
|
||||
assert art.readability is not None and art.readability.error is None
|
||||
assert art.readability.cleaned_html is not None
|
||||
assert len(art.readability.cleaned_text or "") > 50
|
||||
|
||||
# Valida arquivo JSON gravado
|
||||
assert out_file.exists()
|
||||
saved = json.loads(out_file.read_text(encoding="utf-8"))
|
||||
assert saved["total_articles"] == 1
|
||||
assert saved["articles"][0]["extraction_status"] == "success"
|
||||
|
||||
|
||||
def test_e2e_cli_live_execution(tmp_path: Path):
|
||||
"""Valida E2E a execução do CLI real de ponta a ponta."""
|
||||
live_input_file = Path("out/river_plate.json")
|
||||
if not live_input_file.exists():
|
||||
pytest.skip("Arquivo out/river_plate.json não encontrado para teste E2E.")
|
||||
|
||||
out_file = tmp_path / "e2e_cli_live.json"
|
||||
exit_code = main(["-i", str(live_input_file), "-o", str(out_file), "--limit", "1", "-s"])
|
||||
|
||||
assert exit_code == 0
|
||||
assert out_file.exists()
|
||||
data = json.loads(out_file.read_text(encoding="utf-8"))
|
||||
assert data["successful_articles"] == 1
|
||||
@@ -140,9 +140,7 @@ def test_resolve_article_url_fallback():
|
||||
direct_url = "https://www.globo.com/noticia/123"
|
||||
assert resolve_article_url(direct_url) == direct_url
|
||||
|
||||
with patch(
|
||||
"scripts.extract_google_news.gnewsdecoder", return_value={"status": False}
|
||||
):
|
||||
with patch("scripts.extract_google_news.gnewsdecoder", return_value={"status": False}):
|
||||
gn_url = "https://news.google.com/rss/articles/fake_token"
|
||||
assert resolve_article_url(gn_url) == gn_url
|
||||
|
||||
@@ -187,9 +185,7 @@ def test_resolve_articles_urls_batch():
|
||||
|
||||
def test_extract_google_news_orchestration_mocked(sample_rss_xml: str):
|
||||
"""Valida a consolidação do ExtractionResult a partir da busca mockada com URLs resolvidas."""
|
||||
query = SearchQuery(
|
||||
keyword="inteligência artificial", language="pt", locale="BR", max_pages=1
|
||||
)
|
||||
query = SearchQuery(keyword="inteligência artificial", language="pt", locale="BR", max_pages=1)
|
||||
|
||||
with (
|
||||
patch(
|
||||
@@ -221,13 +217,9 @@ def test_cli_execution_stdout(sample_rss_xml: str, capsys: pytest.CaptureFixture
|
||||
"scripts.extract_google_news._fetch_rss_content",
|
||||
return_value=sample_rss_xml,
|
||||
),
|
||||
patch(
|
||||
"scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u
|
||||
),
|
||||
patch("scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u),
|
||||
):
|
||||
exit_code = main(
|
||||
["--query", "inteligencia artificial", "--lang", "pt", "--pretty"]
|
||||
)
|
||||
exit_code = main(["--query", "inteligencia artificial", "--lang", "pt", "--pretty"])
|
||||
assert exit_code == 0
|
||||
|
||||
captured = capsys.readouterr()
|
||||
@@ -248,9 +240,7 @@ def test_cli_execution_file_output(sample_rss_xml: str, tmp_path: Path):
|
||||
"scripts.extract_google_news._fetch_rss_content",
|
||||
return_value=sample_rss_xml,
|
||||
),
|
||||
patch(
|
||||
"scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u
|
||||
),
|
||||
patch("scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u),
|
||||
):
|
||||
exit_code = main(["-q", "IA", "-p", "1", "-o", str(out_file)])
|
||||
assert exit_code == 0
|
||||
|
||||
@@ -1,11 +1,13 @@
|
||||
"""Unit tests for language detection and text normalization."""
|
||||
|
||||
from src.language import detect_language, normalize_text, SUPPORTED_LANGUAGES
|
||||
from src.language import detect_language, normalize_text
|
||||
|
||||
|
||||
def test_normalize_text():
|
||||
assert normalize_text("São Paulo & Petróleo") == "sao paulo & petroleo"
|
||||
assert normalize_text("Über große Veränderungen") == "uber grosse veranderungen" or "uber" in normalize_text("Über")
|
||||
assert normalize_text(
|
||||
"Über große Veränderungen"
|
||||
) == "uber grosse veranderungen" or "uber" in normalize_text("Über")
|
||||
assert normalize_text("Crème brûlée") == "creme brulee"
|
||||
|
||||
|
||||
@@ -24,7 +26,9 @@ def test_detect_english():
|
||||
|
||||
|
||||
def test_detect_spanish():
|
||||
text = "La empresa petrolera anunció una nueva inversión en el sector energético durante este año."
|
||||
text = (
|
||||
"La empresa petrolera anunció una nueva inversión en el sector energético durante este año."
|
||||
)
|
||||
lang, conf = detect_language(text)
|
||||
assert lang == "es"
|
||||
assert conf > 0.5
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
"""Unit tests for ECP models, schema validation, and structured error handling."""
|
||||
|
||||
import pytest
|
||||
|
||||
from src.models import (
|
||||
ECPSnapshot,
|
||||
RelatedEntity,
|
||||
ClassificationResult,
|
||||
ClassificationError,
|
||||
ClassificationResult,
|
||||
DecisionCategory,
|
||||
ECPSnapshot,
|
||||
ErrorCode,
|
||||
)
|
||||
from src.parser import strip_markdown, extract_sentences, extract_evidence_snippets
|
||||
from src.parser import extract_evidence_snippets, strip_markdown
|
||||
|
||||
|
||||
def test_ecp_snapshot_valid():
|
||||
@@ -29,9 +29,9 @@ def test_ecp_snapshot_valid():
|
||||
"weight": 0.9,
|
||||
"aliases": ["Transpetro Logística"],
|
||||
"scope": "logistics",
|
||||
"confidence": 0.95
|
||||
"confidence": 0.95,
|
||||
}
|
||||
]
|
||||
],
|
||||
}
|
||||
snapshot = ECPSnapshot.from_dict(data)
|
||||
assert snapshot.target_entity_id == "ent_123"
|
||||
@@ -48,7 +48,7 @@ def test_ecp_snapshot_defaults():
|
||||
"target_name": "Petrobras",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["energia"]
|
||||
"anchors": ["energia"],
|
||||
}
|
||||
snapshot = ECPSnapshot.from_dict(data)
|
||||
assert snapshot.negative_anchors == []
|
||||
@@ -61,7 +61,7 @@ def test_ecp_snapshot_missing_required():
|
||||
"target_entity_id": "ent_123",
|
||||
"aliases": ["Petrobras"],
|
||||
"domain": "Oil & Gas",
|
||||
"anchors": ["energia"]
|
||||
"anchors": ["energia"],
|
||||
}
|
||||
with pytest.raises(ValueError, match="Missing required field"):
|
||||
ECPSnapshot.from_dict(data)
|
||||
@@ -89,7 +89,7 @@ def test_classification_error_serialization():
|
||||
err = ClassificationError(
|
||||
error_code=ErrorCode.INVALID_ECP_JSON,
|
||||
message="Malformed JSON syntax",
|
||||
details={"path": "snapshot.json"}
|
||||
details={"path": "snapshot.json"},
|
||||
)
|
||||
d = err.to_dict()
|
||||
assert d["error_code"] == "invalid_ecp_json"
|
||||
|
||||
Reference in New Issue
Block a user