feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
@@ -0,0 +1,261 @@
|
|||||||
|
# 🧠 TextNLPClassifierApp
|
||||||
|
|
||||||
|
[](https://www.python.org/)
|
||||||
|
[](LICENSE)
|
||||||
|
[](https://github.com/astral-sh/ruff)
|
||||||
|
[](https://mypy-lang.org/)
|
||||||
|
[-brightgreen.svg)](tests/)
|
||||||
|
|
||||||
|
> Plataforma modular em Python para **Classificação Multilíngue de Inerência de Entidades (NLP/LLM/ECP)** e **Extração Inteligente de Manchetes de Notícias com Evasão Anti-Bot (Google News RSS & Foxcape)**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📑 Tabela de Conteúdos
|
||||||
|
|
||||||
|
- [Visão Geral](#-visão-geral)
|
||||||
|
- [Instalação e Setup](#-instalação-e-setup)
|
||||||
|
- [1. Classificador de Conteúdo e Inerência (NLP / LLM / ECP)](#1--classificador-de-conteúdo-e-inerência-nlp--llm--ecp)
|
||||||
|
- [O que é e Como Funciona](#o-que-é-e-como-funciona)
|
||||||
|
- [Arquitetura de Classificação em 3 Tiers](#arquitetura-de-classificação-em-3-tiers)
|
||||||
|
- [Categorias de Decisão](#categorias-de-decisão)
|
||||||
|
- [Formato do ECP Snapshot e Markdown](#formato-do-ecp-snapshot-e-markdown)
|
||||||
|
- [Exemplos de Uso CLI](#exemplos-de-uso-cli)
|
||||||
|
- [2. Extrator de Manchetes do Google News](#2--extrator-de-manchetes-do-google-news)
|
||||||
|
- [O que é e Como Funciona](#o-que-é-e-como-funciona-1)
|
||||||
|
- [Diferenciais Técnicos](#diferenciais-técnicos)
|
||||||
|
- [Argumentos e Flags de Linha de Comando](#argumentos-e-flags-de-linha-de-comando)
|
||||||
|
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso)
|
||||||
|
- [Estrutura do Projeto](#-estrutura-do-projeto)
|
||||||
|
- [Testes e Qualidade de Código](#-testes-e-qualidade-de-código)
|
||||||
|
- [Licença](#-licença)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 🌟 Visão Geral
|
||||||
|
|
||||||
|
O **TextNLPClassifierApp** reúne ferramentas de engenharia de dados e processamento de linguagem natural:
|
||||||
|
|
||||||
|
1. **`classify.py`**: Motor de classificação semântica e contextual que determina o grau de aderência e inerência de um documento Markdown em relação a uma entidade alvo definida em um **ECP Snapshot (Entity Context Profile)**.
|
||||||
|
2. **`scripts/extract_google_news.py`**: Extrator de notícias por palavra-chave, idioma e região geográfica que utiliza o motor stealth **Foxcape** (em modo headless), decodificação paralela de URLs para os links reais dos portais de notícias e feedback em tempo real.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## ⚙️ Instalação e Setup
|
||||||
|
|
||||||
|
### 1. Clonar o Repositório e Criar Ambiente Virtual
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone <URL_DO_REPOSITORIO>
|
||||||
|
cd TextNLPClassifierApp
|
||||||
|
|
||||||
|
python -m venv .venv
|
||||||
|
# Windows (PowerShell)
|
||||||
|
.venv\Scripts\Activate.ps1
|
||||||
|
# Linux/macOS
|
||||||
|
source .venv/bin/activate
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Instalar Dependências
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Baixar Binários do Navegador Stealth (Camoufox)
|
||||||
|
|
||||||
|
O motor **Foxcape** utiliza o navegador customizado Camoufox para evasão avançada de fingerprinting:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m camoufox fetch
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP)
|
||||||
|
|
||||||
|
### O que é e Como Funciona
|
||||||
|
|
||||||
|
O classificador avalia se um texto em Markdown trata centralmente, contextualmente ou apenas de forma superficial de uma determinada entidade alvo (ex: empresa, figura pública, clube, conceito). A entidade é descrita através de um **ECP Snapshot (Entity Context Profile)** contendo nomes canônicos, apelidos (*aliases*), âncoras temáticas, âncoras negativas e entidades de contexto relacional.
|
||||||
|
|
||||||
|
O sistema suporta nativamente **6 idiomas**: Português (`pt`), Inglês (`en`), Espanhol (`es`), Francês (`fr`), Alemão (`de`) e Italiano (`it`).
|
||||||
|
|
||||||
|
### Arquitetura de Classificação em 3 Tiers
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
MD[Markdown Content] --> Pre[Pré-processamento & Detecção de Idioma]
|
||||||
|
ECP[ECP Snapshot JSON] --> Pre
|
||||||
|
Pre --> T1[Tier 1: Determinístico & Regras NLP]
|
||||||
|
T1 -- Alta Confiança --> Decision[Decisão Final JSON]
|
||||||
|
T1 -- Score Intermediário / Ambíguo --> T2{Tier 2 Habilitado?}
|
||||||
|
T2 -- Sim --> Emb[Tier 2: Similaridade Semântica / Embeddings]
|
||||||
|
T2 -- Não --> Decision
|
||||||
|
Emb -- Inconclusivo --> T3{Tier 3 Habilitado?}
|
||||||
|
T3 -- Sim --> LLM[Tier 3: LLM Inherence Adapter]
|
||||||
|
T3 -- Não --> Decision
|
||||||
|
LLM --> Decision
|
||||||
|
```
|
||||||
|
|
||||||
|
* **Tier 1 (Determinístico / NLP Leve)**: Análise de frequência de termos, detecção de âncoras temáticas no primeiro terço do documento, contagem de aliases e penalização por âncoras negativas.
|
||||||
|
* **Tier 2 (Vetorial / Embeddings)** *(Opcional: `--enable-embeddings`)*: Projeção vetorial e cálculo de cosseno entre o perfil da entidade e os parágrafos do documento.
|
||||||
|
* **Tier 3 (LLM Fallback)** *(Opcional: `--enable-llm`)*: Consulta a modelo de linguagem para desambiguação de casos limiares e sutilezas semânticas.
|
||||||
|
|
||||||
|
### Categorias de Decisão
|
||||||
|
|
||||||
|
| Categoria | Descrição |
|
||||||
|
| :--- | :--- |
|
||||||
|
| `DIRECT_INHERENT` | O conteúdo é centrado e focado diretamente na entidade alvo. |
|
||||||
|
| `CONTEXTUAL_INHERENT` | A entidade é relevante no contexto da discussão, mesmo dividindo foco com outros temas. |
|
||||||
|
| `TANGENTIAL` | A entidade é mencionada apenas de passagem ou em listas ilustrativas. |
|
||||||
|
| `NOT_RELATED` | O conteúdo não possui relação substantiva com a entidade alvo. |
|
||||||
|
|
||||||
|
### Formato do ECP Snapshot e Markdown
|
||||||
|
|
||||||
|
#### Exemplo de ECP Snapshot (`ecp_sample.json`):
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"target_entity_id": "ent_river_plate",
|
||||||
|
"target_name": "River Plate",
|
||||||
|
"aliases": ["Club Atlético River Plate", "Millonario", "El Más Grande"],
|
||||||
|
"domain": "sports/football",
|
||||||
|
"anchors": ["Monumental", "Copa Libertadores", "Marcelo Gallardo", "Copa Sudamericana"],
|
||||||
|
"negative_anchors": ["Boca Juniors vitória", "Flamengo campeão"],
|
||||||
|
"graph_version": "1.0.0",
|
||||||
|
"related_entities": [
|
||||||
|
{
|
||||||
|
"entity_id": "ent_gallardo",
|
||||||
|
"name": "Marcelo Gallardo",
|
||||||
|
"relation_type": "manager",
|
||||||
|
"weight": 0.85,
|
||||||
|
"aliases": ["Muñeco"]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Exemplos de Uso CLI
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Classificação Determinística padrão (Tier 1)
|
||||||
|
python classify.py --ecp tests/fixtures/sample_ecp.json --content tests/fixtures/sample_article.md
|
||||||
|
|
||||||
|
# Salvar resultado em arquivo JSON formatado
|
||||||
|
python classify.py --ecp ecp.json --content artigo.md --output out/resultado_classificacao.json
|
||||||
|
|
||||||
|
# Habilitar camadas adicionais (Embeddings e LLM)
|
||||||
|
python classify.py --ecp ecp.json --content artigo.md --enable-embeddings --enable-llm
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 📰 Extrator de Manchetes do Google News
|
||||||
|
|
||||||
|
### O que é e Como Funciona
|
||||||
|
|
||||||
|
O script [`scripts/extract_google_news.py`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/scripts/extract_google_news.py) é um extrator CLI autônomo projetado para consultar o feed RSS do Google News com máxima velocidade, resiliência e integridade de dados.
|
||||||
|
|
||||||
|
### Diferenciais Técnicos
|
||||||
|
|
||||||
|
* 🛡️ **Foxcape Headless Anti-Bot**: Utiliza o motor `foxcape` com `FoxcapeConfig(headless=True)` e Camoufox para evitar bloqueios 429/403 e captchas em segundo plano sem abrir navegadores visuais.
|
||||||
|
* 🔗 **Resolução Automática de URLs Reais**: Decodifica em paralelo (`ThreadPoolExecutor`) as URLs intermediárias do Google News (`news.google.com/rss/articles/...`) entregando diretamente o link final do portal de notícia (*Olé, TyC Sports, ge, ESPN, BBC, etc.*).
|
||||||
|
* 🧹 **Higienização de HTML**: Sanitiza resumos e descrições, eliminando tags HTML residuais (`<ol>`, `<li>`, `<a>`, `<span>`).
|
||||||
|
* 📊 **Logging Informativo no Terminal**: Emite o status passo a passo no canal `stderr` sem poluir a saída JSON em `stdout` (100% compatível com pipes e `jq`).
|
||||||
|
* 🌎 **Mapeamento de Idioma & Região**: Mapeia automaticamente códigos regionais (`pt` → `BR:pt-BR`, `es` → `AR:es-419`, `en` → `US:en-US`) permitindo customização com `--locale`.
|
||||||
|
|
||||||
|
### Argumentos e Flags de Linha de Comando
|
||||||
|
|
||||||
|
| Flag | Tipo | Obrigatório | Padrão | Descrição |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| `-q`, `--query`, `--keyword` | `str` | **Sim** | — | Termo de pesquisa (ex: `"River Plate"`). |
|
||||||
|
| `-l`, `--lang`, `--language` | `str` | Não | `"pt"` | Código do idioma (`pt`, `es`, `en`, `de`, `fr`, `it`). |
|
||||||
|
| `--locale`, `--country` | `str` | Não | `None` | País/região editorial (`BR`, `AR`, `MX`, `US`, `GB`, `ES`). |
|
||||||
|
| `-p`, `--max-pages` | `int` | Não | `1` | Quantidade de páginas (1 a 10; cada página traz 10 itens). |
|
||||||
|
| `-o`, `--output` | `str` | Não | `None` | Caminho para salvar o arquivo JSON. |
|
||||||
|
| `--pretty` | `flag` | Não | `False` | Formata o JSON no terminal com indentação de 2 espaços. |
|
||||||
|
| `--no-resolve-urls` | `flag` | Não | `False` | Desativa a resolução e mantém as URLs brutas do Google News. |
|
||||||
|
| `-s`, `--silent`, `--quiet` | `flag` | Não | `False` | Suprime os logs de progresso no `stderr`. |
|
||||||
|
|
||||||
|
### Exemplos Práticos de Uso
|
||||||
|
|
||||||
|
#### 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "River Plate" -l es --locale AR -p 2 -o out/river_plate.json
|
||||||
|
```
|
||||||
|
* **Saída no Terminal**:
|
||||||
|
```text
|
||||||
|
[INFO] 🔍 Consultando Google News: 'River Plate' (idioma: es, locale: AR, max_pages: 2)...
|
||||||
|
[INFO] 📥 Feed RSS recebido (162117 bytes).
|
||||||
|
[INFO] 📰 20 artigos extraídos do feed XML.
|
||||||
|
[INFO] 🔗 Decodificando 20 URLs do Google News para os portais reais...
|
||||||
|
[INFO] ✅ 20/20 URLs resolvidas com sucesso para os domínios de origem.
|
||||||
|
[INFO] 💾 Arquivo salvo com sucesso: 'out/river_plate.json' (20 notícias).
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 2. Cruzeiro (Brasil / Português / Formatado no Terminal)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py --query "Cruzeiro" --lang pt --locale BR --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 3. Fórmula 1 (Inglaterra / Inglês)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py --query "Formula 1" --lang en --locale GB --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
#### 4. Filtragem com `jq` em Modo Silencioso
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "inteligência artificial" -s | jq '.items[].url'
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📁 Estrutura do Projeto
|
||||||
|
|
||||||
|
```text
|
||||||
|
TextNLPClassifierApp/
|
||||||
|
├── classify.py # CLI principal do Classificador de Inerência
|
||||||
|
├── scripts/
|
||||||
|
│ ├── __init__.py # Pacote utilitário de scripts
|
||||||
|
│ └── extract_google_news.py # CLI de Extração de Manchetes do Google News
|
||||||
|
├── src/ # Módulos centrais do classificador
|
||||||
|
│ ├── classifier.py # Orquestrador de classificação (Tier 1, 2, 3)
|
||||||
|
│ ├── models.py # Modelos de dados e esquemas (ECPSnapshot, Decision)
|
||||||
|
│ ├── preprocessor.py # Normalização de texto e detecção de idioma
|
||||||
|
│ └── adapters/ # Adaptadores opcionais de Embeddings e LLM
|
||||||
|
├── specs/ # Especificações e planos arquiteturais (Speckit)
|
||||||
|
│ ├── 001-nlp-classifier/ # Especificações do classificador
|
||||||
|
│ └── 002-google-news-extractor/ # Especificações do extrator de notícias
|
||||||
|
├── tests/ # Suíte de testes automatizados
|
||||||
|
│ ├── fixtures/ # Amostras de ECP, Markdown e XML RSS
|
||||||
|
│ ├── test_classifier.py # Testes unitários do classificador
|
||||||
|
│ └── test_extract_google_news.py # Testes unitários e testes E2E ao vivo
|
||||||
|
├── requirements.txt # Dependências do projeto
|
||||||
|
├── pyproject.toml # Configurações de ferramentas (pytest, ruff, mypy)
|
||||||
|
└── README.md # Documentação principal
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 🧪 Testes e Qualidade de Código
|
||||||
|
|
||||||
|
O repositório possui cobertura com testes unitários, testes de integração e testes End-to-End (E2E) com requisição de rede ao vivo:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Executar todos os testes do projeto
|
||||||
|
pytest -v
|
||||||
|
|
||||||
|
# Executar especificamente os testes do Extrator de Notícias
|
||||||
|
pytest tests/test_extract_google_news.py -v
|
||||||
|
|
||||||
|
# Validação e correção automática de formatação com Ruff
|
||||||
|
ruff check --fix .
|
||||||
|
ruff format .
|
||||||
|
|
||||||
|
# Verificação estática de tipos com Mypy
|
||||||
|
mypy scripts/ src/
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 📄 Licença
|
||||||
|
|
||||||
|
Este projeto está licenciado sob os termos da licença **MIT**. Consulte o arquivo `LICENSE` para mais detalhes.
|
||||||
@@ -0,0 +1,561 @@
|
|||||||
|
# Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||||
|
|
||||||
|
> Documento auto-contido. Pode ser copiado e colado em outra sessão sem memória:
|
||||||
|
> ele contém toda a informação necessária para entender, reproduzir ou portar
|
||||||
|
> o extrator de manchetes do Google News deste repositório (`GoogleNewsETL`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Visão geral (arquitetura)
|
||||||
|
|
||||||
|
O extrator segue **Clean Architecture** em camadas:
|
||||||
|
|
||||||
|
```
|
||||||
|
┌──────────────────────────────────────────────────────────────┐
|
||||||
|
│ application/ (casos de uso + DTOs) │
|
||||||
|
│ extract_news_use_case.py → orquestra o fluxo │
|
||||||
|
│ dtos/extract_news_dto.py → contratos de entrada/saída │
|
||||||
|
├──────────────────────────────────────────────────────────────┤
|
||||||
|
│ domain/ (entidades, value objects, portas, serviços) │
|
||||||
|
│ entities/news_article.py → entidade NewsArticle │
|
||||||
|
│ entities/search_query.py → value object SearchQuery │
|
||||||
|
│ ports/news_extractor_port.py → interface NewsExtractorPort│
|
||||||
|
│ ports/url_resolver_port.py → interface UrlResolverPort │
|
||||||
|
│ services/rate_limiter_service.py → throttling │
|
||||||
|
├──────────────────────────────────────────────────────────────┤
|
||||||
|
│ infrastructure/ (adaptadores concretos) │
|
||||||
|
│ adapters/google_news_extractor_adapter.py → ★ o extrator │
|
||||||
|
│ adapters/playwright_url_resolver_adapter.py → resolve URLs│
|
||||||
|
└──────────────────────────────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
**Princípio chave:** o caso de uso depende da *porta* (`NewsExtractorPort`), nunca do adaptador concreto. O `GoogleNewsExtractorAdapter` implementa essa porta. Isso permite trocar o mecanismo de raspagem (RSS, Playwright, etc.) sem tocar no domínio.
|
||||||
|
|
||||||
|
**Fluxo resumido:**
|
||||||
|
|
||||||
|
```
|
||||||
|
InputDTO ──► ExtractNewsUseCase.execute() ──► SearchQuery (validação)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
NewsExtractorPort.extract(query)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
GoogleNewsExtractorAdapter.extract()
|
||||||
|
├─ _fetch_rss() → parseia feed RSS
|
||||||
|
└─ url_resolver.resolve_batch() → URLs finais
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
list[NewsArticle]
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
ExtractNewsOutputDTO (JSON)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Entrada
|
||||||
|
|
||||||
|
### 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
class ExtractNewsInputDTO(BaseModel):
|
||||||
|
"""DTO de Entrada do Caso de Uso de Extração."""
|
||||||
|
|
||||||
|
keyword: str = Field(..., description="Palavra ou expressão de busca")
|
||||||
|
language: str = Field(..., description="Código do idioma (ex: 'es', 'pt', 'en')")
|
||||||
|
max_pages: int = Field(default=3, ge=1, le=10, description="Quantidade de páginas para extrair (1 a 10)")
|
||||||
|
```
|
||||||
|
|
||||||
|
| Campo | Tipo | Obrigatório | Descrição |
|
||||||
|
|------------|------|-------------|--------------------------------------------|
|
||||||
|
| `keyword` | str | sim | Palavra/expressão de busca |
|
||||||
|
| `language` | str | sim | Código do idioma (ex: `pt`, `en`, `es`) |
|
||||||
|
| `max_pages`| int | não (def=3) | Páginas a extrair, entre **1 e 10** (10 itens/página) |
|
||||||
|
|
||||||
|
### 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`)
|
||||||
|
|
||||||
|
O use case **nunca usa o DTO cru**: converte-o em `SearchQuery`, que valida as regras de domínio no `__post_init__`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class SearchQuery:
|
||||||
|
"""Value Object representando os parâmetros validados de consulta."""
|
||||||
|
|
||||||
|
keyword: str
|
||||||
|
language: str
|
||||||
|
max_pages: int = 3
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
if not self.keyword or not self.keyword.strip():
|
||||||
|
raise InvalidSearchQueryError("A palavra-chave não pode ser vazia.")
|
||||||
|
|
||||||
|
if not self.language or len(self.language.strip()) < 2:
|
||||||
|
raise InvalidSearchQueryError("O idioma deve conter pelo menos 2 caracteres (ex: 'es', 'pt', 'en').")
|
||||||
|
|
||||||
|
if self.max_pages < 1 or self.max_pages > 10:
|
||||||
|
raise InvalidSearchQueryError("O número máximo de páginas deve estar entre 1 e 10.")
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean_keyword(self) -> str:
|
||||||
|
return self.keyword.strip()
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean_language(self) -> str:
|
||||||
|
return self.language.strip().lower()
|
||||||
|
```
|
||||||
|
|
||||||
|
- Campos congelados (`frozen=True`) → imutáveis.
|
||||||
|
- `clean_keyword` / `clean_language` são os valores normalizados usados na busca.
|
||||||
|
- Falha de validação lança `InvalidSearchQueryError` (exceção de domínio).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Processamento — passo a passo
|
||||||
|
|
||||||
|
### 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`)
|
||||||
|
|
||||||
|
É o orquestrador completo. Ele faz 3 coisas:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class ExtractNewsUseCase:
|
||||||
|
def __init__(self, extractor: NewsExtractorPort | None = None) -> None:
|
||||||
|
if extractor is None:
|
||||||
|
from googlenews_etl.infrastructure.adapters.google_news_extractor_adapter import (
|
||||||
|
GoogleNewsExtractorAdapter,
|
||||||
|
)
|
||||||
|
self.extractor = GoogleNewsExtractorAdapter()
|
||||||
|
else:
|
||||||
|
self.extractor = extractor
|
||||||
|
|
||||||
|
def execute(self, input_dto: ExtractNewsInputDTO) -> ExtractNewsOutputDTO:
|
||||||
|
# 1. Validação de Domínio (Value Object)
|
||||||
|
search_query = SearchQuery(
|
||||||
|
keyword=input_dto.keyword,
|
||||||
|
language=input_dto.language,
|
||||||
|
max_pages=input_dto.max_pages,
|
||||||
|
)
|
||||||
|
|
||||||
|
# 2. Execução da Extração via Porta (Desacoplada)
|
||||||
|
articles = self.extractor.extract(search_query)
|
||||||
|
|
||||||
|
# 3. Mapeamento de Entidades de Domínio -> Output DTO
|
||||||
|
article_dtos = [
|
||||||
|
NewsArticleDTO(
|
||||||
|
titulo=art.title,
|
||||||
|
subtitulo=art.subtitle,
|
||||||
|
quando_publicado=art.published_at,
|
||||||
|
url=art.url,
|
||||||
|
pagina=art.page,
|
||||||
|
)
|
||||||
|
for art in articles
|
||||||
|
]
|
||||||
|
|
||||||
|
return ExtractNewsOutputDTO(
|
||||||
|
query=search_query.clean_keyword,
|
||||||
|
language=search_query.clean_language,
|
||||||
|
total_paginas=search_query.max_pages,
|
||||||
|
total_itens=len(article_dtos),
|
||||||
|
scraped_at=datetime.now(UTC).isoformat(),
|
||||||
|
items=article_dtos,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Notas importantes:
|
||||||
|
- **Injeção de dependência com fallback:** se não receber um extrator, o use case instancia `GoogleNewsExtractorAdapter()` por padrão (import lazy dentro do `__init__`).
|
||||||
|
- A extração passa pela **porta** `NewsExtractorPort.extract(query)` — desacoplamento da infraestrutura.
|
||||||
|
- A saída é montada com `scraped_at` = timestamp UTC ISO do momento da raspagem.
|
||||||
|
|
||||||
|
### 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
class NewsExtractorPort(ABC):
|
||||||
|
"""
|
||||||
|
Porta (Interface) para o serviço de extração de notícias do Google News.
|
||||||
|
Permite desacoplar totalmente o mecanismo de raspagem (HTTP, Playwright, RSS, RabbitMQ, etc.) do domínio.
|
||||||
|
"""
|
||||||
|
|
||||||
|
@abstractmethod
|
||||||
|
def extract(self, query: SearchQuery) -> list[NewsArticle]:
|
||||||
|
"""Extrai as notícias correspondentes aos critérios de busca."""
|
||||||
|
pass
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.3 O adaptador concreto — o coração do extrator (`googlenews_etl/infrastructure/adapters/google_news_extractor_adapter.py`)
|
||||||
|
|
||||||
|
#### 3.3.1 Inicialização: sessão HTTP com impersonação de browser
|
||||||
|
|
||||||
|
```python
|
||||||
|
class GoogleNewsExtractorAdapter(NewsExtractorPort):
|
||||||
|
DEFAULT_HEADERS = {
|
||||||
|
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
|
||||||
|
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/132.0.0.0 Safari/537.36",
|
||||||
|
}
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
rate_limiter: RateLimiterService | None = None,
|
||||||
|
url_resolver: UrlResolverPort | None = None,
|
||||||
|
impersonate: str = "chrome120",
|
||||||
|
resolve_final_urls: bool = True,
|
||||||
|
) -> None:
|
||||||
|
self.impersonate = impersonate
|
||||||
|
self.rate_limiter = rate_limiter or RateLimiterService(min_delay_seconds=0.5, max_delay_seconds=1.0)
|
||||||
|
self.url_resolver = url_resolver or PlaywrightUrlResolverAdapter()
|
||||||
|
self.resolve_final_urls = resolve_final_urls
|
||||||
|
self.session = requests.Session(impersonate=self.impersonate)
|
||||||
|
self.session.headers.update(self.DEFAULT_HEADERS)
|
||||||
|
```
|
||||||
|
|
||||||
|
- Usa `curl_cffi` com `impersonate="chrome120"` — imita o TLS fingerprint do Chrome para evitar bloqueios.
|
||||||
|
- `RateLimiterService` com delay aleatório 0.5–1.0s entre requisições (configurável).
|
||||||
|
- `resolve_final_urls=True` por padrão → resolve as URLs intermediárias do Google para as URLs reais dos veículos.
|
||||||
|
|
||||||
|
#### 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _get_hl_gl(self, lang_raw: str) -> tuple[str, str]:
|
||||||
|
"""Mapeia dinamicamente código de idioma e país para os parâmetros hl e gl do Google News."""
|
||||||
|
lang_clean = lang_raw.lower().replace("-", "_")
|
||||||
|
|
||||||
|
locale_map = {
|
||||||
|
"pt": ("pt-BR", "BR"),
|
||||||
|
"pt_br": ("pt-BR", "BR"),
|
||||||
|
"es": ("es-419", "AR"),
|
||||||
|
"es_mx": ("es-419", "MX"),
|
||||||
|
"es_es": ("es", "ES"),
|
||||||
|
"en": ("en-US", "US"),
|
||||||
|
"en_gb": ("en-GB", "GB"),
|
||||||
|
"en_uk": ("en-GB", "GB"),
|
||||||
|
"en_us": ("en-US", "US"),
|
||||||
|
"de": ("de", "DE"),
|
||||||
|
"de_de": ("de", "DE"),
|
||||||
|
"it": ("it", "IT"),
|
||||||
|
"it_it": ("it", "IT"),
|
||||||
|
"fr": ("fr", "FR"),
|
||||||
|
}
|
||||||
|
|
||||||
|
if lang_clean in locale_map:
|
||||||
|
return locale_map[lang_clean]
|
||||||
|
|
||||||
|
parts = lang_clean.split("_")
|
||||||
|
if len(parts) == 2:
|
||||||
|
return (f"{parts[0]}-{parts[1].upper()}", parts[1].upper())
|
||||||
|
|
||||||
|
return (lang_clean, lang_clean.upper())
|
||||||
|
```
|
||||||
|
|
||||||
|
- `hl` = idioma da interface, `gl` = país da região (ex: `pt` → `pt-BR`/`BR`).
|
||||||
|
- Fallback genérico: `xx_yy` → `xx-YY`/`YY`; senão `xx`/`XX`.
|
||||||
|
|
||||||
|
#### 3.3.3 ★ Extração do RSS — o ponto fundamental (`_fetch_rss`)
|
||||||
|
|
||||||
|
**Este é o trecho que extrai de fato as manchetes com título, URL e data de publicação:**
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _fetch_rss(self, query: SearchQuery) -> list[NewsArticle]:
|
||||||
|
encoded_query = urllib.parse.quote_plus(query.clean_keyword)
|
||||||
|
hl, gl = self._get_hl_gl(query.clean_language)
|
||||||
|
rss_url = f"https://news.google.com/rss/search?q={encoded_query}&hl={hl}&gl={gl}&ceid={gl}:{hl}"
|
||||||
|
|
||||||
|
response = self.session.get(rss_url)
|
||||||
|
response.raise_for_status()
|
||||||
|
|
||||||
|
soup = BeautifulSoup(response.content, "xml")
|
||||||
|
rss_items = soup.find_all("item")
|
||||||
|
|
||||||
|
articles: list[NewsArticle] = []
|
||||||
|
max_allowed = query.max_pages * 10
|
||||||
|
|
||||||
|
for idx, item in enumerate(rss_items[:max_allowed]):
|
||||||
|
page_number = (idx // 10) + 1
|
||||||
|
title = item.find("title").text if item.find("title") else ""
|
||||||
|
link = item.find("link").text if item.find("link") else ""
|
||||||
|
pub_date = item.find("pubDate").text if item.find("pubDate") else ""
|
||||||
|
desc_raw = item.find("description").text if item.find("description") else ""
|
||||||
|
|
||||||
|
desc_soup = BeautifulSoup(desc_raw, "html.parser")
|
||||||
|
snippet = desc_soup.get_text(separator=" ", strip=True) if desc_raw else None
|
||||||
|
|
||||||
|
if title and link:
|
||||||
|
articles.append(
|
||||||
|
NewsArticle(
|
||||||
|
title=title,
|
||||||
|
subtitle=snippet if snippet != title else None,
|
||||||
|
published_at=pub_date,
|
||||||
|
url=link,
|
||||||
|
page=page_number,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
return articles
|
||||||
|
```
|
||||||
|
|
||||||
|
**Pontos fundamentais deste trecho:**
|
||||||
|
|
||||||
|
| # | Mecanismo | Detalhe |
|
||||||
|
|---|-----------|---------|
|
||||||
|
| 1 | **URL do RSS** | `https://news.google.com/rss/search?q={keyword}&hl={hl}&gl={gl}&ceid={gl}:{hl}` — RSS oficial de busca do Google News |
|
||||||
|
| 2 | **Parsing XML** | `BeautifulSoup(response.content, "xml")` + `find_all("item")` (formato RSS padrão) |
|
||||||
|
| 3 | **Limite** | `max_pages * 10` itens — cada "página" do Google News = 10 itens |
|
||||||
|
| 4 | **Nº da página** | `page_number = (idx // 10) + 1` — agrupa os itens em páginas de 10 |
|
||||||
|
| 5 | **Campos extraídos por item** | `title`, `link`, `pubDate`, `description` (XML do RSS) |
|
||||||
|
| 6 | **Limpeza do resumo** | `description` contém HTML — `BeautifulSoup(desc_raw, "html.parser")` + `get_text(separator=" ", strip=True)` remove tags |
|
||||||
|
| 7 | **Filtro** | item só entra se tiver `title` **e** `link` |
|
||||||
|
| 8 | **Dedupe de subtítulo** | `subtitle = snippet if snippet != title else None` — se o resumo for igual ao título, fica `None` |
|
||||||
|
|
||||||
|
#### 3.3.4 Orquestração com resolução de URLs (`extract`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
def extract(self, query: SearchQuery) -> list[NewsArticle]:
|
||||||
|
# Busca direta das manchetes pelo RSS oficial do Google News
|
||||||
|
all_articles = self._fetch_rss(query)
|
||||||
|
|
||||||
|
# Resolução paralela das URLs finais dos veículos se ativado
|
||||||
|
if self.resolve_final_urls and all_articles:
|
||||||
|
raw_urls = [a.url for a in all_articles]
|
||||||
|
resolved_urls = self.url_resolver.resolve_batch(raw_urls)
|
||||||
|
|
||||||
|
resolved_articles: list[NewsArticle] = []
|
||||||
|
for idx, article in enumerate(all_articles):
|
||||||
|
new_url = resolved_urls[idx] if idx < len(resolved_urls) else article.url
|
||||||
|
resolved_articles.append(dataclasses.replace(article, url=new_url))
|
||||||
|
|
||||||
|
all_articles = resolved_articles
|
||||||
|
|
||||||
|
return all_articles
|
||||||
|
```
|
||||||
|
|
||||||
|
- As URLs do RSS do Google News são intermediárias (`news.google.com/rss/articles/CBMi...`).
|
||||||
|
- `resolve_batch` resolve **em paralelo** (Playwright headless) para as URLs diretas dos veículos.
|
||||||
|
- `dataclasses.replace(article, url=new_url)` preserva os demais campos.
|
||||||
|
|
||||||
|
### 3.4 Resolução de URLs (`googlenews_etl/infrastructure/adapters/playwright_url_resolver_adapter.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
class PlaywrightUrlResolverAdapter(UrlResolverPort):
|
||||||
|
def __init__(self, timeout_ms: int = 6000, max_concurrent: int = 5) -> None:
|
||||||
|
self.timeout_ms = timeout_ms
|
||||||
|
self.max_concurrent = max_concurrent
|
||||||
|
|
||||||
|
async def _resolve_single_async(self, context, semaphore, url: str) -> str:
|
||||||
|
if not url or "news.google.com/rss/articles/" not in url:
|
||||||
|
return url
|
||||||
|
|
||||||
|
async with semaphore:
|
||||||
|
page = await context.new_page()
|
||||||
|
target_url = url
|
||||||
|
|
||||||
|
def handle_request(req):
|
||||||
|
nonlocal target_url
|
||||||
|
u = req.url
|
||||||
|
if not any(
|
||||||
|
x in u
|
||||||
|
for x in [
|
||||||
|
"google.", "gstatic.", "googleapis.", "googletagmanager.",
|
||||||
|
"w3.org", "schema.org",
|
||||||
|
]
|
||||||
|
):
|
||||||
|
if not target_url or target_url == url:
|
||||||
|
if u.startswith("http"):
|
||||||
|
target_url = u
|
||||||
|
|
||||||
|
page.on("request", handle_request)
|
||||||
|
try:
|
||||||
|
await page.goto(url, wait_until="commit", timeout=self.timeout_ms)
|
||||||
|
for _ in range(12):
|
||||||
|
await asyncio.sleep(0.25)
|
||||||
|
if "google.com" not in page.url:
|
||||||
|
target_url = page.url
|
||||||
|
break
|
||||||
|
if target_url and target_url != url:
|
||||||
|
break
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
finally:
|
||||||
|
await page.close()
|
||||||
|
|
||||||
|
return target_url or url
|
||||||
|
```
|
||||||
|
|
||||||
|
- Abre cada URL em uma página headless e **captura o primeiro request não-Google** (o redirecionamento para o veículo).
|
||||||
|
- `Semaphore(max_concurrent=5)` limita concorrência; timeout de 6s por URL.
|
||||||
|
- Intercepta requisições (`page.on("request")`) filtrando domínios de Google/telemetria.
|
||||||
|
|
||||||
|
### 3.5 Rate limiter (`googlenews_etl/domain/services/rate_limiter_service.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
class RateLimiterService:
|
||||||
|
def __init__(self, min_delay_seconds: float = 1.0, max_delay_seconds: float = 2.5) -> None:
|
||||||
|
self.min_delay = min_delay_seconds
|
||||||
|
self.max_delay = max_delay_seconds
|
||||||
|
|
||||||
|
def wait(self) -> float:
|
||||||
|
"""Aplica uma pausa aleatória dentro dos limites configurados e retorna o tempo aguardado."""
|
||||||
|
delay = random.uniform(self.min_delay, self.max_delay)
|
||||||
|
time.sleep(delay)
|
||||||
|
return delay
|
||||||
|
```
|
||||||
|
|
||||||
|
- Delay **aleatório** entre requisições (evita padrão detectável/anti-bot).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Entidade de domínio da notícia (`googlenews_etl/domain/entities/news_article.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class NewsArticle:
|
||||||
|
"""Entidade do Domínio representando uma notícia extraída do Google News."""
|
||||||
|
|
||||||
|
title: str
|
||||||
|
url: str
|
||||||
|
page: int
|
||||||
|
subtitle: str | None = None
|
||||||
|
published_at: str | None = None
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
if not self.title or not self.title.strip():
|
||||||
|
raise ValueError("O título da notícia não pode ser vazio.")
|
||||||
|
if not self.url or not self.url.strip():
|
||||||
|
raise ValueError("A URL da notícia não pode ser vazia.")
|
||||||
|
```
|
||||||
|
|
||||||
|
| Campo | Tipo | Origem no RSS |
|
||||||
|
|----------------|------------|----------------------------------|
|
||||||
|
| `title` | str | `<title>` do `<item>` |
|
||||||
|
| `url` | str | `<link>` do `<item>` (resolvida depois) |
|
||||||
|
| `page` | int | Calculado: `(idx // 10) + 1` |
|
||||||
|
| `subtitle` | str\|None | `<description>` limpo (ou `None`)|
|
||||||
|
| `published_at` | str\|None | `<pubDate>` do `<item>` |
|
||||||
|
|
||||||
|
**Nota:** `published_at` é mantido como **string crua** do RSS (formato RFC 822, ex: `Thu, 20 Aug 2026 10:00:00 GMT`). Nenhum parse de data acontece no extrator.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Saída
|
||||||
|
|
||||||
|
### 5.1 DTOs de saída (`googlenews_etl/application/dtos/extract_news_dto.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
class NewsArticleDTO(BaseModel):
|
||||||
|
"""DTO individual para representação de cada artigo."""
|
||||||
|
|
||||||
|
titulo: str
|
||||||
|
subtitulo: str | None = None
|
||||||
|
quando_publicado: str | None = None
|
||||||
|
url: str
|
||||||
|
pagina: int
|
||||||
|
|
||||||
|
|
||||||
|
class ExtractNewsOutputDTO(BaseModel):
|
||||||
|
"""DTO de Saída estruturado com o resultado consolidador do ETL."""
|
||||||
|
|
||||||
|
query: str
|
||||||
|
language: str
|
||||||
|
total_paginas: int
|
||||||
|
total_itens: int
|
||||||
|
scraped_at: str
|
||||||
|
items: list[NewsArticleDTO]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5.2 Exemplo de saída JSON
|
||||||
|
|
||||||
|
Entrada: `{"keyword": "inteligencia artificial", "language": "pt", "max_pages": 1}`
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"query": "inteligencia artificial",
|
||||||
|
"language": "pt",
|
||||||
|
"total_paginas": 1,
|
||||||
|
"total_itens": 10,
|
||||||
|
"scraped_at": "2026-08-20T14:32:10.482930+00:00",
|
||||||
|
"items": [
|
||||||
|
{
|
||||||
|
"titulo": "Empresas aceleram adoção de inteligência artificial no Brasil",
|
||||||
|
"subtitulo": "Levantamento mostra crescimento de 40% no uso de IA generativa...",
|
||||||
|
"quando_publicado": "Thu, 20 Aug 2026 09:12:00 GMT",
|
||||||
|
"url": "https://exemplo.com.br/noticia/123",
|
||||||
|
"pagina": 1
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Resumo do fluxo completo (entrada → saída)
|
||||||
|
|
||||||
|
```
|
||||||
|
1. ExtractNewsInputDTO(keyword, language, max_pages)
|
||||||
|
│
|
||||||
|
2. SearchQuery (valida: keyword não vazia, language ≥ 2 chars, 1 ≤ max_pages ≤ 10)
|
||||||
|
│
|
||||||
|
3. NewsExtractorPort.extract(query) ← porta (abstração)
|
||||||
|
│
|
||||||
|
4. GoogleNewsExtractorAdapter.extract(query)
|
||||||
|
│ a) _get_hl_gl(language) → (hl, gl) ex: "pt" → ("pt-BR", "BR")
|
||||||
|
│ b) _fetch_rss(query):
|
||||||
|
│ URL: https://news.google.com/rss/search?q=...&hl=...&gl=...&ceid=...
|
||||||
|
│ GET com curl_cffi (impersonate chrome120)
|
||||||
|
│ BeautifulSoup XML → itens
|
||||||
|
│ extrai title, link, pubDate, description (HTML limpo)
|
||||||
|
│ limita a max_pages * 10 itens, pagina = (idx // 10) + 1
|
||||||
|
│ c) resolve_batch(urls) via Playwright → URLs finais dos veículos
|
||||||
|
│
|
||||||
|
5. list[NewsArticle] (entidade imutável com validação title/url não vazios)
|
||||||
|
│
|
||||||
|
6. Mapeamento → list[NewsArticleDTO] (titulo, subtitulo, quando_publicado, url, pagina)
|
||||||
|
│
|
||||||
|
7. ExtractNewsOutputDTO (query, language, total_paginas, total_itens, scraped_at, items)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Dependências para rodar (o que NÃO está nos 4 arquivos centrais)
|
||||||
|
|
||||||
|
Para portar a extração a outro repositório, além dos 4 arquivos centrais
|
||||||
|
(`extract_news_use_case.py`, `google_news_extractor_adapter.py`, `news_article.py`,
|
||||||
|
`extract_news_dto.py`), é preciso:
|
||||||
|
|
||||||
|
| Componente | Arquivo | Papel |
|
||||||
|
|------------|---------|-------|
|
||||||
|
| `NewsExtractorPort` | `domain/ports/news_extractor_port.py` | Interface que o adaptador implementa |
|
||||||
|
| `SearchQuery` | `domain/entities/search_query.py` | Value object validado usado na busca |
|
||||||
|
| `UrlResolverPort` | `domain/ports/url_resolver_port.py` | Interface do resolvedor de URLs |
|
||||||
|
| `PlaywrightUrlResolverAdapter` | `infrastructure/adapters/playwright_url_resolver_adapter.py` | Resolução das URLs intermediárias |
|
||||||
|
| `RateLimiterService` | `domain/services/rate_limiter_service.py` | Throttling entre requisições |
|
||||||
|
| `InvalidSearchQueryError` | `domain/exceptions/domain_exceptions.py` | Exceção de validação do `SearchQuery` |
|
||||||
|
|
||||||
|
### Dependências de bibliotecas (requirements)
|
||||||
|
|
||||||
|
- `curl_cffi` — sessão HTTP com impersonação de TLS do Chrome
|
||||||
|
- `beautifulsoup4` — parsing XML/HTML do feed
|
||||||
|
- `pydantic` — DTOs de entrada/saída
|
||||||
|
- `playwright` — resolução de URLs (navegador headless)
|
||||||
|
- Python ≥ 3.11 (uso de `X | None` em anotações de tipo)
|
||||||
|
|
||||||
|
### Comportamentos observáveis (honestidade)
|
||||||
|
|
||||||
|
- `published_at` é string crua do RSS (RFC 822), **não** parseada.
|
||||||
|
- `subtitle` vem do `description` do item com HTML removido; vira `None` se igual ao título.
|
||||||
|
- O item só é mantido se tiver `title` e `link` não vazios.
|
||||||
|
- `resolve_final_urls` pode ser desligado (`False`) para pular a etapa Playwright.
|
||||||
|
- O rate limiter existe, mas no `_fetch_rss` atual o delay é aplicado apenas por construção
|
||||||
|
do `RateLimiterService` (o método `wait()` está disponível para chamadas sequenciais).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Como chamar (exemplo de uso)
|
||||||
|
|
||||||
|
```python
|
||||||
|
from googlenews_etl.application.dtos.extract_news_dto import ExtractNewsInputDTO
|
||||||
|
from googlenews_etl.application.use_cases.extract_news_use_case import ExtractNewsUseCase
|
||||||
|
|
||||||
|
dto_in = ExtractNewsInputDTO(keyword="inteligencia artificial", language="pt", max_pages=1)
|
||||||
|
resultado = ExtractNewsUseCase().execute(dto_in)
|
||||||
|
|
||||||
|
print(resultado.total_itens) # ex: 10
|
||||||
|
print(resultado.items[0].titulo) # título da primeira manchete
|
||||||
|
print(resultado.items[0].url) # URL final resolvida
|
||||||
|
print(resultado.items[0].quando_publicado) # data crua do RSS
|
||||||
|
```
|
||||||
@@ -0,0 +1,739 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "ecp-nlp-poc-v1",
|
||||||
|
"profile_kind": "entity_context_profile",
|
||||||
|
"profile_status": "poc_reference",
|
||||||
|
"target_entity_id": "ecp_sao_paulo_fc",
|
||||||
|
"target_name": "São Paulo Futebol Clube",
|
||||||
|
"canonical_name": "São Paulo Futebol Clube",
|
||||||
|
"qid": "Q38568",
|
||||||
|
"entity_type": "sports_club",
|
||||||
|
"domains": [
|
||||||
|
"sports",
|
||||||
|
"football",
|
||||||
|
"brazil_football"
|
||||||
|
],
|
||||||
|
"primary_language": "pt-BR",
|
||||||
|
"languages": [
|
||||||
|
"pt-BR",
|
||||||
|
"pt"
|
||||||
|
],
|
||||||
|
"geo": {
|
||||||
|
"country": "BR",
|
||||||
|
"country_name": "Brazil",
|
||||||
|
"region": "SP",
|
||||||
|
"region_name": "São Paulo",
|
||||||
|
"city": "São Paulo"
|
||||||
|
},
|
||||||
|
"official": {
|
||||||
|
"website": "https://www.saopaulofc.net/",
|
||||||
|
"domains": [
|
||||||
|
"saopaulofc.net"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"identity": {
|
||||||
|
"short_description": "Brazilian professional football club based in São Paulo, Brazil.",
|
||||||
|
"founded_year": 1930,
|
||||||
|
"sport": "association_football",
|
||||||
|
"colors": [
|
||||||
|
"red",
|
||||||
|
"white",
|
||||||
|
"black"
|
||||||
|
],
|
||||||
|
"common_nicknames": [
|
||||||
|
"Tricolor Paulista",
|
||||||
|
"Clube da Fé",
|
||||||
|
"Soberano"
|
||||||
|
],
|
||||||
|
"home_venue": "MorumBIS",
|
||||||
|
"home_venue_aliases": [
|
||||||
|
"Morumbi",
|
||||||
|
"Estádio do Morumbi",
|
||||||
|
"Estádio Cícero Pompeu de Toledo",
|
||||||
|
"Cícero Pompeu de Toledo",
|
||||||
|
"MorumBIS"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"aliases": [
|
||||||
|
"São Paulo Futebol Clube",
|
||||||
|
"São Paulo FC",
|
||||||
|
"São Paulo F.C.",
|
||||||
|
"SPFC",
|
||||||
|
"São Paulo",
|
||||||
|
"Tricolor Paulista",
|
||||||
|
"Clube da Fé",
|
||||||
|
"Soberano"
|
||||||
|
],
|
||||||
|
"alias_rules": [
|
||||||
|
{
|
||||||
|
"value": "São Paulo Futebol Clube",
|
||||||
|
"strength": "strong",
|
||||||
|
"ambiguous": false,
|
||||||
|
"requires_context": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "São Paulo FC",
|
||||||
|
"strength": "strong",
|
||||||
|
"ambiguous": false,
|
||||||
|
"requires_context": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "São Paulo F.C.",
|
||||||
|
"strength": "strong",
|
||||||
|
"ambiguous": false,
|
||||||
|
"requires_context": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "SPFC",
|
||||||
|
"strength": "strong",
|
||||||
|
"ambiguous": false,
|
||||||
|
"requires_context": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "Tricolor Paulista",
|
||||||
|
"strength": "strong",
|
||||||
|
"ambiguous": false,
|
||||||
|
"requires_context": [
|
||||||
|
"football"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "Clube da Fé",
|
||||||
|
"strength": "medium",
|
||||||
|
"ambiguous": true,
|
||||||
|
"requires_context": [
|
||||||
|
"football",
|
||||||
|
"São Paulo FC",
|
||||||
|
"SPFC",
|
||||||
|
"Tricolor Paulista"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "Soberano",
|
||||||
|
"strength": "medium",
|
||||||
|
"ambiguous": true,
|
||||||
|
"requires_context": [
|
||||||
|
"football",
|
||||||
|
"São Paulo FC",
|
||||||
|
"SPFC",
|
||||||
|
"Tricolor Paulista"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"value": "São Paulo",
|
||||||
|
"strength": "weak",
|
||||||
|
"ambiguous": true,
|
||||||
|
"requires_context": [
|
||||||
|
"football",
|
||||||
|
"club",
|
||||||
|
"team",
|
||||||
|
"match",
|
||||||
|
"goal",
|
||||||
|
"coach",
|
||||||
|
"stadium",
|
||||||
|
"championship",
|
||||||
|
"SPFC",
|
||||||
|
"Tricolor Paulista",
|
||||||
|
"Morumbi"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"anchors": [
|
||||||
|
"futebol",
|
||||||
|
"clube",
|
||||||
|
"time",
|
||||||
|
"equipe",
|
||||||
|
"tricolor",
|
||||||
|
"torcida",
|
||||||
|
"campeonato",
|
||||||
|
"partida",
|
||||||
|
"jogo",
|
||||||
|
"gol",
|
||||||
|
"técnico",
|
||||||
|
"elenco",
|
||||||
|
"jogador",
|
||||||
|
"estádio",
|
||||||
|
"Morumbi",
|
||||||
|
"MorumBIS",
|
||||||
|
"Estádio do Morumbi",
|
||||||
|
"Cícero Pompeu de Toledo",
|
||||||
|
"Libertadores",
|
||||||
|
"Copa Libertadores",
|
||||||
|
"Campeonato Brasileiro",
|
||||||
|
"Brasileirão",
|
||||||
|
"Copa do Brasil",
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Paulistão",
|
||||||
|
"Mundial de Clubes",
|
||||||
|
"Rogério Ceni",
|
||||||
|
"Raí",
|
||||||
|
"Telê Santana",
|
||||||
|
"Luis Fabiano",
|
||||||
|
"Luís Fabiano",
|
||||||
|
"Muricy Ramalho",
|
||||||
|
"Calleri",
|
||||||
|
"Luciano",
|
||||||
|
"Lucas Moura"
|
||||||
|
],
|
||||||
|
"anchor_groups": {
|
||||||
|
"strong": [
|
||||||
|
"SPFC",
|
||||||
|
"São Paulo FC",
|
||||||
|
"São Paulo Futebol Clube",
|
||||||
|
"Tricolor Paulista",
|
||||||
|
"Morumbi",
|
||||||
|
"MorumBIS",
|
||||||
|
"Estádio do Morumbi",
|
||||||
|
"Cícero Pompeu de Toledo",
|
||||||
|
"Rogério Ceni",
|
||||||
|
"Raí",
|
||||||
|
"Telê Santana"
|
||||||
|
],
|
||||||
|
"medium": [
|
||||||
|
"Libertadores",
|
||||||
|
"Copa Libertadores",
|
||||||
|
"Campeonato Brasileiro",
|
||||||
|
"Brasileirão",
|
||||||
|
"Copa do Brasil",
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Paulistão",
|
||||||
|
"Mundial de Clubes",
|
||||||
|
"Luis Fabiano",
|
||||||
|
"Luís Fabiano",
|
||||||
|
"Muricy Ramalho"
|
||||||
|
],
|
||||||
|
"weak": [
|
||||||
|
"futebol",
|
||||||
|
"clube",
|
||||||
|
"time",
|
||||||
|
"equipe",
|
||||||
|
"tricolor",
|
||||||
|
"torcida",
|
||||||
|
"campeonato",
|
||||||
|
"partida",
|
||||||
|
"jogo",
|
||||||
|
"gol",
|
||||||
|
"técnico",
|
||||||
|
"elenco",
|
||||||
|
"jogador",
|
||||||
|
"estádio"
|
||||||
|
],
|
||||||
|
"volatile": [
|
||||||
|
"Calleri",
|
||||||
|
"Luciano",
|
||||||
|
"Lucas Moura"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"negative_anchors": [
|
||||||
|
"cidade de São Paulo",
|
||||||
|
"município de São Paulo",
|
||||||
|
"estado de São Paulo",
|
||||||
|
"governo de São Paulo",
|
||||||
|
"governo do estado de São Paulo",
|
||||||
|
"prefeitura de São Paulo",
|
||||||
|
"prefeito de São Paulo",
|
||||||
|
"capital paulista",
|
||||||
|
"região metropolitana de São Paulo",
|
||||||
|
"trânsito em São Paulo",
|
||||||
|
"chuva em São Paulo",
|
||||||
|
"violência em São Paulo",
|
||||||
|
"segurança pública em São Paulo",
|
||||||
|
"evento em São Paulo",
|
||||||
|
"show em São Paulo",
|
||||||
|
"restaurante em São Paulo",
|
||||||
|
"bairro de São Paulo",
|
||||||
|
"bolsa de valores de São Paulo",
|
||||||
|
"B3",
|
||||||
|
"São Paulo Fashion Week",
|
||||||
|
"Universidade de São Paulo",
|
||||||
|
"USP"
|
||||||
|
],
|
||||||
|
"negative_anchor_groups": {
|
||||||
|
"geo_political": [
|
||||||
|
"cidade de São Paulo",
|
||||||
|
"município de São Paulo",
|
||||||
|
"estado de São Paulo",
|
||||||
|
"governo de São Paulo",
|
||||||
|
"governo do estado de São Paulo",
|
||||||
|
"prefeitura de São Paulo",
|
||||||
|
"prefeito de São Paulo",
|
||||||
|
"capital paulista",
|
||||||
|
"região metropolitana de São Paulo"
|
||||||
|
],
|
||||||
|
"urban_news": [
|
||||||
|
"trânsito em São Paulo",
|
||||||
|
"chuva em São Paulo",
|
||||||
|
"violência em São Paulo",
|
||||||
|
"segurança pública em São Paulo",
|
||||||
|
"bairro de São Paulo"
|
||||||
|
],
|
||||||
|
"non_sports_events": [
|
||||||
|
"evento em São Paulo",
|
||||||
|
"show em São Paulo",
|
||||||
|
"restaurante em São Paulo",
|
||||||
|
"São Paulo Fashion Week"
|
||||||
|
],
|
||||||
|
"institutions_or_finance": [
|
||||||
|
"bolsa de valores de São Paulo",
|
||||||
|
"B3",
|
||||||
|
"Universidade de São Paulo",
|
||||||
|
"USP"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"graph_version": "spfc-graph-snapshot-v1",
|
||||||
|
"related_entities": [
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_morumbi",
|
||||||
|
"name": "MorumBIS",
|
||||||
|
"canonical_name": "Estádio Cícero Pompeu de Toledo",
|
||||||
|
"aliases": [
|
||||||
|
"Morumbi",
|
||||||
|
"MorumBIS",
|
||||||
|
"Estádio do Morumbi",
|
||||||
|
"Cícero Pompeu de Toledo",
|
||||||
|
"Estádio Cícero Pompeu de Toledo"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_venue",
|
||||||
|
"relation_type": "HOME_STADIUM",
|
||||||
|
"weight": 0.95,
|
||||||
|
"confidence": 0.95,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"stadium",
|
||||||
|
"match",
|
||||||
|
"club",
|
||||||
|
"torcida",
|
||||||
|
"jogo",
|
||||||
|
"campeonato",
|
||||||
|
"São Paulo FC",
|
||||||
|
"SPFC"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_rogerio_ceni",
|
||||||
|
"name": "Rogério Ceni",
|
||||||
|
"canonical_name": "Rogério Ceni",
|
||||||
|
"aliases": [
|
||||||
|
"Rogério Ceni",
|
||||||
|
"Ceni",
|
||||||
|
"M1TO",
|
||||||
|
"Mito"
|
||||||
|
],
|
||||||
|
"entity_type": "athlete_or_coach",
|
||||||
|
"relation_type": "HAS_IDOL",
|
||||||
|
"weight": 0.9,
|
||||||
|
"confidence": 0.95,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"goalkeeper",
|
||||||
|
"idol",
|
||||||
|
"coach",
|
||||||
|
"history",
|
||||||
|
"SPFC",
|
||||||
|
"tricolor"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_rai",
|
||||||
|
"name": "Raí",
|
||||||
|
"canonical_name": "Raí Souza Vieira de Oliveira",
|
||||||
|
"aliases": [
|
||||||
|
"Raí",
|
||||||
|
"Rai Souza Vieira de Oliveira",
|
||||||
|
"Raí Souza"
|
||||||
|
],
|
||||||
|
"entity_type": "athlete",
|
||||||
|
"relation_type": "HAS_IDOL",
|
||||||
|
"weight": 0.85,
|
||||||
|
"confidence": 0.9,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"idol",
|
||||||
|
"midfielder",
|
||||||
|
"history",
|
||||||
|
"SPFC",
|
||||||
|
"tricolor"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_tele_santana",
|
||||||
|
"name": "Telê Santana",
|
||||||
|
"canonical_name": "Telê Santana",
|
||||||
|
"aliases": [
|
||||||
|
"Telê Santana",
|
||||||
|
"Tele Santana"
|
||||||
|
],
|
||||||
|
"entity_type": "coach",
|
||||||
|
"relation_type": "HISTORIC_COACH",
|
||||||
|
"weight": 0.85,
|
||||||
|
"confidence": 0.9,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"coach",
|
||||||
|
"history",
|
||||||
|
"Libertadores",
|
||||||
|
"Mundial",
|
||||||
|
"SPFC",
|
||||||
|
"tricolor"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_muricy_ramalho",
|
||||||
|
"name": "Muricy Ramalho",
|
||||||
|
"canonical_name": "Muricy Ramalho",
|
||||||
|
"aliases": [
|
||||||
|
"Muricy Ramalho",
|
||||||
|
"Muricy"
|
||||||
|
],
|
||||||
|
"entity_type": "coach",
|
||||||
|
"relation_type": "HISTORIC_COACH",
|
||||||
|
"weight": 0.78,
|
||||||
|
"confidence": 0.85,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"coach",
|
||||||
|
"Brasileirão",
|
||||||
|
"SPFC",
|
||||||
|
"tricolor"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_luis_fabiano",
|
||||||
|
"name": "Luis Fabiano",
|
||||||
|
"canonical_name": "Luís Fabiano",
|
||||||
|
"aliases": [
|
||||||
|
"Luis Fabiano",
|
||||||
|
"Luís Fabiano",
|
||||||
|
"Fabuloso"
|
||||||
|
],
|
||||||
|
"entity_type": "athlete",
|
||||||
|
"relation_type": "FORMER_PLAYER",
|
||||||
|
"weight": 0.75,
|
||||||
|
"confidence": 0.85,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"striker",
|
||||||
|
"player",
|
||||||
|
"goal",
|
||||||
|
"SPFC",
|
||||||
|
"tricolor"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_corinthians",
|
||||||
|
"name": "Sport Club Corinthians Paulista",
|
||||||
|
"canonical_name": "Sport Club Corinthians Paulista",
|
||||||
|
"aliases": [
|
||||||
|
"Corinthians",
|
||||||
|
"Sport Club Corinthians Paulista",
|
||||||
|
"Timão"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_club",
|
||||||
|
"relation_type": "RIVAL_OF",
|
||||||
|
"weight": 0.75,
|
||||||
|
"confidence": 0.9,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"derby",
|
||||||
|
"rival",
|
||||||
|
"clássico",
|
||||||
|
"Majestoso",
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Brasileirão"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_palmeiras",
|
||||||
|
"name": "Sociedade Esportiva Palmeiras",
|
||||||
|
"canonical_name": "Sociedade Esportiva Palmeiras",
|
||||||
|
"aliases": [
|
||||||
|
"Palmeiras",
|
||||||
|
"Sociedade Esportiva Palmeiras",
|
||||||
|
"Verdão"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_club",
|
||||||
|
"relation_type": "RIVAL_OF",
|
||||||
|
"weight": 0.72,
|
||||||
|
"confidence": 0.9,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"derby",
|
||||||
|
"rival",
|
||||||
|
"clássico",
|
||||||
|
"Choque-Rei",
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Brasileirão"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_santos_fc",
|
||||||
|
"name": "Santos Futebol Clube",
|
||||||
|
"canonical_name": "Santos Futebol Clube",
|
||||||
|
"aliases": [
|
||||||
|
"Santos",
|
||||||
|
"Santos FC",
|
||||||
|
"Santos Futebol Clube",
|
||||||
|
"Peixe"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_club",
|
||||||
|
"relation_type": "RIVAL_OF",
|
||||||
|
"weight": 0.68,
|
||||||
|
"confidence": 0.85,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"derby",
|
||||||
|
"rival",
|
||||||
|
"clássico",
|
||||||
|
"San-São",
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Brasileirão"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_libertadores",
|
||||||
|
"name": "Copa Libertadores da América",
|
||||||
|
"canonical_name": "Copa Libertadores da América",
|
||||||
|
"aliases": [
|
||||||
|
"Libertadores",
|
||||||
|
"Copa Libertadores",
|
||||||
|
"Copa Libertadores da América",
|
||||||
|
"CONMEBOL Libertadores"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_competition",
|
||||||
|
"relation_type": "PLAYS_COMPETITION",
|
||||||
|
"weight": 0.65,
|
||||||
|
"confidence": 0.85,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"competition",
|
||||||
|
"continental",
|
||||||
|
"CONMEBOL",
|
||||||
|
"title",
|
||||||
|
"SPFC"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_campeonato_brasileiro",
|
||||||
|
"name": "Campeonato Brasileiro Série A",
|
||||||
|
"canonical_name": "Campeonato Brasileiro Série A",
|
||||||
|
"aliases": [
|
||||||
|
"Campeonato Brasileiro",
|
||||||
|
"Brasileirão",
|
||||||
|
"Brasileiro Série A",
|
||||||
|
"Série A"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_competition",
|
||||||
|
"relation_type": "PLAYS_COMPETITION",
|
||||||
|
"weight": 0.62,
|
||||||
|
"confidence": 0.85,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"competition",
|
||||||
|
"Brazil",
|
||||||
|
"league",
|
||||||
|
"club",
|
||||||
|
"table",
|
||||||
|
"match"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_copa_do_brasil",
|
||||||
|
"name": "Copa do Brasil",
|
||||||
|
"canonical_name": "Copa do Brasil",
|
||||||
|
"aliases": [
|
||||||
|
"Copa do Brasil"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_competition",
|
||||||
|
"relation_type": "PLAYS_COMPETITION",
|
||||||
|
"weight": 0.58,
|
||||||
|
"confidence": 0.8,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"competition",
|
||||||
|
"Brazil",
|
||||||
|
"knockout",
|
||||||
|
"club",
|
||||||
|
"match"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"entity_id": "ecp_campeonato_paulista",
|
||||||
|
"name": "Campeonato Paulista",
|
||||||
|
"canonical_name": "Campeonato Paulista",
|
||||||
|
"aliases": [
|
||||||
|
"Campeonato Paulista",
|
||||||
|
"Paulistão",
|
||||||
|
"Paulista"
|
||||||
|
],
|
||||||
|
"entity_type": "sports_competition",
|
||||||
|
"relation_type": "PLAYS_COMPETITION",
|
||||||
|
"weight": 0.58,
|
||||||
|
"confidence": 0.8,
|
||||||
|
"scope_terms": [
|
||||||
|
"football",
|
||||||
|
"competition",
|
||||||
|
"São Paulo state",
|
||||||
|
"club",
|
||||||
|
"derby"
|
||||||
|
],
|
||||||
|
"classification_effect": "contextual_inherent_candidate"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"classification_profile": {
|
||||||
|
"labels": [
|
||||||
|
"DIRECT_INHERENT",
|
||||||
|
"CONTEXTUAL_INHERENT",
|
||||||
|
"TANGENTIAL",
|
||||||
|
"NOT_RELATED",
|
||||||
|
"UNCERTAIN"
|
||||||
|
],
|
||||||
|
"direct_inherent": {
|
||||||
|
"description": "Content is directly about São Paulo Futebol Clube.",
|
||||||
|
"strong_positive_signals": [
|
||||||
|
"canonical_name",
|
||||||
|
"strong_alias",
|
||||||
|
"SPFC",
|
||||||
|
"Tricolor Paulista",
|
||||||
|
"official_domain",
|
||||||
|
"club-specific anchor"
|
||||||
|
],
|
||||||
|
"minimum_expected_context": [
|
||||||
|
"football",
|
||||||
|
"club",
|
||||||
|
"match",
|
||||||
|
"player",
|
||||||
|
"coach",
|
||||||
|
"stadium",
|
||||||
|
"competition"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"contextual_inherent": {
|
||||||
|
"description": "Content is mainly about a strongly related entity, but relevant to São Paulo Futebol Clube.",
|
||||||
|
"required_signals": [
|
||||||
|
"related_entity_match",
|
||||||
|
"relation_weight >= 0.65",
|
||||||
|
"sports_or_football_context"
|
||||||
|
],
|
||||||
|
"examples": [
|
||||||
|
"article about Rogério Ceni as São Paulo idol or coach",
|
||||||
|
"article about Raí in relation to São Paulo history",
|
||||||
|
"article about MorumBIS as São Paulo home venue",
|
||||||
|
"article about São Paulo rivalry with Corinthians or Palmeiras"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"tangential": {
|
||||||
|
"description": "Content mentions São Paulo FC or a related term, but the entity is not meaningfully part of the subject.",
|
||||||
|
"examples": [
|
||||||
|
"generic Brasileirão table with São Paulo only listed among many clubs",
|
||||||
|
"article about another club that only mentions upcoming match against São Paulo"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"not_related": {
|
||||||
|
"description": "Content is about another São Paulo meaning or unrelated subject.",
|
||||||
|
"strong_negative_signals": [
|
||||||
|
"city_or_state_context",
|
||||||
|
"government_context",
|
||||||
|
"weather_or_traffic_context",
|
||||||
|
"university_context",
|
||||||
|
"finance_context",
|
||||||
|
"fashion_or_event_context"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"uncertain": {
|
||||||
|
"description": "Evidence is weak, conflicting, or insufficient.",
|
||||||
|
"fallback_recommendation": "send_to_llm_judge"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"scoring_hints": {
|
||||||
|
"strong_alias_weight": 0.45,
|
||||||
|
"ambiguous_alias_weight": 0.12,
|
||||||
|
"strong_anchor_weight": 0.25,
|
||||||
|
"medium_anchor_weight": 0.15,
|
||||||
|
"weak_anchor_weight": 0.05,
|
||||||
|
"negative_anchor_penalty": 0.4,
|
||||||
|
"related_entity_weight_multiplier": 0.5,
|
||||||
|
"official_domain_weight": 0.5,
|
||||||
|
"direct_threshold": 0.75,
|
||||||
|
"contextual_threshold": 0.65,
|
||||||
|
"tangential_threshold": 0.35,
|
||||||
|
"uncertain_band": [
|
||||||
|
0.45,
|
||||||
|
0.65
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"output_contract": {
|
||||||
|
"label": "DIRECT_INHERENT | CONTEXTUAL_INHERENT | TANGENTIAL | NOT_RELATED | UNCERTAIN",
|
||||||
|
"confidence": "number from 0.0 to 1.0",
|
||||||
|
"matched_signals": "array of matched positive signals",
|
||||||
|
"negative_signals": "array of matched negative signals",
|
||||||
|
"related_entity_matches": "array of matched related entities",
|
||||||
|
"explanation": "short human-readable explanation",
|
||||||
|
"needs_llm_fallback": "boolean"
|
||||||
|
},
|
||||||
|
"quality": {
|
||||||
|
"profile_confidence": 0.86,
|
||||||
|
"intended_use": "local deterministic NLP POC fixture",
|
||||||
|
"not_intended_for": [
|
||||||
|
"final production ECP schema",
|
||||||
|
"canonical graph storage",
|
||||||
|
"legal/compliance decisioning",
|
||||||
|
"fully automated publishing without review"
|
||||||
|
],
|
||||||
|
"known_limitations": [
|
||||||
|
"Current player anchors may become stale.",
|
||||||
|
"The alias 'São Paulo' is highly ambiguous and must never be enough alone for DIRECT_INHERENT.",
|
||||||
|
"Generic football terms are weak signals only.",
|
||||||
|
"Related entities should usually produce CONTEXTUAL_INHERENT, not DIRECT_INHERENT.",
|
||||||
|
"This profile is optimized for Portuguese content first."
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"sources": [
|
||||||
|
{
|
||||||
|
"type": "wikidata",
|
||||||
|
"url": "https://www.wikidata.org/wiki/Q38568",
|
||||||
|
"supports": [
|
||||||
|
"qid",
|
||||||
|
"canonical identity",
|
||||||
|
"entity type",
|
||||||
|
"inception"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "official_site",
|
||||||
|
"url": "https://www.saopaulofc.net/",
|
||||||
|
"supports": [
|
||||||
|
"official domain",
|
||||||
|
"club identity",
|
||||||
|
"sports sections"
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "official_site",
|
||||||
|
"url": "https://www.saopaulofc.net/institucional/sobre-o-sao-paulo-fc/",
|
||||||
|
"supports": [
|
||||||
|
"history",
|
||||||
|
"founded year",
|
||||||
|
"club narrative",
|
||||||
|
"MorumBIS",
|
||||||
|
"historic figures"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"metadata": {
|
||||||
|
"created_for": "ecp-nlp-classifier-poc",
|
||||||
|
"created_at": "2026-08-20",
|
||||||
|
"graph_snapshot": "spfc-graph-snapshot-v1",
|
||||||
|
"review_status": "example_ready"
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -39,15 +39,15 @@
|
|||||||
"37": "Plan Setup",
|
"37": "Plan Setup",
|
||||||
"38": "Task Setup",
|
"38": "Task Setup",
|
||||||
"39": "Graphify Workflows",
|
"39": "Graphify Workflows",
|
||||||
"40": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
"40": "main",
|
||||||
"41": "1. Technical Decisions & Tradeoffs",
|
"41": "1. Technical Decisions & Tradeoffs",
|
||||||
"42": "1. Input Schemas",
|
"42": "1. Input Schemas",
|
||||||
"43": "2. Basic CLI Usage Examples",
|
"43": "2. Basic CLI Usage Examples",
|
||||||
"44": "2. Standard Streams & Exit Codes",
|
"44": "2. Standard Streams & Exit Codes",
|
||||||
"45": "ECPSnapshot",
|
"45": "ClassificationResult",
|
||||||
"46": "test_models.py",
|
"46": "InherenceClassifier",
|
||||||
"47": "detect_language",
|
"47": "detect_language",
|
||||||
"48": "main",
|
"48": "test_models.py",
|
||||||
"49": "content_northvolt_de.md",
|
"49": "content_northvolt_de.md",
|
||||||
"50": "content_presal_pt.md",
|
"50": "content_presal_pt.md",
|
||||||
"51": "content_tangential_es.md",
|
"51": "content_tangential_es.md",
|
||||||
@@ -78,5 +78,29 @@
|
|||||||
"76": "pt/not_related.md",
|
"76": "pt/not_related.md",
|
||||||
"77": "pt/tangential.md",
|
"77": "pt/tangential.md",
|
||||||
"78": "tests/__init__.py",
|
"78": "tests/__init__.py",
|
||||||
"79": "text-nlp-classifier"
|
"79": "text-nlp-classifier",
|
||||||
|
"80": "get_hl_gl_ceid",
|
||||||
|
"81": "Extrator de Notícias do Google News — Guia Completo de Funcionamento",
|
||||||
|
"82": "extract_google_news.py",
|
||||||
|
"83": "ExtractionResult",
|
||||||
|
"84": "test_extract_google_news.py",
|
||||||
|
"85": "Implementation Tasks: Google News Headlines Extractor",
|
||||||
|
"86": "Feature Specification: Google News Headlines Extractor",
|
||||||
|
"87": "2. Cenários Práticos de Uso",
|
||||||
|
"88": "Implementation Plan: Google News Headlines Extractor",
|
||||||
|
"89": "scripts/__init__.py",
|
||||||
|
"90": "SearchQuery",
|
||||||
|
"91": "1. Technical Decisions & Tradeoffs",
|
||||||
|
"92": "General Readiness Checklist: Google News Headlines Extractor",
|
||||||
|
"93": "1. Entidades de Domínio & DTOs",
|
||||||
|
"94": "Specification Quality Checklist: Google News Headlines Extractor",
|
||||||
|
"95": "CLI Contract: Google News Headlines Extractor",
|
||||||
|
"96": "readiness.md",
|
||||||
|
"97": "🧠 TextNLPClassifierApp",
|
||||||
|
"98": "build_parser",
|
||||||
|
"99": "sample_rss_xml",
|
||||||
|
"100": "classifier.py",
|
||||||
|
"101": "ECPSnapshot",
|
||||||
|
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||||
|
"103": "main"
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1 +1 @@
|
|||||||
{"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "b493e66e33005daa", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "6aa00d5a83295f11", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "635b0327b73603cf", "46": "aebae10575e16ffb", "47": "7f0025cb1cabb27d", "48": "cbb35e88e6b7e1ac", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6"}
|
{"0": "36bdb6f09c457f7c", "1": "8c5bf6244cf710c6", "2": "efbcc9c62a3ee78b", "3": "8599153989b07faa", "4": "b5952a1f7fee9f20", "5": "b493e66e33005daa", "6": "80f79e9e2011a3e3", "7": "4654167fd211d027", "8": "50acfa00fe353440", "9": "c6d2f770737823f1", "10": "44f2ca451aea24be", "11": "feaac5ab67a8c17a", "12": "b71bd92e5edbf2e0", "13": "219d65ba6d2689e4", "14": "8e30bb8112fd02d1", "15": "03906ab80b99db85", "16": "5d51c60ba1bc2be0", "17": "a1da914f522dcd21", "18": "fbad840891b90569", "19": "0686ff2d6fe29fb3", "20": "060baa9e1924b465", "21": "a5c8f2c3080b8243", "22": "0d76852f1d29eeb1", "23": "6ff68619f2d72924", "24": "3da11675eee7ec46", "25": "a6696589e9556f97", "26": "6c752999e8a4d4b6", "27": "2d4e13ea2111d750", "28": "4b60cb0ee1ac186a", "29": "f56fbca9bb8235ec", "30": "c7beed940704509f", "31": "38be2d254fb31ae8", "32": "ee5596fcf7e7c0b3", "33": "e4d4e0a440bc599f", "34": "c897e49c001acdae", "35": "3aad272a2cf5d495", "36": "0a197439d306b956", "37": "f43acf5c8b1329af", "38": "6775efafc9b33338", "39": "8176a164778526f9", "40": "62a40bb7296050db", "41": "0322ff824966a4d8", "42": "784c9e3d336a7f53", "43": "4b8bb6c3f7b64856", "44": "18c0ff3e6225bcb2", "45": "0172259097d0974e", "46": "d2ae82bc1907fef3", "47": "7be46d99756bf41e", "48": "a94841cd5849d3f5", "49": "0d0f9f015921feef", "50": "8d0c81e5ca23e9a6", "51": "f79963571b9c15ee", "52": "5935824c825606cb", "53": "9685f9cbe158e50b", "54": "3d5ab759f350bc79", "55": "d549f24931a990e9", "56": "3cc031dcb648797c", "57": "a0ab88e6c629251d", "58": "76bd6412e2a22ecd", "59": "54827845564490c9", "60": "0a9736c416c0c6b9", "61": "77358620ac528153", "62": "3b0c585df09df48a", "63": "7e78cd3b28828c20", "64": "1c0c958231735f61", "65": "60b0f81225f62f69", "66": "920754c65cc94b88", "67": "df911472140a9b94", "68": "8e17bc11bcea91b9", "69": "7e905b75e4f28b95", "70": "a28424eca5d36c55", "71": "2cdb53d5b6051ab6", "72": "e42fbd3dc744e730", "73": "7fe2cac980de160c", "74": "2b1343a6a9db1487", "75": "54a1bb232f1d4ceb", "76": "442ba11d31ec0e0a", "77": "852a25b8b95bf8d1", "78": "1810ab370b9cd608", "79": "0fc5dca02a3f02f6", "80": "10ea6ef8e60efb58", "81": "a38f84ae3d895236", "82": "5efe702565242e04", "83": "f74492db7d32ba1e", "84": "c0d4860b83f7cb84", "85": "f8bfd0cfe9e8b478", "86": "410d15a346bd5894", "87": "6b41d288cfd834ab", "88": "5aa6db96312a8811", "89": "80225792bb62ba04", "90": "e553a45aca39b563", "91": "d4579c5b7aa2742a", "92": "7b9ba7c3bff11361", "93": "71cd9c1fa4a857f0", "94": "34cd980be3c32d21", "95": "970093453f3b7d90", "96": "9e96780a2b7c4bd6", "97": "3befc42bd6078583", "98": "32f0d209b4ebd09b", "99": "8968e9e7d55afcbe", "100": "c14800d0ab2e5026", "101": "d7e7752d280777b1", "102": "6aa00d5a83295f11", "103": "cbb35e88e6b7e1ac"}
|
||||||
@@ -4,7 +4,7 @@
|
|||||||
"2": "SpecKit Utilities",
|
"2": "SpecKit Utilities",
|
||||||
"3": "Graphify Commands",
|
"3": "Graphify Commands",
|
||||||
"4": "speckit-analyze/SKILL.md",
|
"4": "speckit-analyze/SKILL.md",
|
||||||
"5": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
"5": "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||||
"6": "Feature Specification Template",
|
"6": "Feature Specification Template",
|
||||||
"7": "Graphify Rules",
|
"7": "Graphify Rules",
|
||||||
"8": "Implementation Planning",
|
"8": "Implementation Planning",
|
||||||
@@ -39,15 +39,15 @@
|
|||||||
"37": "Plan Setup",
|
"37": "Plan Setup",
|
||||||
"38": "Task Setup",
|
"38": "Task Setup",
|
||||||
"39": "Graphify Workflows",
|
"39": "Graphify Workflows",
|
||||||
"40": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
"40": "main",
|
||||||
"41": "1. Technical Decisions & Tradeoffs",
|
"41": "1. Technical Decisions & Tradeoffs",
|
||||||
"42": "1. Input Schemas",
|
"42": "1. Input Schemas",
|
||||||
"43": "2. Basic CLI Usage Examples",
|
"43": "2. Basic CLI Usage Examples",
|
||||||
"44": "2. Standard Streams & Exit Codes",
|
"44": "2. Standard Streams & Exit Codes",
|
||||||
"45": "ECPSnapshot",
|
"45": "ECPSnapshot",
|
||||||
"46": "classifier.py",
|
"46": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||||
"47": "detect_language",
|
"47": "detect_language",
|
||||||
"48": "main",
|
"48": "test_models.py",
|
||||||
"49": "content_northvolt_de.md",
|
"49": "content_northvolt_de.md",
|
||||||
"50": "content_presal_pt.md",
|
"50": "content_presal_pt.md",
|
||||||
"51": "content_tangential_es.md",
|
"51": "content_tangential_es.md",
|
||||||
@@ -78,5 +78,24 @@
|
|||||||
"76": "pt/not_related.md",
|
"76": "pt/not_related.md",
|
||||||
"77": "pt/tangential.md",
|
"77": "pt/tangential.md",
|
||||||
"78": "tests/__init__.py",
|
"78": "tests/__init__.py",
|
||||||
"79": "text-nlp-classifier"
|
"79": "text-nlp-classifier",
|
||||||
|
"80": "get_hl_gl_ceid",
|
||||||
|
"81": "Extrator de Notícias do Google News — Guia Completo de Funcionamento",
|
||||||
|
"82": "extract_google_news.py",
|
||||||
|
"83": "ExtractionResult",
|
||||||
|
"84": "test_extract_google_news.py",
|
||||||
|
"85": "Implementation Tasks: Google News Headlines Extractor",
|
||||||
|
"86": "Feature Specification: Google News Headlines Extractor",
|
||||||
|
"87": "2. Cenários Práticos de Uso",
|
||||||
|
"88": "Implementation Plan: Google News Headlines Extractor",
|
||||||
|
"89": "scripts/__init__.py",
|
||||||
|
"90": "SearchQuery",
|
||||||
|
"91": "1. Technical Decisions & Tradeoffs",
|
||||||
|
"92": "General Readiness Checklist: Google News Headlines Extractor",
|
||||||
|
"93": "1. Entidades de Domínio & DTOs",
|
||||||
|
"94": "Specification Quality Checklist: Google News Headlines Extractor",
|
||||||
|
"95": "CLI Contract: Google News Headlines Extractor",
|
||||||
|
"96": "readiness.md",
|
||||||
|
"98": "build_parser",
|
||||||
|
"99": "sample_rss_xml"
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,16 +1,16 @@
|
|||||||
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
||||||
|
|
||||||
## Corpus Check
|
## Corpus Check
|
||||||
- 132 files · ~53,988 words
|
- 147 files · ~65,826 words
|
||||||
- Verdict: corpus is large enough that graph structure adds value.
|
- Verdict: corpus is large enough that graph structure adds value.
|
||||||
|
|
||||||
## Summary
|
## Summary
|
||||||
- 579 nodes · 674 edges · 80 communities (43 shown, 37 thin omitted)
|
- 775 nodes · 931 edges · 99 communities (61 shown, 38 thin omitted)
|
||||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 28 edges (avg confidence: 0.95)
|
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
|
||||||
- Token cost: 0 input · 0 output
|
- Token cost: 0 input · 0 output
|
||||||
|
|
||||||
## Graph Freshness
|
## Graph Freshness
|
||||||
- Built from commit: `d371b81a`
|
- Built from commit: `67cc40f9`
|
||||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||||
- Run `graphify update .` after code changes (no API cost).
|
- Run `graphify update .` after code changes (no API cost).
|
||||||
|
|
||||||
@@ -20,7 +20,7 @@
|
|||||||
- SpecKit Utilities
|
- SpecKit Utilities
|
||||||
- Graphify Commands
|
- Graphify Commands
|
||||||
- speckit-analyze/SKILL.md
|
- speckit-analyze/SKILL.md
|
||||||
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||||
- Feature Specification Template
|
- Feature Specification Template
|
||||||
- Graphify Rules
|
- Graphify Rules
|
||||||
- Implementation Planning
|
- Implementation Planning
|
||||||
@@ -51,15 +51,15 @@
|
|||||||
- Media Transcription
|
- Media Transcription
|
||||||
- Extraction Specification
|
- Extraction Specification
|
||||||
- Graphify Workflows
|
- Graphify Workflows
|
||||||
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
- main
|
||||||
- 1. Technical Decisions & Tradeoffs
|
- 1. Technical Decisions & Tradeoffs
|
||||||
- 1. Input Schemas
|
- 1. Input Schemas
|
||||||
- 2. Basic CLI Usage Examples
|
- 2. Basic CLI Usage Examples
|
||||||
- 2. Standard Streams & Exit Codes
|
- 2. Standard Streams & Exit Codes
|
||||||
- ECPSnapshot
|
- ECPSnapshot
|
||||||
- classifier.py
|
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
||||||
- detect_language
|
- detect_language
|
||||||
- main
|
- test_models.py
|
||||||
- content_northvolt_de.md
|
- content_northvolt_de.md
|
||||||
- content_presal_pt.md
|
- content_presal_pt.md
|
||||||
- content_tangential_es.md
|
- content_tangential_es.md
|
||||||
@@ -91,35 +91,53 @@
|
|||||||
- pt/tangential.md
|
- pt/tangential.md
|
||||||
- tests/__init__.py
|
- tests/__init__.py
|
||||||
- text-nlp-classifier
|
- text-nlp-classifier
|
||||||
|
- get_hl_gl_ceid
|
||||||
|
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||||
|
- extract_google_news.py
|
||||||
|
- ExtractionResult
|
||||||
|
- test_extract_google_news.py
|
||||||
|
- Implementation Tasks: Google News Headlines Extractor
|
||||||
|
- Feature Specification: Google News Headlines Extractor
|
||||||
|
- 2. Cenários Práticos de Uso
|
||||||
|
- Implementation Plan: Google News Headlines Extractor
|
||||||
|
- scripts/__init__.py
|
||||||
|
- SearchQuery
|
||||||
|
- 1. Technical Decisions & Tradeoffs
|
||||||
|
- General Readiness Checklist: Google News Headlines Extractor
|
||||||
|
- 1. Entidades de Domínio & DTOs
|
||||||
|
- Specification Quality Checklist: Google News Headlines Extractor
|
||||||
|
- CLI Contract: Google News Headlines Extractor
|
||||||
|
- build_parser
|
||||||
|
- sample_rss_xml
|
||||||
|
|
||||||
## God Nodes (most connected - your core abstractions)
|
## God Nodes (most connected - your core abstractions)
|
||||||
1. `ECPSnapshot` - 27 edges
|
1. `ECPSnapshot` - 31 edges
|
||||||
2. `InherenceClassifier` - 21 edges
|
2. `InherenceClassifier` - 25 edges
|
||||||
3. `ClassificationResult` - 17 edges
|
3. `DecisionCategory` - 18 edges
|
||||||
4. `LocalEmbeddingsAdapter` - 14 edges
|
4. `ClassificationResult` - 17 edges
|
||||||
5. `LLMFallbackAdapter` - 14 edges
|
5. `main()` - 14 edges
|
||||||
6. `detect_language()` - 14 edges
|
6. `LocalEmbeddingsAdapter` - 14 edges
|
||||||
7. `DecisionCategory` - 14 edges
|
7. `LLMFallbackAdapter` - 14 edges
|
||||||
8. `main()` - 13 edges
|
8. `detect_language()` - 14 edges
|
||||||
9. `Tasks: [FEATURE NAME]` - 13 edges
|
9. `Tasks: [FEATURE NAME]` - 13 edges
|
||||||
10. `BaseNLPAdapter` - 12 edges
|
10. `SearchQuery` - 12 edges
|
||||||
|
|
||||||
## Surprising Connections (you probably didn't know these)
|
## Surprising Connections (you probably didn't know these)
|
||||||
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
||||||
classify.py → src/models.py
|
classify.py → src/models.py
|
||||||
- `main()` --uses--> `ErrorCode` [INFERRED]
|
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
|
||||||
classify.py → src/models.py
|
tests/test_extract_google_news.py → scripts/extract_google_news.py
|
||||||
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
||||||
tests/test_models.py → src/models.py
|
tests/test_models.py → src/models.py
|
||||||
- `petrobras_ecp()` --uses--> `ECPSnapshot` [INFERRED]
|
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
|
||||||
tests/test_classifier.py → src/models.py
|
tests/test_models.py → src/models.py
|
||||||
- `emit_error()` --uses--> `ErrorCode` [INFERRED]
|
- `test_ecp_snapshot_missing_required()` --uses--> `ECPSnapshot` [INFERRED]
|
||||||
classify.py → src/models.py
|
tests/test_models.py → src/models.py
|
||||||
|
|
||||||
## Import Cycles
|
## Import Cycles
|
||||||
- None detected.
|
- None detected.
|
||||||
|
|
||||||
## Communities (80 total, 37 thin omitted)
|
## Communities (99 total, 38 thin omitted)
|
||||||
|
|
||||||
### Community 0 - "Task Planning"
|
### Community 0 - "Task Planning"
|
||||||
Cohesion: 0.07
|
Cohesion: 0.07
|
||||||
@@ -141,9 +159,9 @@ Nodes (24): For /graphify add and --watch, For /graphify query, For the commit h
|
|||||||
Cohesion: 0.08
|
Cohesion: 0.08
|
||||||
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
|
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
|
||||||
|
|
||||||
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
### Community 5 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||||
Cohesion: 0.06
|
Cohesion: 0.05
|
||||||
Nodes (33): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+25 more)
|
Nodes (34): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+26 more)
|
||||||
|
|
||||||
### Community 6 - "Feature Specification Template"
|
### Community 6 - "Feature Specification Template"
|
||||||
Cohesion: 0.15
|
Cohesion: 0.15
|
||||||
@@ -241,9 +259,9 @@ Nodes (3): For --cluster-only, For --update (incremental re-extraction), graphif
|
|||||||
Cohesion: 0.50
|
Cohesion: 0.50
|
||||||
Nodes (3): Boundaries, Output, Scan
|
Nodes (3): Boundaries, Output, Scan
|
||||||
|
|
||||||
### Community 40 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
### Community 40 - "main"
|
||||||
Cohesion: 0.14
|
Cohesion: 0.19
|
||||||
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
|
||||||
|
|
||||||
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
||||||
Cohesion: 0.22
|
Cohesion: 0.22
|
||||||
@@ -262,40 +280,108 @@ Cohesion: 0.29
|
|||||||
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
||||||
|
|
||||||
### Community 45 - "ECPSnapshot"
|
### Community 45 - "ECPSnapshot"
|
||||||
Cohesion: 0.07
|
Cohesion: 0.06
|
||||||
Nodes (26): ABC, Any, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+18 more)
|
Nodes (55): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+47 more)
|
||||||
|
|
||||||
### Community 46 - "classifier.py"
|
### Community 46 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||||
Cohesion: 0.09
|
Cohesion: 0.15
|
||||||
Nodes (39): emit_error(), Enum, count_phrase_occurrences(), InherenceClassifier, match_phrase_in_text(), Core deterministic classification engine (Tier 1 core)., Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text. (+31 more)
|
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
|
||||||
|
|
||||||
### Community 47 - "detect_language"
|
### Community 47 - "detect_language"
|
||||||
Cohesion: 0.19
|
Cohesion: 0.14
|
||||||
Nodes (16): detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization., Normalize text by converting to lowercase and stripping combining diacritical…, Tokenize text into lowercase alphanumeric words., Detect the ISO-639-1 language code of text among supported languages (pt, en,…, Unit tests for language detection and text normalization. (+8 more)
|
Nodes (21): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+13 more)
|
||||||
|
|
||||||
### Community 48 - "main"
|
### Community 48 - "test_models.py"
|
||||||
Cohesion: 0.31
|
Cohesion: 0.09
|
||||||
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
|
Nodes (29): emit_error(), main(), parse_args(), Namespace, ClassificationError, ErrorCode, Any, extract_evidence_snippets() (+21 more)
|
||||||
|
|
||||||
|
### Community 80 - "get_hl_gl_ceid"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
|
||||||
|
|
||||||
|
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
|
||||||
|
Cohesion: 0.08
|
||||||
|
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
|
||||||
|
|
||||||
|
### Community 82 - "extract_google_news.py"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
|
||||||
|
|
||||||
|
### Community 83 - "ExtractionResult"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
|
||||||
|
|
||||||
|
### Community 84 - "test_extract_google_news.py"
|
||||||
|
Cohesion: 0.21
|
||||||
|
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
|
||||||
|
|
||||||
|
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.14
|
||||||
|
Nodes (14): Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Tasks: Google News Headlines Extractor, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Extração Básica de Notícias por Assunto e Idioma (Priority: P1) 🌟 MVP, Phase 4: User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2) (+6 more)
|
||||||
|
|
||||||
|
### Community 86 - "Feature Specification: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.18
|
||||||
|
Nodes (11): Clarifications, Edge Cases, Feature Specification: Google News Headlines Extractor, Functional Requirements, Requirements *(mandatory)*, Session 2026-08-20, Success Criteria *(mandatory)*, User Scenarios & Testing *(mandatory)* (+3 more)
|
||||||
|
|
||||||
|
### Community 87 - "2. Cenários Práticos de Uso"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (9): 1. Pré-requisitos e Instalação, 2. Cenários Práticos de Uso, 3. Validação dos Testes Automatizados e Linter, Cenário 1: River Plate — Argentina (Espanhol / 2 Páginas / Salvar em Arquivo), Cenário 2: Cruzeiro — Brasil (Português / Formatado no Terminal), Cenário 3: Fórmula 1 — Reino Unido (Inglês), Cenário 4: Integração em Pipeline com `jq` (Modo Silencioso), Cenário 5: Extração Rápida com Links Brutos (Sem Resolução de URLs) (+1 more)
|
||||||
|
|
||||||
|
### Community 88 - "Implementation Plan: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (7): Architecture & Pipeline, Documentation (this feature), Implementation Plan: Google News Headlines Extractor, Project Structure, Source Code, Summary, Technical Context
|
||||||
|
|
||||||
|
### Community 90 - "SearchQuery"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (6): Value Object com parâmetros de busca validados., SearchQuery, Valida a consolidação do ExtractionResult a partir da busca mockada com URLs…, Valida as regras de negócio e limites de SearchQuery., test_extract_google_news_orchestration_mocked(), test_search_query_validation()
|
||||||
|
|
||||||
|
### Community 91 - "1. Technical Decisions & Tradeoffs"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (7): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Requisição e Scraping com `foxcape` em Modo Headless, Decision 2: Endpoint RSS do Google News vs. Scraping de DOM, Decision 3: Mapeamento de Idioma e Locale (`hl`, `gl`, `ceid`), Decision 4: Resolução de URLs do Google News via `googlenewsdecoder`, Decision 5: Logging em Tempo Real no `stderr` e Segregação de Streams, Research: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 92 - "General Readiness Checklist: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (7): CLI Interface & Parameter Contracts, Data Sanitization & Article Extraction, Error Handling & Edge Cases, General Readiness Checklist: Google News Headlines Extractor, Non-Functional & Operational Readiness, Notes, Scraping Engine & Feed Mapping
|
||||||
|
|
||||||
|
### Community 93 - "1. Entidades de Domínio & DTOs"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (6): 1.1 SearchQuery (Parâmetros da Busca), 1.2 NewsArticle (Item de Notícia), 1.3 ExtractionResult (Saída Estruturada Consolidada), 1. Entidades de Domínio & DTOs, 2. Esquema JSON de Saída, Data Model: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 94 - "Specification Quality Checklist: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.33
|
||||||
|
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 95 - "CLI Contract: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.33
|
||||||
|
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
|
||||||
|
|
||||||
|
### Community 98 - "build_parser"
|
||||||
|
Cohesion: 0.67
|
||||||
|
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
|
||||||
|
|
||||||
|
### Community 99 - "sample_rss_xml"
|
||||||
|
Cohesion: 0.67
|
||||||
|
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
|
||||||
|
|
||||||
## Knowledge Gaps
|
## Knowledge Gaps
|
||||||
- **291 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+286 more)
|
- **361 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+356 more)
|
||||||
These have ≤1 connection - possible missing edges or undocumented components.
|
These have ≤1 connection - possible missing edges or undocumented components.
|
||||||
- **37 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||||
|
|
||||||
## Suggested Questions
|
## Suggested Questions
|
||||||
_Questions this graph is uniquely positioned to answer:_
|
_Questions this graph is uniquely positioned to answer:_
|
||||||
|
|
||||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `classifier.py`?**
|
- **Why does `main()` connect `test_models.py` to `main`, `ECPSnapshot`?**
|
||||||
_High betweenness centrality (0.011) - this node is a cross-community bridge._
|
_High betweenness centrality (0.031) - this node is a cross-community bridge._
|
||||||
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?**
|
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `test_models.py`, `detect_language`?**
|
||||||
_High betweenness centrality (0.007) - this node is a cross-community bridge._
|
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||||
- **Why does `InherenceClassifier` connect `classifier.py` to `main`, `ECPSnapshot`?**
|
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
|
||||||
_High betweenness centrality (0.006) - this node is a cross-community bridge._
|
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||||
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
||||||
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
||||||
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
|
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
|
||||||
|
- **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?**
|
||||||
|
_`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
|
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
|
||||||
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
|
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Are the 3 inferred relationships involving `LocalEmbeddingsAdapter` (e.g. with `ClassificationResult` and `ECPSnapshot`) actually correct?**
|
|
||||||
_`LocalEmbeddingsAdapter` has 3 INFERRED edges - model-reasoned connections that need verification._
|
|
||||||
+6111
-912
File diff suppressed because it is too large
Load Diff
@@ -294,15 +294,15 @@
|
|||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"classify.py": {
|
"classify.py": {
|
||||||
"mtime": 1787195933.7872643,
|
"mtime": 1787197110.3753626,
|
||||||
"seen": 1787196463.095295,
|
"seen": 1787197159.9828603,
|
||||||
"ast_hash": "5fa90ecfb87bdf15ee99165a3be7c63b",
|
"ast_hash": "d09a35a5e25d42f6ecc537d5b67bef8e",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"pyproject.toml": {
|
"pyproject.toml": {
|
||||||
"mtime": 1787195827.3305018,
|
"mtime": 1787234982.768605,
|
||||||
"seen": 1787196463.0953014,
|
"seen": 1787235020.6850271,
|
||||||
"ast_hash": "f2503e96d08f5a4e41b13352e81709c6",
|
"ast_hash": "26f5f979658067bf447f375f044a8f21",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/__init__.py": {
|
"src/__init__.py": {
|
||||||
@@ -336,9 +336,9 @@
|
|||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/classifier.py": {
|
"src/classifier.py": {
|
||||||
"mtime": 1787196413.3774571,
|
"mtime": 1787197085.0250447,
|
||||||
"seen": 1787196463.095319,
|
"seen": 1787197159.985785,
|
||||||
"ast_hash": "d7bab620ae29badf0d3779fab44552bc",
|
"ast_hash": "a2e70f968109d213fdab0ae81ac819ff",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/language.py": {
|
"src/language.py": {
|
||||||
@@ -420,9 +420,9 @@
|
|||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"requirements.txt": {
|
"requirements.txt": {
|
||||||
"mtime": 1787195817.4782183,
|
"mtime": 1787236111.876247,
|
||||||
"seen": 1787196463.1020837,
|
"seen": 1787236262.0105975,
|
||||||
"ast_hash": "5612073a1e7034d7c25765379baa5da4",
|
"ast_hash": "9fb7e3f8ecb04ff30d650c772f509557",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"tests/fixtures/benchmark_24/de/contextual.md": {
|
"tests/fixtures/benchmark_24/de/contextual.md": {
|
||||||
@@ -568,5 +568,89 @@
|
|||||||
"seen": 1787196463.1031141,
|
"seen": 1787196463.1031141,
|
||||||
"ast_hash": "e05ab20a5190cfb5b9d41da64d5273cf",
|
"ast_hash": "e05ab20a5190cfb5b9d41da64d5273cf",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_adversarial.py": {
|
||||||
|
"mtime": 1787197141.2340336,
|
||||||
|
"seen": 1787197159.9881344,
|
||||||
|
"ast_hash": "5e8daf517ee32c261617d96264ef0473",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/extract_google_news.py": {
|
||||||
|
"mtime": 1787236453.5231855,
|
||||||
|
"seen": 1787236509.1147037,
|
||||||
|
"ast_hash": "4216de490a87378d21dab707a3166679",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_extract_google_news.py": {
|
||||||
|
"mtime": 1787236301.513448,
|
||||||
|
"seen": 1787236353.5582018,
|
||||||
|
"ast_hash": "1fc9a108a8abfecf1d37c8642a885e9e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/googlenews_extractor_guia_completo.md": {
|
||||||
|
"mtime": 1787231605.0830677,
|
||||||
|
"seen": 1787234772.8112028,
|
||||||
|
"ast_hash": "59db8e652000ba6087a00289a0ce0959",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/checklists/readiness.md": {
|
||||||
|
"mtime": 1787234305.932687,
|
||||||
|
"seen": 1787234772.8123577,
|
||||||
|
"ast_hash": "a1200a6959d41e256a73403dd5c4e1f6",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/checklists/requirements.md": {
|
||||||
|
"mtime": 1787232386.6763985,
|
||||||
|
"seen": 1787234772.812359,
|
||||||
|
"ast_hash": "39cd86af76c9f74baa2984418870b677",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/contracts/cli_contract.md": {
|
||||||
|
"mtime": 1787236744.3139226,
|
||||||
|
"seen": 1787236817.4979818,
|
||||||
|
"ast_hash": "76a61ad31df989244edc5e739081bf7b",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/data-model.md": {
|
||||||
|
"mtime": 1787234040.8134317,
|
||||||
|
"seen": 1787234772.8123617,
|
||||||
|
"ast_hash": "e8380c98d2f6418f60ac8a511ecdb637",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/plan.md": {
|
||||||
|
"mtime": 1787236725.9825225,
|
||||||
|
"seen": 1787236817.4983613,
|
||||||
|
"ast_hash": "a787c974414ae6c173d4c39e76113c31",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/quickstart.md": {
|
||||||
|
"mtime": 1787236764.0206513,
|
||||||
|
"seen": 1787236817.4983652,
|
||||||
|
"ast_hash": "ada22d34a0fa8341f88136cf42f055a9",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/research.md": {
|
||||||
|
"mtime": 1787236785.0466182,
|
||||||
|
"seen": 1787236817.498368,
|
||||||
|
"ast_hash": "60837e6c7463c4413911ceca0087374c",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/spec.md": {
|
||||||
|
"mtime": 1787236709.0783155,
|
||||||
|
"seen": 1787236817.4983711,
|
||||||
|
"ast_hash": "dfd963f7ca351c93e09f74af9b11d19e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/tasks.md": {
|
||||||
|
"mtime": 1787236802.4808047,
|
||||||
|
"seen": 1787236817.4983742,
|
||||||
|
"ast_hash": "427f24026deaa7c9aca95d99e7587ae6",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/__init__.py": {
|
||||||
|
"mtime": 1787234973.5782337,
|
||||||
|
"seen": 1787235020.6850297,
|
||||||
|
"ast_hash": "627a6c953b250083ebe76ee74d605bb5",
|
||||||
|
"semantic_hash": ""
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
+149
-38
@@ -1,16 +1,16 @@
|
|||||||
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
||||||
|
|
||||||
## Corpus Check
|
## Corpus Check
|
||||||
- 133 files · ~54,759 words
|
- 148 files · ~67,357 words
|
||||||
- Verdict: corpus is large enough that graph structure adds value.
|
- Verdict: corpus is large enough that graph structure adds value.
|
||||||
|
|
||||||
## Summary
|
## Summary
|
||||||
- 595 nodes · 704 edges · 80 communities (43 shown, 37 thin omitted)
|
- 802 nodes · 957 edges · 104 communities (66 shown, 38 thin omitted)
|
||||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 31 edges (avg confidence: 0.95)
|
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
|
||||||
- Token cost: 0 input · 0 output
|
- Token cost: 0 input · 0 output
|
||||||
|
|
||||||
## Graph Freshness
|
## Graph Freshness
|
||||||
- Built from commit: `d371b81a`
|
- Built from commit: `67cc40f9`
|
||||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||||
- Run `graphify update .` after code changes (no API cost).
|
- Run `graphify update .` after code changes (no API cost).
|
||||||
|
|
||||||
@@ -51,15 +51,15 @@
|
|||||||
- Media Transcription
|
- Media Transcription
|
||||||
- Extraction Specification
|
- Extraction Specification
|
||||||
- Graphify Workflows
|
- Graphify Workflows
|
||||||
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
- main
|
||||||
- 1. Technical Decisions & Tradeoffs
|
- 1. Technical Decisions & Tradeoffs
|
||||||
- 1. Input Schemas
|
- 1. Input Schemas
|
||||||
- 2. Basic CLI Usage Examples
|
- 2. Basic CLI Usage Examples
|
||||||
- 2. Standard Streams & Exit Codes
|
- 2. Standard Streams & Exit Codes
|
||||||
- ECPSnapshot
|
- ClassificationResult
|
||||||
- test_models.py
|
- InherenceClassifier
|
||||||
- detect_language
|
- detect_language
|
||||||
- main
|
- test_models.py
|
||||||
- content_northvolt_de.md
|
- content_northvolt_de.md
|
||||||
- content_presal_pt.md
|
- content_presal_pt.md
|
||||||
- content_tangential_es.md
|
- content_tangential_es.md
|
||||||
@@ -91,35 +91,58 @@
|
|||||||
- pt/tangential.md
|
- pt/tangential.md
|
||||||
- tests/__init__.py
|
- tests/__init__.py
|
||||||
- text-nlp-classifier
|
- text-nlp-classifier
|
||||||
|
- get_hl_gl_ceid
|
||||||
|
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||||
|
- extract_google_news.py
|
||||||
|
- ExtractionResult
|
||||||
|
- test_extract_google_news.py
|
||||||
|
- Implementation Tasks: Google News Headlines Extractor
|
||||||
|
- Feature Specification: Google News Headlines Extractor
|
||||||
|
- 2. Cenários Práticos de Uso
|
||||||
|
- Implementation Plan: Google News Headlines Extractor
|
||||||
|
- scripts/__init__.py
|
||||||
|
- SearchQuery
|
||||||
|
- 1. Technical Decisions & Tradeoffs
|
||||||
|
- General Readiness Checklist: Google News Headlines Extractor
|
||||||
|
- 1. Entidades de Domínio & DTOs
|
||||||
|
- Specification Quality Checklist: Google News Headlines Extractor
|
||||||
|
- CLI Contract: Google News Headlines Extractor
|
||||||
|
- 🧠 TextNLPClassifierApp
|
||||||
|
- build_parser
|
||||||
|
- sample_rss_xml
|
||||||
|
- classifier.py
|
||||||
|
- ECPSnapshot
|
||||||
|
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
||||||
|
- main
|
||||||
|
|
||||||
## God Nodes (most connected - your core abstractions)
|
## God Nodes (most connected - your core abstractions)
|
||||||
1. `ECPSnapshot` - 31 edges
|
1. `ECPSnapshot` - 31 edges
|
||||||
2. `InherenceClassifier` - 25 edges
|
2. `InherenceClassifier` - 25 edges
|
||||||
3. `DecisionCategory` - 18 edges
|
3. `DecisionCategory` - 18 edges
|
||||||
4. `ClassificationResult` - 17 edges
|
4. `ClassificationResult` - 17 edges
|
||||||
5. `LocalEmbeddingsAdapter` - 14 edges
|
5. `main()` - 14 edges
|
||||||
6. `LLMFallbackAdapter` - 14 edges
|
6. `LocalEmbeddingsAdapter` - 14 edges
|
||||||
7. `detect_language()` - 14 edges
|
7. `LLMFallbackAdapter` - 14 edges
|
||||||
8. `main()` - 13 edges
|
8. `detect_language()` - 14 edges
|
||||||
9. `Tasks: [FEATURE NAME]` - 13 edges
|
9. `Tasks: [FEATURE NAME]` - 13 edges
|
||||||
10. `BaseNLPAdapter` - 12 edges
|
10. `SearchQuery` - 12 edges
|
||||||
|
|
||||||
## Surprising Connections (you probably didn't know these)
|
## Surprising Connections (you probably didn't know these)
|
||||||
- `emit_error()` --uses--> `ErrorCode` [INFERRED]
|
|
||||||
classify.py → src/models.py
|
|
||||||
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
||||||
classify.py → src/models.py
|
classify.py → src/models.py
|
||||||
- `main()` --uses--> `ErrorCode` [INFERRED]
|
- `main()` --uses--> `ErrorCode` [INFERRED]
|
||||||
classify.py → src/models.py
|
classify.py → src/models.py
|
||||||
- `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
|
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
|
||||||
tests/test_models.py → src/models.py
|
tests/test_extract_google_news.py → scripts/extract_google_news.py
|
||||||
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
|
- `classifier()` --uses--> `InherenceClassifier` [INFERRED]
|
||||||
|
tests/test_benchmark_24.py → src/classifier.py
|
||||||
|
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
||||||
tests/test_models.py → src/models.py
|
tests/test_models.py → src/models.py
|
||||||
|
|
||||||
## Import Cycles
|
## Import Cycles
|
||||||
- None detected.
|
- None detected.
|
||||||
|
|
||||||
## Communities (80 total, 37 thin omitted)
|
## Communities (104 total, 38 thin omitted)
|
||||||
|
|
||||||
### Community 0 - "Task Planning"
|
### Community 0 - "Task Planning"
|
||||||
Cohesion: 0.07
|
Cohesion: 0.07
|
||||||
@@ -241,9 +264,9 @@ Nodes (3): For --cluster-only, For --update (incremental re-extraction), graphif
|
|||||||
Cohesion: 0.50
|
Cohesion: 0.50
|
||||||
Nodes (3): Boundaries, Output, Scan
|
Nodes (3): Boundaries, Output, Scan
|
||||||
|
|
||||||
### Community 40 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
### Community 40 - "main"
|
||||||
Cohesion: 0.14
|
Cohesion: 0.19
|
||||||
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
|
||||||
|
|
||||||
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
||||||
Cohesion: 0.22
|
Cohesion: 0.22
|
||||||
@@ -261,36 +284,124 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
|
|||||||
Cohesion: 0.29
|
Cohesion: 0.29
|
||||||
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
||||||
|
|
||||||
### Community 45 - "ECPSnapshot"
|
### Community 45 - "ClassificationResult"
|
||||||
Cohesion: 0.05
|
Cohesion: 0.10
|
||||||
Nodes (63): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+55 more)
|
Nodes (18): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., LocalEmbeddingsAdapter (+10 more)
|
||||||
|
|
||||||
### Community 46 - "test_models.py"
|
### Community 46 - "InherenceClassifier"
|
||||||
Cohesion: 0.12
|
Cohesion: 0.12
|
||||||
Nodes (17): Any, emit_error(), ClassificationError, extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences. (+9 more)
|
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more)
|
||||||
|
|
||||||
### Community 47 - "detect_language"
|
### Community 47 - "detect_language"
|
||||||
Cohesion: 0.19
|
Cohesion: 0.14
|
||||||
Nodes (16): detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization., Normalize text by converting to lowercase and stripping combining diacritical…, Tokenize text into lowercase alphanumeric words., Detect the ISO-639-1 language code of text among supported languages (pt, en,…, Unit tests for language detection and text normalization. (+8 more)
|
Nodes (20): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization. (+12 more)
|
||||||
|
|
||||||
### Community 48 - "main"
|
### Community 48 - "test_models.py"
|
||||||
|
Cohesion: 0.21
|
||||||
|
Nodes (12): Classify inherence of content against an ECP snapshot., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+4 more)
|
||||||
|
|
||||||
|
### Community 80 - "get_hl_gl_ceid"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
|
||||||
|
|
||||||
|
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
|
||||||
|
Cohesion: 0.08
|
||||||
|
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
|
||||||
|
|
||||||
|
### Community 82 - "extract_google_news.py"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
|
||||||
|
|
||||||
|
### Community 83 - "ExtractionResult"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
|
||||||
|
|
||||||
|
### Community 84 - "test_extract_google_news.py"
|
||||||
|
Cohesion: 0.21
|
||||||
|
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
|
||||||
|
|
||||||
|
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.14
|
||||||
|
Nodes (14): Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Tasks: Google News Headlines Extractor, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Extração Básica de Notícias por Assunto e Idioma (Priority: P1) 🌟 MVP, Phase 4: User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2) (+6 more)
|
||||||
|
|
||||||
|
### Community 86 - "Feature Specification: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.18
|
||||||
|
Nodes (11): Clarifications, Edge Cases, Feature Specification: Google News Headlines Extractor, Functional Requirements, Requirements *(mandatory)*, Session 2026-08-20, Success Criteria *(mandatory)*, User Scenarios & Testing *(mandatory)* (+3 more)
|
||||||
|
|
||||||
|
### Community 87 - "2. Cenários Práticos de Uso"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (9): 1. Pré-requisitos e Instalação, 2. Cenários Práticos de Uso, 3. Validação dos Testes Automatizados e Linter, Cenário 1: River Plate — Argentina (Espanhol / 2 Páginas / Salvar em Arquivo), Cenário 2: Cruzeiro — Brasil (Português / Formatado no Terminal), Cenário 3: Fórmula 1 — Reino Unido (Inglês), Cenário 4: Integração em Pipeline com `jq` (Modo Silencioso), Cenário 5: Extração Rápida com Links Brutos (Sem Resolução de URLs) (+1 more)
|
||||||
|
|
||||||
|
### Community 88 - "Implementation Plan: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (7): Architecture & Pipeline, Documentation (this feature), Implementation Plan: Google News Headlines Extractor, Project Structure, Source Code, Summary, Technical Context
|
||||||
|
|
||||||
|
### Community 90 - "SearchQuery"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (6): Value Object com parâmetros de busca validados., SearchQuery, Valida a consolidação do ExtractionResult a partir da busca mockada com URLs…, Valida as regras de negócio e limites de SearchQuery., test_extract_google_news_orchestration_mocked(), test_search_query_validation()
|
||||||
|
|
||||||
|
### Community 91 - "1. Technical Decisions & Tradeoffs"
|
||||||
|
Cohesion: 0.25
|
||||||
|
Nodes (7): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Requisição e Scraping com `foxcape` em Modo Headless, Decision 2: Endpoint RSS do Google News vs. Scraping de DOM, Decision 3: Mapeamento de Idioma e Locale (`hl`, `gl`, `ceid`), Decision 4: Resolução de URLs do Google News via `googlenewsdecoder`, Decision 5: Logging em Tempo Real no `stderr` e Segregação de Streams, Research: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 92 - "General Readiness Checklist: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (7): CLI Interface & Parameter Contracts, Data Sanitization & Article Extraction, Error Handling & Edge Cases, General Readiness Checklist: Google News Headlines Extractor, Non-Functional & Operational Readiness, Notes, Scraping Engine & Feed Mapping
|
||||||
|
|
||||||
|
### Community 93 - "1. Entidades de Domínio & DTOs"
|
||||||
|
Cohesion: 0.29
|
||||||
|
Nodes (6): 1.1 SearchQuery (Parâmetros da Busca), 1.2 NewsArticle (Item de Notícia), 1.3 ExtractionResult (Saída Estruturada Consolidada), 1. Entidades de Domínio & DTOs, 2. Esquema JSON de Saída, Data Model: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 94 - "Specification Quality Checklist: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.33
|
||||||
|
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Google News Headlines Extractor
|
||||||
|
|
||||||
|
### Community 95 - "CLI Contract: Google News Headlines Extractor"
|
||||||
|
Cohesion: 0.33
|
||||||
|
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
|
||||||
|
|
||||||
|
### Community 97 - "🧠 TextNLPClassifierApp"
|
||||||
|
Cohesion: 0.07
|
||||||
|
Nodes (26): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências, 3. Baixar Binários do Navegador Stealth (Camoufox), 3. Fórmula 1 (Inglaterra / Inglês) (+18 more)
|
||||||
|
|
||||||
|
### Community 98 - "build_parser"
|
||||||
|
Cohesion: 0.67
|
||||||
|
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
|
||||||
|
|
||||||
|
### Community 99 - "sample_rss_xml"
|
||||||
|
Cohesion: 0.67
|
||||||
|
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
|
||||||
|
|
||||||
|
### Community 100 - "classifier.py"
|
||||||
|
Cohesion: 0.23
|
||||||
|
Nodes (9): emit_error(), Enum, Core deterministic classification engine (Tier 1 core)., ClassificationError, ErrorCode, MatchedGraphEntity, Data models and validation schemas for Multilingual NLP Entity Inherence…, str (+1 more)
|
||||||
|
|
||||||
|
### Community 101 - "ECPSnapshot"
|
||||||
|
Cohesion: 0.20
|
||||||
|
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more)
|
||||||
|
|
||||||
|
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||||
|
Cohesion: 0.14
|
||||||
|
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
||||||
|
|
||||||
|
### Community 103 - "main"
|
||||||
Cohesion: 0.31
|
Cohesion: 0.31
|
||||||
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
|
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
|
||||||
|
|
||||||
## Knowledge Gaps
|
## Knowledge Gaps
|
||||||
- **291 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+286 more)
|
- **381 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+376 more)
|
||||||
These have ≤1 connection - possible missing edges or undocumented components.
|
These have ≤1 connection - possible missing edges or undocumented components.
|
||||||
- **37 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||||
|
|
||||||
## Suggested Questions
|
## Suggested Questions
|
||||||
_Questions this graph is uniquely positioned to answer:_
|
_Questions this graph is uniquely positioned to answer:_
|
||||||
|
|
||||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `test_models.py`?**
|
- **Why does `main()` connect `main` to `main`, `classifier.py`, `ECPSnapshot`, `InherenceClassifier`?**
|
||||||
_High betweenness centrality (0.015) - this node is a cross-community bridge._
|
_High betweenness centrality (0.029) - this node is a cross-community bridge._
|
||||||
- **Why does `InherenceClassifier` connect `ECPSnapshot` to `main`?**
|
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `classifier.py`, `main`, `ClassificationResult`, `InherenceClassifier`, `test_models.py`?**
|
||||||
_High betweenness centrality (0.008) - this node is a cross-community bridge._
|
_High betweenness centrality (0.020) - this node is a cross-community bridge._
|
||||||
- **Why does `detect_language()` connect `detect_language` to `ECPSnapshot`?**
|
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
|
||||||
_High betweenness centrality (0.007) - this node is a cross-community bridge._
|
_High betweenness centrality (0.020) - this node is a cross-community bridge._
|
||||||
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
||||||
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||||
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
||||||
|
|||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_specs_002_google_news_extractor_plan_md", "label": "plan.md", "file_type": "document", "node_kind": "page", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "label": "Implementation Plan: Google News Headlines Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_summary", "label": "Summary", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_technical_context", "label": "Technical Context", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L11"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_architecture_pipeline", "label": "Architecture & Pipeline", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L28"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_project_structure", "label": "Project Structure", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L45"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_documentation_this_feature", "label": "Documentation (this feature)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L47"}, {"id": "$graphify-root$_specs_002_google_news_extractor_plan_source_code", "label": "Source Code", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L63"}], "edges": [{"source": "$graphify-root$_specs_002_google_news_extractor_plan_md", "target": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_md", "target": "$graphify-root$_specs_002_google_news_extractor_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L3", "weight": 1.0, "target_file": "$graphify-root$/specs/002-google-news-extractor/spec.md"}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_plan_summary", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_plan_technical_context", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L11", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_plan_architecture_pipeline", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L28", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_implementation_plan_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_plan_project_structure", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L45", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_project_structure", "target": "$graphify-root$_specs_002_google_news_extractor_plan_documentation_this_feature", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L47", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_plan_project_structure", "target": "$graphify-root$_specs_002_google_news_extractor_plan_source_code", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/plan.md", "source_location": "L63", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "pkg_text_nlp_classifier", "label": "text-nlp-classifier", "file_type": "code", "type": "package", "ecosystem": "python", "source_file": "pyproject.toml", "source_location": "L1", "version": "0.1.0"}], "edges": []}
|
||||||
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_md", "label": "cli_contract.md", "file_type": "document", "node_kind": "page", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "label": "CLI Contract: Google News Headlines Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "label": "1. Comando e Argumentos", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L3"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_sintaxe", "label": "Sintaxe", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L5"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_argumentos_de_linha_de_comando", "label": "Argumentos de Linha de Comando", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L10"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_2_c\u00f3digos_de_sa\u00edda_exit_codes", "label": "2. C\u00f3digos de Sa\u00edda (Exit Codes)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L25"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_3_protocolo_de_streams_stdout_stderr", "label": "3. Protocolo de Streams (Stdout / Stderr)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L35"}], "edges": [{"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_md", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L3", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_sintaxe", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L5", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_argumentos_de_linha_de_comando", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L10", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_2_c\u00f3digos_de_sa\u00edda_exit_codes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L25", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_3_protocolo_de_streams_stdout_stderr", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L35", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_specs_002_google_news_extractor_data_model_md", "label": "data-model.md", "file_type": "document", "node_kind": "page", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_data_model_google_news_headlines_extractor", "label": "Data Model: Google News Headlines Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_1_entidades_de_dom\u00ednio_dtos", "label": "1. Entidades de Dom\u00ednio & DTOs", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L3"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_1_1_searchquery_par\u00e2metros_da_busca", "label": "1.1 SearchQuery (Par\u00e2metros da Busca)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L5"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_1_2_newsarticle_item_de_not\u00edcia", "label": "1.2 NewsArticle (Item de Not\u00edcia)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L15"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_1_3_extractionresult_sa\u00edda_estruturada_consolidada", "label": "1.3 ExtractionResult (Sa\u00edda Estruturada Consolidada)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L26"}, {"id": "$graphify-root$_specs_002_google_news_extractor_data_model_2_esquema_json_de_sa\u00edda", "label": "2. Esquema JSON de Sa\u00edda", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L41"}], "edges": [{"source": "$graphify-root$_specs_002_google_news_extractor_data_model_md", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_data_model_google_news_headlines_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_data_model_data_model_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_1_entidades_de_dom\u00ednio_dtos", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L3", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_data_model_1_entidades_de_dom\u00ednio_dtos", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_1_1_searchquery_par\u00e2metros_da_busca", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L5", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_data_model_1_entidades_de_dom\u00ednio_dtos", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_1_2_newsarticle_item_de_not\u00edcia", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L15", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_data_model_1_entidades_de_dom\u00ednio_dtos", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_1_3_extractionresult_sa\u00edda_estruturada_consolidada", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L26", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_data_model_data_model_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_data_model_2_esquema_json_de_sa\u00edda", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/data-model.md", "source_location": "L41", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_md", "label": "cli_contract.md", "file_type": "document", "node_kind": "page", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "label": "CLI Contract: Google News Headlines Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "label": "1. Comando e Argumentos", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L3"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_sintaxe", "label": "Sintaxe", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L5"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_argumentos_de_linha_de_comando", "label": "Argumentos de Linha de Comando", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L10"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_2_c\u00f3digos_de_sa\u00edda_exit_codes", "label": "2. C\u00f3digos de Sa\u00edda (Exit Codes)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L23"}, {"id": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_3_protocolo_de_streams_stdout_stderr", "label": "3. Protocolo de Streams (Stdout / Stderr)", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L33"}], "edges": [{"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_md", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L3", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_sintaxe", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L5", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_1_comando_e_argumentos", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_argumentos_de_linha_de_comando", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L10", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_2_c\u00f3digos_de_sa\u00edda_exit_codes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L23", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_cli_contract_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_contracts_cli_contract_3_protocolo_de_streams_stdout_stderr", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/contracts/cli_contract.md", "source_location": "L33", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_md", "label": "requirements.md", "file_type": "document", "node_kind": "page", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "label": "Specification Quality Checklist: Google News Headlines Extractor", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L1"}, {"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_content_quality", "label": "Content Quality", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L7"}, {"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_requirement_completeness", "label": "Requirement Completeness", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L14"}, {"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_feature_readiness", "label": "Feature Readiness", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L25"}, {"id": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_notes", "label": "Notes", "file_type": "document", "node_kind": "heading", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L32"}], "edges": [{"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_md", "target": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L1", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_md", "target": "$graphify-root$_specs_002_google_news_extractor_spec_md", "relation": "references", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L5", "weight": 1.0, "target_file": "$graphify-root$/specs/002-google-news-extractor/spec.md"}, {"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_content_quality", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L7", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_requirement_completeness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L14", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_feature_readiness", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L25", "weight": 1.0}, {"source": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_specification_quality_checklist_google_news_headlines_extractor", "target": "$graphify-root$_specs_002_google_news_extractor_checklists_requirements_notes", "relation": "contains", "confidence": "EXTRACTED", "source_file": "specs/002-google-news-extractor/checklists/requirements.md", "source_location": "L32", "weight": 1.0}], "input_tokens": 0, "output_tokens": 0}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
+1
@@ -0,0 +1 @@
|
|||||||
|
{"nodes": [{"id": "$graphify-root$_scripts_init_py", "label": "__init__.py", "file_type": "code", "source_file": "scripts/__init__.py", "source_location": "L1"}, {"id": "$graphify-root$_scripts_init_rationale_1", "label": "M\u00f3dulo de scripts utilit\u00e1rios do projeto.", "file_type": "rationale", "source_file": "scripts/__init__.py", "source_location": "L1"}], "edges": [{"source": "$graphify-root$_scripts_init_rationale_1", "target": "$graphify-root$_scripts_init_py", "relation": "rationale_for", "confidence": "EXTRACTED", "source_file": "scripts/__init__.py", "source_location": "L1", "weight": 1.0}], "raw_calls": []}
|
||||||
+1
File diff suppressed because one or more lines are too long
+1
File diff suppressed because one or more lines are too long
Vendored
+1
-1
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+6409
-1125
File diff suppressed because it is too large
Load Diff
@@ -300,9 +300,9 @@
|
|||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"pyproject.toml": {
|
"pyproject.toml": {
|
||||||
"mtime": 1787195827.3305018,
|
"mtime": 1787234982.768605,
|
||||||
"seen": 1787196463.0953014,
|
"seen": 1787235020.6850271,
|
||||||
"ast_hash": "f2503e96d08f5a4e41b13352e81709c6",
|
"ast_hash": "26f5f979658067bf447f375f044a8f21",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"src/__init__.py": {
|
"src/__init__.py": {
|
||||||
@@ -420,9 +420,9 @@
|
|||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"requirements.txt": {
|
"requirements.txt": {
|
||||||
"mtime": 1787195817.4782183,
|
"mtime": 1787236111.876247,
|
||||||
"seen": 1787196463.1020837,
|
"seen": 1787236262.0105975,
|
||||||
"ast_hash": "5612073a1e7034d7c25765379baa5da4",
|
"ast_hash": "9fb7e3f8ecb04ff30d650c772f509557",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
},
|
},
|
||||||
"tests/fixtures/benchmark_24/de/contextual.md": {
|
"tests/fixtures/benchmark_24/de/contextual.md": {
|
||||||
@@ -574,5 +574,89 @@
|
|||||||
"seen": 1787197159.9881344,
|
"seen": 1787197159.9881344,
|
||||||
"ast_hash": "5e8daf517ee32c261617d96264ef0473",
|
"ast_hash": "5e8daf517ee32c261617d96264ef0473",
|
||||||
"semantic_hash": ""
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/extract_google_news.py": {
|
||||||
|
"mtime": 1787236453.5231855,
|
||||||
|
"seen": 1787236509.1147037,
|
||||||
|
"ast_hash": "4216de490a87378d21dab707a3166679",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"tests/test_extract_google_news.py": {
|
||||||
|
"mtime": 1787236301.513448,
|
||||||
|
"seen": 1787236353.5582018,
|
||||||
|
"ast_hash": "1fc9a108a8abfecf1d37c8642a885e9e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"docs/googlenews_extractor_guia_completo.md": {
|
||||||
|
"mtime": 1787231605.0830677,
|
||||||
|
"seen": 1787234772.8112028,
|
||||||
|
"ast_hash": "59db8e652000ba6087a00289a0ce0959",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/checklists/readiness.md": {
|
||||||
|
"mtime": 1787234305.932687,
|
||||||
|
"seen": 1787234772.8123577,
|
||||||
|
"ast_hash": "a1200a6959d41e256a73403dd5c4e1f6",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/checklists/requirements.md": {
|
||||||
|
"mtime": 1787232386.6763985,
|
||||||
|
"seen": 1787234772.812359,
|
||||||
|
"ast_hash": "39cd86af76c9f74baa2984418870b677",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/contracts/cli_contract.md": {
|
||||||
|
"mtime": 1787236744.3139226,
|
||||||
|
"seen": 1787236817.4979818,
|
||||||
|
"ast_hash": "76a61ad31df989244edc5e739081bf7b",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/data-model.md": {
|
||||||
|
"mtime": 1787234040.8134317,
|
||||||
|
"seen": 1787234772.8123617,
|
||||||
|
"ast_hash": "e8380c98d2f6418f60ac8a511ecdb637",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/plan.md": {
|
||||||
|
"mtime": 1787236725.9825225,
|
||||||
|
"seen": 1787236817.4983613,
|
||||||
|
"ast_hash": "a787c974414ae6c173d4c39e76113c31",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/quickstart.md": {
|
||||||
|
"mtime": 1787236764.0206513,
|
||||||
|
"seen": 1787236817.4983652,
|
||||||
|
"ast_hash": "ada22d34a0fa8341f88136cf42f055a9",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/research.md": {
|
||||||
|
"mtime": 1787236785.0466182,
|
||||||
|
"seen": 1787236817.498368,
|
||||||
|
"ast_hash": "60837e6c7463c4413911ceca0087374c",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/spec.md": {
|
||||||
|
"mtime": 1787236709.0783155,
|
||||||
|
"seen": 1787236817.4983711,
|
||||||
|
"ast_hash": "dfd963f7ca351c93e09f74af9b11d19e",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"specs/002-google-news-extractor/tasks.md": {
|
||||||
|
"mtime": 1787236802.4808047,
|
||||||
|
"seen": 1787236817.4983742,
|
||||||
|
"ast_hash": "427f24026deaa7c9aca95d99e7587ae6",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"scripts/__init__.py": {
|
||||||
|
"mtime": 1787234973.5782337,
|
||||||
|
"seen": 1787235020.6850297,
|
||||||
|
"ast_hash": "627a6c953b250083ebe76ee74d605bb5",
|
||||||
|
"semantic_hash": ""
|
||||||
|
},
|
||||||
|
"README.md": {
|
||||||
|
"mtime": 1787237002.3288481,
|
||||||
|
"seen": 1787237013.7487457,
|
||||||
|
"ast_hash": "1b624e85d8ce1521683cbea2664dd2d2",
|
||||||
|
"semantic_hash": ""
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -16,7 +16,11 @@ test = [
|
|||||||
]
|
]
|
||||||
|
|
||||||
[tool.pytest.ini_options]
|
[tool.pytest.ini_options]
|
||||||
|
pythonpath = [".", "src"]
|
||||||
testpaths = ["tests"]
|
testpaths = ["tests"]
|
||||||
python_files = ["test_*.py"]
|
python_files = ["test_*.py"]
|
||||||
python_classes = ["Test*"]
|
python_classes = ["Test*"]
|
||||||
python_functions = ["test_*"]
|
python_functions = ["test_*"]
|
||||||
|
|
||||||
|
[tool.pyright]
|
||||||
|
extraPaths = [".", "src"]
|
||||||
|
|||||||
@@ -1,4 +1,8 @@
|
|||||||
pytest>=7.0.0
|
pytest>=7.0.0
|
||||||
|
foxcape>=0.1.1
|
||||||
|
beautifulsoup4>=4.12.0
|
||||||
|
googlenewsdecoder>=0.1.7
|
||||||
|
selectolax>=0.3.27
|
||||||
# Optional Tier 2 / Tier 3 dependencies (not required for POC core execution)
|
# Optional Tier 2 / Tier 3 dependencies (not required for POC core execution)
|
||||||
# sentence-transformers>=2.2.0
|
# sentence-transformers>=2.2.0
|
||||||
# httpx>=0.24.0
|
# httpx>=0.24.0
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
"""Módulo de scripts utilitários do projeto."""
|
||||||
@@ -0,0 +1,467 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""
|
||||||
|
Extrator de Manchetes do Google News via RSS.
|
||||||
|
|
||||||
|
Script CLI autônomo para busca, extração e estruturação de notícias
|
||||||
|
por termo/assunto, idioma e localização geográfica (locale), utilizando
|
||||||
|
a biblioteca Foxcape para raspagem indetectável e decodificação automática
|
||||||
|
das URLs intermediárias do Google News para as URLs finais dos veículos.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import concurrent.futures
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import urllib.error
|
||||||
|
import urllib.parse
|
||||||
|
import urllib.request
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from bs4 import BeautifulSoup, FeatureNotFound
|
||||||
|
from foxcape import Foxcape, FoxcapeConfig
|
||||||
|
from googlenewsdecoder import gnewsdecoder # type: ignore[import-untyped]
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class SearchQuery:
|
||||||
|
"""Value Object com parâmetros de busca validados."""
|
||||||
|
|
||||||
|
keyword: str
|
||||||
|
language: str = "pt"
|
||||||
|
locale: str | None = None
|
||||||
|
max_pages: int = 1
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
if not self.keyword or not self.keyword.strip():
|
||||||
|
raise ValueError("A palavra-chave não pode ser vazia.")
|
||||||
|
|
||||||
|
if not self.language or len(self.language.strip()) < 2:
|
||||||
|
raise ValueError(
|
||||||
|
"O idioma deve conter pelo menos 2 caracteres (ex: 'pt', 'en', 'es')."
|
||||||
|
)
|
||||||
|
|
||||||
|
if self.max_pages < 1 or self.max_pages > 10:
|
||||||
|
raise ValueError("O número máximo de páginas deve estar entre 1 e 10.")
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean_keyword(self) -> str:
|
||||||
|
return self.keyword.strip()
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean_language(self) -> str:
|
||||||
|
return self.language.strip().lower()
|
||||||
|
|
||||||
|
@property
|
||||||
|
def clean_locale(self) -> str | None:
|
||||||
|
return self.locale.strip().upper() if self.locale else None
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class NewsArticle:
|
||||||
|
"""Entidade que representa uma notícia extraída."""
|
||||||
|
|
||||||
|
titulo: str
|
||||||
|
url: str
|
||||||
|
pagina: int
|
||||||
|
subtitulo: str | None = None
|
||||||
|
quando_publicado: str | None = None
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"titulo": self.titulo,
|
||||||
|
"subtitulo": self.subtitulo,
|
||||||
|
"quando_publicado": self.quando_publicado,
|
||||||
|
"url": self.url,
|
||||||
|
"pagina": self.pagina,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ExtractionResult:
|
||||||
|
"""Resultado consolidado da extração."""
|
||||||
|
|
||||||
|
query: str
|
||||||
|
language: str
|
||||||
|
locale: str
|
||||||
|
total_paginas: int
|
||||||
|
total_itens: int
|
||||||
|
scraped_at: str
|
||||||
|
items: list[NewsArticle] = field(default_factory=list)
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"query": self.query,
|
||||||
|
"language": self.language,
|
||||||
|
"locale": self.locale,
|
||||||
|
"total_paginas": self.total_paginas,
|
||||||
|
"total_itens": self.total_itens,
|
||||||
|
"scraped_at": self.scraped_at,
|
||||||
|
"items": [item.to_dict() for item in self.items],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def get_hl_gl_ceid(lang_raw: str, locale: str | None = None) -> tuple[str, str, str]:
|
||||||
|
"""
|
||||||
|
Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News.
|
||||||
|
"""
|
||||||
|
lang_clean = lang_raw.lower().replace("-", "_")
|
||||||
|
|
||||||
|
locale_map: dict[str, tuple[str, str]] = {
|
||||||
|
"pt": ("pt-BR", "BR"),
|
||||||
|
"pt_br": ("pt-BR", "BR"),
|
||||||
|
"es": ("es-419", "AR"),
|
||||||
|
"es_mx": ("es-419", "MX"),
|
||||||
|
"es_es": ("es", "ES"),
|
||||||
|
"en": ("en-US", "US"),
|
||||||
|
"en_gb": ("en-GB", "GB"),
|
||||||
|
"en_uk": ("en-GB", "GB"),
|
||||||
|
"en_us": ("en-US", "US"),
|
||||||
|
"de": ("de", "DE"),
|
||||||
|
"de_de": ("de", "DE"),
|
||||||
|
"it": ("it", "IT"),
|
||||||
|
"it_it": ("it", "IT"),
|
||||||
|
"fr": ("fr", "FR"),
|
||||||
|
"fr_fr": ("fr", "FR"),
|
||||||
|
}
|
||||||
|
|
||||||
|
if lang_clean in locale_map:
|
||||||
|
hl, gl = locale_map[lang_clean]
|
||||||
|
else:
|
||||||
|
parts = lang_clean.split("_")
|
||||||
|
if len(parts) == 2:
|
||||||
|
hl, gl = f"{parts[0]}-{parts[1].upper()}", parts[1].upper()
|
||||||
|
else:
|
||||||
|
hl, gl = lang_clean, lang_clean.upper()
|
||||||
|
|
||||||
|
if locale:
|
||||||
|
gl = locale.strip().upper()
|
||||||
|
if lang_clean == "es" and gl == "ES":
|
||||||
|
hl = "es"
|
||||||
|
elif lang_clean == "en" and gl == "GB":
|
||||||
|
hl = "en-GB"
|
||||||
|
|
||||||
|
ceid = f"{gl}:{hl}"
|
||||||
|
return hl, gl, ceid
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_text_for_comparison(text: str) -> str:
|
||||||
|
"""Remove pontuação e espaços extras para comparação de redundância."""
|
||||||
|
return re.sub(r"[^\w\s]", "", text).lower().strip()
|
||||||
|
|
||||||
|
|
||||||
|
def parse_google_news_rss(xml_content: str, max_pages: int = 1) -> list[NewsArticle]:
|
||||||
|
"""
|
||||||
|
Parseia o XML do RSS do Google News e extrai os itens estruturados.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
soup = BeautifulSoup(xml_content, "xml")
|
||||||
|
except FeatureNotFound:
|
||||||
|
soup = BeautifulSoup(xml_content, "html.parser")
|
||||||
|
|
||||||
|
rss_items = soup.find_all("item")
|
||||||
|
articles: list[NewsArticle] = []
|
||||||
|
max_allowed = max_pages * 10
|
||||||
|
|
||||||
|
for idx, item in enumerate(rss_items[:max_allowed]):
|
||||||
|
page_number = (idx // 10) + 1
|
||||||
|
|
||||||
|
title_tag = item.find("title")
|
||||||
|
link_tag = item.find("link")
|
||||||
|
pubdate_tag = item.find("pubDate")
|
||||||
|
desc_tag = item.find("description")
|
||||||
|
|
||||||
|
title = title_tag.get_text(strip=True) if title_tag else ""
|
||||||
|
link = link_tag.get_text(strip=True) if link_tag else ""
|
||||||
|
pub_date = pubdate_tag.get_text(strip=True) if pubdate_tag else None
|
||||||
|
desc_raw = desc_tag.get_text() if desc_tag else ""
|
||||||
|
|
||||||
|
# Limpeza de HTML do resumo / snippet
|
||||||
|
snippet = None
|
||||||
|
if desc_raw:
|
||||||
|
desc_soup = BeautifulSoup(desc_raw, "html.parser")
|
||||||
|
clean_snippet = desc_soup.get_text(separator=" ", strip=True)
|
||||||
|
if clean_snippet:
|
||||||
|
norm_snippet = _normalize_text_for_comparison(clean_snippet)
|
||||||
|
norm_title = _normalize_text_for_comparison(title)
|
||||||
|
if norm_snippet != norm_title:
|
||||||
|
snippet = clean_snippet
|
||||||
|
|
||||||
|
if title and link:
|
||||||
|
articles.append(
|
||||||
|
NewsArticle(
|
||||||
|
titulo=title,
|
||||||
|
subtitulo=snippet,
|
||||||
|
quando_publicado=pub_date,
|
||||||
|
url=link,
|
||||||
|
pagina=page_number,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
return articles
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_article_url(url: str) -> str:
|
||||||
|
"""Resolve a URL intermediária do Google News para a URL real do veículo."""
|
||||||
|
if not url or "news.google.com" not in url:
|
||||||
|
return url
|
||||||
|
try:
|
||||||
|
res = gnewsdecoder(url)
|
||||||
|
if isinstance(res, dict) and res.get("status") and res.get("decoded_url"):
|
||||||
|
return str(res["decoded_url"])
|
||||||
|
except (ValueError, OSError, RuntimeError, TypeError):
|
||||||
|
pass
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_articles_urls(
|
||||||
|
articles: list[NewsArticle], max_workers: int = 5, verbose: bool = True
|
||||||
|
) -> list[NewsArticle]:
|
||||||
|
"""Resolve em paralelo as URLs intermediárias do Google News para os links finais dos veículos."""
|
||||||
|
if not articles:
|
||||||
|
return articles
|
||||||
|
|
||||||
|
if verbose:
|
||||||
|
sys.stderr.write(
|
||||||
|
f"[INFO] 🔗 Decodificando {len(articles)} URLs do Google News para os portais reais...\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
def _resolve_single(art: NewsArticle) -> NewsArticle:
|
||||||
|
return NewsArticle(
|
||||||
|
titulo=art.titulo,
|
||||||
|
subtitulo=art.subtitulo,
|
||||||
|
quando_publicado=art.quando_publicado,
|
||||||
|
url=resolve_article_url(art.url),
|
||||||
|
pagina=art.pagina,
|
||||||
|
)
|
||||||
|
|
||||||
|
with concurrent.futures.ThreadPoolExecutor(max_workers=max_workers) as executor:
|
||||||
|
resolved = list(executor.map(_resolve_single, articles))
|
||||||
|
|
||||||
|
if verbose:
|
||||||
|
resolved_count = sum(1 for a in resolved if "news.google.com" not in a.url)
|
||||||
|
sys.stderr.write(
|
||||||
|
f"[INFO] ✅ {resolved_count}/{len(resolved)} URLs resolvidas com sucesso para os domínios de origem.\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
return resolved
|
||||||
|
|
||||||
|
|
||||||
|
def _fetch_rss_content(rss_url: str) -> str:
|
||||||
|
"""
|
||||||
|
Realiza a requisição ao feed RSS utilizando a biblioteca Foxcape como
|
||||||
|
motor primário de extração e evasão anti-bot em modo headless.
|
||||||
|
"""
|
||||||
|
config = FoxcapeConfig(
|
||||||
|
headless=True,
|
||||||
|
humanize=False,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
result = Foxcape.fetch(
|
||||||
|
rss_url,
|
||||||
|
config=config,
|
||||||
|
timeout_ms=15000,
|
||||||
|
human_delay=False,
|
||||||
|
)
|
||||||
|
if result and result.html:
|
||||||
|
return str(result.html)
|
||||||
|
raise RuntimeError("Foxcape não retornou conteúdo para a URL informada.")
|
||||||
|
except (RuntimeError, OSError, TimeoutError, ValueError) as exc:
|
||||||
|
# Fallback de contingência caso o ambiente não possua suporte a subprocessos de navegador
|
||||||
|
try:
|
||||||
|
headers = {
|
||||||
|
"User-Agent": (
|
||||||
|
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
|
||||||
|
"AppleWebKit/537.36 (KHTML, like Gecko) Chrome/132.0.0.0 Safari/537.36"
|
||||||
|
),
|
||||||
|
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
|
||||||
|
"Accept-Language": "pt-BR,pt;q=0.9,en-US;q=0.8,en;q=0.7,es;q=0.6",
|
||||||
|
}
|
||||||
|
req = urllib.request.Request(rss_url, headers=headers)
|
||||||
|
with urllib.request.urlopen(req, timeout=10) as response:
|
||||||
|
return response.read().decode("utf-8", errors="replace")
|
||||||
|
except (
|
||||||
|
urllib.error.URLError,
|
||||||
|
urllib.error.HTTPError,
|
||||||
|
OSError,
|
||||||
|
TimeoutError,
|
||||||
|
) as fallback_exc:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"Erro na extração Foxcape: {exc} | Contingência: {fallback_exc}"
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
def extract_google_news(
|
||||||
|
query: SearchQuery, resolve_urls: bool = True, verbose: bool = True
|
||||||
|
) -> ExtractionResult:
|
||||||
|
"""
|
||||||
|
Orquestra a consulta ao Google News e retorna o resultado consolidado
|
||||||
|
com URLs resolvidas para os portais de notícias.
|
||||||
|
"""
|
||||||
|
encoded_query = urllib.parse.quote_plus(query.clean_keyword)
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid(query.clean_language, query.clean_locale)
|
||||||
|
rss_url = f"https://news.google.com/rss/search?q={encoded_query}&hl={hl}&gl={gl}&ceid={ceid}"
|
||||||
|
|
||||||
|
if verbose:
|
||||||
|
sys.stderr.write(
|
||||||
|
f"[INFO] 🔍 Consultando Google News: '{query.clean_keyword}' "
|
||||||
|
f"(idioma: {query.clean_language}, locale: {gl}, max_pages: {query.max_pages})...\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
xml_content = _fetch_rss_content(rss_url)
|
||||||
|
if verbose:
|
||||||
|
sys.stderr.write(f"[INFO] 📥 Feed RSS recebido ({len(xml_content)} bytes).\n")
|
||||||
|
|
||||||
|
articles = parse_google_news_rss(xml_content, max_pages=query.max_pages)
|
||||||
|
if verbose:
|
||||||
|
sys.stderr.write(f"[INFO] 📰 {len(articles)} artigos extraídos do feed XML.\n")
|
||||||
|
|
||||||
|
if resolve_urls and articles:
|
||||||
|
articles = resolve_articles_urls(articles, verbose=verbose)
|
||||||
|
|
||||||
|
return ExtractionResult(
|
||||||
|
query=query.clean_keyword,
|
||||||
|
language=query.clean_language,
|
||||||
|
locale=gl,
|
||||||
|
total_paginas=query.max_pages,
|
||||||
|
total_itens=len(articles),
|
||||||
|
scraped_at=datetime.now(timezone.utc).isoformat(),
|
||||||
|
items=articles,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
"""Cria e configura o parser de argumentos CLI."""
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
prog="extract_google_news",
|
||||||
|
description="Extrator de manchetes do Google News RSS por assunto, idioma e região.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-q",
|
||||||
|
"--query",
|
||||||
|
"--keyword",
|
||||||
|
dest="query",
|
||||||
|
required=True,
|
||||||
|
help="Termo ou expressão de busca (obrigatório).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-l",
|
||||||
|
"--lang",
|
||||||
|
"--language",
|
||||||
|
dest="lang",
|
||||||
|
default="pt",
|
||||||
|
help="Código do idioma (padrão: pt).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--locale",
|
||||||
|
"--country",
|
||||||
|
dest="locale",
|
||||||
|
default=None,
|
||||||
|
help="Código do país/região (ex: BR, US, MX, ES).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-p",
|
||||||
|
"--max-pages",
|
||||||
|
dest="max_pages",
|
||||||
|
type=int,
|
||||||
|
default=1,
|
||||||
|
help="Número de páginas (1 a 10, com 10 itens por página; padrão: 1).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-o",
|
||||||
|
"--output",
|
||||||
|
dest="output",
|
||||||
|
default=None,
|
||||||
|
help="Caminho do arquivo para salvar a saída JSON.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--pretty",
|
||||||
|
action="store_true",
|
||||||
|
help="Formata a saída JSON com indentação legível.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--no-resolve-urls",
|
||||||
|
dest="resolve_urls",
|
||||||
|
action="store_false",
|
||||||
|
default=True,
|
||||||
|
help="Desativa a decodificação para as URLs diretas dos portais de notícias.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-s",
|
||||||
|
"--silent",
|
||||||
|
"--quiet",
|
||||||
|
dest="silent",
|
||||||
|
action="store_true",
|
||||||
|
help="Suprime as mensagens informativas de progresso.",
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: list[str] | None = None) -> int:
|
||||||
|
"""Ponto de entrada do script CLI."""
|
||||||
|
if hasattr(sys.stdout, "reconfigure"):
|
||||||
|
try:
|
||||||
|
sys.stdout.reconfigure(encoding="utf-8")
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
if hasattr(sys.stderr, "reconfigure"):
|
||||||
|
try:
|
||||||
|
sys.stderr.reconfigure(encoding="utf-8")
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
parser = build_parser()
|
||||||
|
|
||||||
|
try:
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
query = SearchQuery(
|
||||||
|
keyword=args.query,
|
||||||
|
language=args.lang,
|
||||||
|
locale=args.locale,
|
||||||
|
max_pages=args.max_pages,
|
||||||
|
)
|
||||||
|
except (ValueError, argparse.ArgumentError) as err:
|
||||||
|
sys.stderr.write(f"Erro de validação: {err}\n")
|
||||||
|
return 1
|
||||||
|
except SystemExit as err:
|
||||||
|
return err.code if isinstance(err.code, int) else 1
|
||||||
|
|
||||||
|
verbose = not getattr(args, "silent", False)
|
||||||
|
|
||||||
|
try:
|
||||||
|
result = extract_google_news(
|
||||||
|
query, resolve_urls=args.resolve_urls, verbose=verbose
|
||||||
|
)
|
||||||
|
json_output = json.dumps(
|
||||||
|
result.to_dict(),
|
||||||
|
ensure_ascii=False,
|
||||||
|
indent=2 if args.pretty else None,
|
||||||
|
)
|
||||||
|
|
||||||
|
if args.output:
|
||||||
|
out_path = Path(args.output)
|
||||||
|
out_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
out_path.write_text(json_output + "\n", encoding="utf-8")
|
||||||
|
if verbose:
|
||||||
|
sys.stderr.write(
|
||||||
|
f"[INFO] 💾 Arquivo salvo com sucesso: '{args.output}' ({result.total_itens} notícias).\n"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
sys.stdout.write(json_output + "\n")
|
||||||
|
|
||||||
|
return 0
|
||||||
|
except RuntimeError as err:
|
||||||
|
sys.stderr.write(f"Erro na extração: {err}\n")
|
||||||
|
return 2
|
||||||
|
except OSError as err:
|
||||||
|
sys.stderr.write(f"Erro de arquivo/sistema: {err}\n")
|
||||||
|
return 2
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
# General Readiness Checklist: Google News Headlines Extractor
|
||||||
|
|
||||||
|
**Purpose**: Validate the completeness, clarity, consistency, and testability of requirements for the standalone Google News CLI extractor before implementation.
|
||||||
|
**Created**: 2026-08-20
|
||||||
|
**Validated**: 2026-08-20 (Audited against guide, spec, research, plan, data-model, and cli_contract)
|
||||||
|
**Feature**: [spec.md](../spec.md) | [plan.md](../plan.md) | [cli_contract.md](../contracts/cli_contract.md) | [data-model.md](../data-model.md) | [research.md](../research.md)
|
||||||
|
|
||||||
|
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
|
||||||
|
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||||
|
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
|
||||||
|
|
||||||
|
## CLI Interface & Parameter Contracts
|
||||||
|
|
||||||
|
- [x] CHK001 - Are all required CLI arguments (e.g., `--query` / `--keyword`) and their aliases explicitly specified? [Completeness, Spec §FR-003, Contract §1]
|
||||||
|
- [x] CHK002 - Are default values clearly defined for optional parameters (`--lang`, `--locale`, `--max-pages`)? [Clarity, Spec §FR-003, Contract §1]
|
||||||
|
- [x] CHK003 - Is the allowed range for pagination (`1` to `10` pages) explicitly bounded and unambiguous? [Clarity, Spec §FR-006, DataModel §1.1]
|
||||||
|
- [x] CHK004 - Are exit codes defined for each distinct execution outcome (success, validation failure, scraping/network error)? [Completeness, Contract §2]
|
||||||
|
- [x] CHK005 - Are stream separation requirements (clean JSON on `stdout`, diagnostics/errors on `stderr`) strictly documented? [Consistency, Spec §FR-009, Contract §3]
|
||||||
|
|
||||||
|
## Scraping Engine & Feed Mapping
|
||||||
|
|
||||||
|
- [x] CHK006 - Is the integration role of the `foxcape` library and its anti-bot evasion responsibilities clearly documented? [Completeness, Spec §FR-002, Plan §Summary]
|
||||||
|
- [x] CHK007 - Is the RSS search URL template and its encoding rules (`q`, `hl`, `gl`, `ceid`) precisely defined? [Clarity, Research §Decision 2, Research §Decision 3]
|
||||||
|
- [x] CHK008 - Is the locale inference fallback rule (when `--locale` is omitted) explicitly specified across standard language codes? [Consistency, Spec §FR-005, Research §Decision 3]
|
||||||
|
- [x] CHK009 - Is the override behavior when a custom `--locale` is passed alongside a different language documented? [Completeness, Spec §User Story 2, Research §Decision 3]
|
||||||
|
|
||||||
|
## Data Sanitization & Article Extraction
|
||||||
|
|
||||||
|
- [x] CHK010 - Are the target extraction fields (`titulo`, `subtitulo`, `quando_publicado`, `url`, `pagina`) mapped to specific RSS XML tags? [Completeness, Spec §FR-007, DataModel §1.2]
|
||||||
|
- [x] CHK011 - Is the HTML tag stripping requirement for description/snippets specified with unambiguous criteria? [Clarity, Spec §FR-008, SC-003]
|
||||||
|
- [x] CHK012 - Is the deduplication rule for subtitles identical to titles clearly documented with expected null/empty behavior? [Consistency, Spec §Edge Cases, DataModel §1.2]
|
||||||
|
- [x] CHK013 - Are filtering rules specified for discarding incomplete items missing title or link? [Coverage, Spec §Edge Cases, SC-002]
|
||||||
|
- [x] CHK014 - Is the consolidated JSON schema defined with all mandatory metadata fields (`query`, `language`, `locale`, `total_itens`, `scraped_at`)? [Completeness, Spec §FR-009, DataModel §2]
|
||||||
|
|
||||||
|
## Error Handling & Edge Cases
|
||||||
|
|
||||||
|
- [x] CHK015 - Are validation requirements specified for empty or whitespace-only search keywords? [Coverage, Spec §Edge Cases, Spec §FR-010]
|
||||||
|
- [x] CHK016 - Is the behavior for queries yielding zero matching news articles specified without raising unhandled errors? [Coverage, Spec §User Story 1, SC-004]
|
||||||
|
- [x] CHK017 - Are network interruption or upstream HTTP block error handling requirements documented? [Coverage, Spec §Edge Cases, Contract §2]
|
||||||
|
- [x] CHK018 - Are character encoding and URL parameter escaping requirements documented for special characters and accents? [Clarity, Spec §Edge Cases]
|
||||||
|
|
||||||
|
## Non-Functional & Operational Readiness
|
||||||
|
|
||||||
|
- [x] CHK019 - Is the performance target (< 2.0s for standard single-page extractions) quantified with measurable criteria? [Measurability, Spec §SC-001, Plan §Technical Context]
|
||||||
|
- [x] CHK020 - Are JSON output compatibility requirements with terminal streaming tools (e.g., `jq`, shell pipes) defined? [Completeness, Spec §SC-005, Contract §3]
|
||||||
|
- [x] CHK021 - Is the execution isolation assumption (pure RSS feed extraction without mandatory Playwright browser runtime) clearly stated? [Traceability, Spec §Assumptions, Research §Decision 1]
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied
|
||||||
|
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation
|
||||||
|
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers
|
||||||
|
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`
|
||||||
|
- Add comments or findings inline
|
||||||
|
- Link to relevant resources or documentation
|
||||||
|
- Items are numbered sequentially for easy reference
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
# Specification Quality Checklist: Google News Headlines Extractor
|
||||||
|
|
||||||
|
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||||
|
**Created**: 2026-08-20
|
||||||
|
**Feature**: [spec.md](../spec.md)
|
||||||
|
|
||||||
|
## Content Quality
|
||||||
|
|
||||||
|
- [x] No implementation details (languages, frameworks, APIs)
|
||||||
|
- [x] Focused on user value and business needs
|
||||||
|
- [x] Written for non-technical stakeholders
|
||||||
|
- [x] All mandatory sections completed
|
||||||
|
|
||||||
|
## Requirement Completeness
|
||||||
|
|
||||||
|
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||||
|
- [x] Requirements are testable and unambiguous
|
||||||
|
- [x] Success criteria are measurable
|
||||||
|
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||||
|
- [x] All acceptance scenarios are defined
|
||||||
|
- [x] Edge cases are identified
|
||||||
|
- [x] Scope is clearly bounded
|
||||||
|
- [x] Dependencies and assumptions identified
|
||||||
|
|
||||||
|
## Feature Readiness
|
||||||
|
|
||||||
|
- [x] All functional requirements have clear acceptance criteria
|
||||||
|
- [x] User scenarios cover primary flows
|
||||||
|
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||||
|
- [x] No implementation details leak into specification
|
||||||
|
|
||||||
|
## Notes
|
||||||
|
|
||||||
|
- Feature spec ready for planning phase (`/speckit-plan`).
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# CLI Contract: Google News Headlines Extractor
|
||||||
|
|
||||||
|
## 1. Comando e Argumentos
|
||||||
|
|
||||||
|
### Sintaxe
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py --query <KEYWORD> [--lang <LANG>] [--locale <LOCALE>] [--max-pages <PAGES>] [--output <FILE>] [--pretty] [--no-resolve-urls] [--silent]
|
||||||
|
```
|
||||||
|
|
||||||
|
### Argumentos de Linha de Comando
|
||||||
|
|
||||||
|
| Flag / Argumento | Tipo | Obrigatório | Padrão | Descrição |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| `-q`, `--query`, `--keyword` | `str` | **Sim** | — | Termo ou expressão de pesquisa no Google News. |
|
||||||
|
| `-l`, `--lang`, `--language` | `str` | Não | `"pt"` | Código do idioma (ex: `pt`, `en`, `es`, `de`, `fr`, `it`). |
|
||||||
|
| `--locale`, `--country` | `str` | Não | `None` | Código do país/região (ex: `BR`, `US`, `GB`, `MX`, `ES`, `AR`). |
|
||||||
|
| `-p`, `--max-pages` | `int` | Não | `1` | Quantidade de páginas a extrair (1 a 10, onde cada página possui até 10 itens). |
|
||||||
|
| `-o`, `--output` | `str` | Não | `None` | Caminho de arquivo opcional para salvar o JSON resultante diretamente (cria diretórios pais automaticamente). |
|
||||||
|
| `--pretty` | `flag` | Não | `False` | Formata o JSON emitido no stdout com indentação legível (2 espaços). |
|
||||||
|
| `--no-resolve-urls` | `flag` | Não | `False` | Desativa a decodificação automática para as URLs originais dos veículos (mantém os links brutos do feed). |
|
||||||
|
| `-s`, `--silent`, `--quiet` | `flag` | Não | `False` | Suprime mensagens informativas de progresso emitidas no `stderr`. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Códigos de Saída (Exit Codes)
|
||||||
|
|
||||||
|
| Código | Significado | Descrição |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `0` | **Sucesso** | Extração concluída com êxito (mesmo que 0 notícias sejam encontradas). |
|
||||||
|
| `1` | **Erro de Validação** | Parâmetro obrigatório ausente, valor inválido ou flag desconhecida. |
|
||||||
|
| `2` | **Erro de Rede/Scraping/I/O** | Falha de conectividade, bloqueio não recuperável ou erro de gravação. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Protocolo de Streams (Stdout / Stderr)
|
||||||
|
|
||||||
|
- **`stdout`**: Exclusivo para o payload JSON estruturado de saída. Permite redirecionamento direto para pipes e arquivos:
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "tecnologia" | jq '.items[].titulo'
|
||||||
|
```
|
||||||
|
- **`stderr`**: Exclusivo para logs informativos de progresso e mensagens de erro:
|
||||||
|
```text
|
||||||
|
[INFO] 🔍 Consultando Google News: 'River Plate' (idioma: es, locale: AR, max_pages: 2)...
|
||||||
|
[INFO] 📥 Feed RSS recebido (162117 bytes).
|
||||||
|
[INFO] 📰 20 artigos extraídos do feed XML.
|
||||||
|
[INFO] 🔗 Decodificando 20 URLs do Google News para os portais reais...
|
||||||
|
[INFO] ✅ 20/20 URLs resolvidas com sucesso para os domínios de origem.
|
||||||
|
[INFO] 💾 Arquivo salvo com sucesso: 'out/river_plate.json' (20 notícias).
|
||||||
|
```
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
# Data Model: Google News Headlines Extractor
|
||||||
|
|
||||||
|
## 1. Entidades de Domínio & DTOs
|
||||||
|
|
||||||
|
### 1.1 SearchQuery (Parâmetros da Busca)
|
||||||
|
Representa os parâmetros de entrada sanitizados e validados para a consulta ao feed do Google News.
|
||||||
|
|
||||||
|
| Atributo | Tipo | Obrigatório | Padrão | Validação / Regra |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| `keyword` | `str` | Sim | — | Não vazio, sem espaços em branco apenas. |
|
||||||
|
| `language` | `str` | Não | `"pt"` | Mínimo 2 caracteres, normalizado em minúsculas. |
|
||||||
|
| `locale` | `str | None` | Não | `None` | Código ISO alpha-2 de país ou inferido do idioma. |
|
||||||
|
| `max_pages` | `int` | Não | `1` | Intervalo entre `1` e `10` (correspondente a 10 até 100 itens). |
|
||||||
|
|
||||||
|
### 1.2 NewsArticle (Item de Notícia)
|
||||||
|
Representa uma notícia individual extraída do feed RSS.
|
||||||
|
|
||||||
|
| Atributo | Tipo | Descrição | Origem no Feed |
|
||||||
|
| :--- | :--- | :--- | :--- |
|
||||||
|
| `titulo` | `str` | Título da manchete | Nó `<title>` |
|
||||||
|
| `subtitulo` | `str | None` | Resumo textual limpo de tags HTML (ou `None` se ausente/idêntico ao título) | Nó `<description>` sanitizado |
|
||||||
|
| `quando_publicado` | `str | None` | Data original de publicação do feed (RFC 822) | Nó `<pubDate>` |
|
||||||
|
| `url` | `str` | Link de acesso à notícia | Nó `<link>` |
|
||||||
|
| `pagina` | `int` | Número da página lógica calculada | `(idx // 10) + 1` |
|
||||||
|
|
||||||
|
### 1.3 ExtractionResult (Saída Estruturada Consolidada)
|
||||||
|
Pacote consolidado emitido para consumo/stdout.
|
||||||
|
|
||||||
|
| Atributo | Tipo | Descrição |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `query` | `str` | Palavra-chave/expressão pesquisada |
|
||||||
|
| `language` | `str` | Código do idioma utilizado na busca |
|
||||||
|
| `locale` | `str` | Código da região/país utilizado |
|
||||||
|
| `total_paginas` | `int` | Total de páginas lógicas solicitadas/extraídas |
|
||||||
|
| `total_itens` | `int` | Quantidade total de notícias retornadas na lista |
|
||||||
|
| `scraped_at` | `str` | Data/hora ISO 8601 UTC do momento da extração |
|
||||||
|
| `items` | `list[NewsArticle]` | Lista ordenada de notícias |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Esquema JSON de Saída
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"$schema": "http://json-schema.org/draft-07/schema#",
|
||||||
|
"title": "ExtractionResult",
|
||||||
|
"type": "object",
|
||||||
|
"required": ["query", "language", "locale", "total_paginas", "total_itens", "scraped_at", "items"],
|
||||||
|
"properties": {
|
||||||
|
"query": { "type": "string" },
|
||||||
|
"language": { "type": "string" },
|
||||||
|
"locale": { "type": "string" },
|
||||||
|
"total_paginas": { "type": "integer", "minimum": 1, "maximum": 10 },
|
||||||
|
"total_itens": { "type": "integer", "minimum": 0 },
|
||||||
|
"scraped_at": { "type": "string", "format": "date-time" },
|
||||||
|
"items": {
|
||||||
|
"type": "array",
|
||||||
|
"items": {
|
||||||
|
"type": "object",
|
||||||
|
"required": ["titulo", "url", "pagina"],
|
||||||
|
"properties": {
|
||||||
|
"titulo": { "type": "string" },
|
||||||
|
"subtitulo": { "type": ["string", "null"] },
|
||||||
|
"quando_publicado": { "type": ["string", "null"] },
|
||||||
|
"url": { "type": "string", "format": "uri" },
|
||||||
|
"pagina": { "type": "integer", "minimum": 1 }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
# Implementation Plan: Google News Headlines Extractor
|
||||||
|
|
||||||
|
**Branch**: `002-google-news-extractor` | **Date**: 2026-08-20 | **Spec**: [spec.md](./spec.md)
|
||||||
|
|
||||||
|
**Input**: Feature specification from `specs/002-google-news-extractor/spec.md`
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
Implementar um script CLI Python simples, robusto e autônomo ([`scripts/extract_google_news.py`](file:///c:/Users/aferr/Projects/AFTech/DunaMedia/TextNLPClassifierApp/scripts/extract_google_news.py)) para extrair manchetes do Google News RSS por assunto, idioma e região geográfica (*locale*). A solução utiliza a biblioteca [`foxcape`](https://pypi.org/project/foxcape/) em modo `headless=True` para garantir evasão anti-bot stealth, a biblioteca [`googlenewsdecoder`](https://pypi.org/project/googlenewsdecoder/) para resolver automaticamente as URLs intermediárias para os links finais dos portais de notícias em paralelo (`ThreadPoolExecutor`), e logging em tempo real via `sys.stderr`.
|
||||||
|
|
||||||
|
## Technical Context
|
||||||
|
|
||||||
|
**Language/Version**: Python >= 3.10 (3.11 / 3.12 / 3.13 / 3.14)
|
||||||
|
**Primary Dependencies**:
|
||||||
|
* `foxcape>=0.1.1` (scraping stealth & evasão anti-bot Camoufox em modo headless)
|
||||||
|
* `googlenewsdecoder>=0.1.7` (decodificação de URLs intermediárias do Google News)
|
||||||
|
* `selectolax>=0.3.27` (parser ultra-rápido de nós e atributos)
|
||||||
|
* `beautifulsoup4>=4.12.0` (limpeza HTML e higienização de resumos/snippets)
|
||||||
|
* `argparse` (parser CLI na biblioteca padrão)
|
||||||
|
* `concurrent.futures` (resolução paralela em pool de threads)
|
||||||
|
|
||||||
|
**Storage**: N/A (stateless; saída via `stdout` ou arquivo especificado por flag `-o / --output`)
|
||||||
|
**Testing**: `pytest` com fixtures de feeds RSS, testes unitários mockados e testes End-to-End (E2E) ao vivo
|
||||||
|
**Target Platform**: Multiplataforma (Windows / Linux / macOS)
|
||||||
|
**Project Type**: Standalone CLI script + módulo utilitário de scripts
|
||||||
|
**Constraints**: Separação estrita de streams (`stdout` para JSON e `stderr` para logs informativos/erros)
|
||||||
|
|
||||||
|
## Architecture & Pipeline
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
A[SearchQuery CLI] --> B[Foxcape Headless Fetch]
|
||||||
|
B --> C[parse_google_news_rss XML/HTML]
|
||||||
|
C --> D[ThreadPoolExecutor URL Resolution]
|
||||||
|
D --> E[ExtractionResult JSON Output]
|
||||||
|
E --> F[stdout / File Output]
|
||||||
|
```
|
||||||
|
|
||||||
|
1. **Ingestão & Validação**: `SearchQuery` valida palavra-chave não-vazia, idioma e limites de páginas (1 a 10).
|
||||||
|
2. **Coleta RSS via Foxcape**: `Foxcape.fetch(url, config=FoxcapeConfig(headless=True))` busca o feed com evasões anti-bot.
|
||||||
|
3. **Higienização XML/HTML**: `parse_google_news_rss` extrai nós, limpa resumos com `BeautifulSoup` e deduplica subtítulos redundantes.
|
||||||
|
4. **Decodificação de Links**: `resolve_articles_urls` decodifica as URLs intermediárias do Google News em paralelo via `googlenewsdecoder`.
|
||||||
|
5. **Emissão Estruturada**: JSON formatado no `stdout` ou gravado em arquivo (`--output`).
|
||||||
|
|
||||||
|
## Project Structure
|
||||||
|
|
||||||
|
### Documentation (this feature)
|
||||||
|
|
||||||
|
```text
|
||||||
|
specs/002-google-news-extractor/
|
||||||
|
├── spec.md # Especificação de requisitos da feature
|
||||||
|
├── plan.md # Este plano de implementação
|
||||||
|
├── research.md # Pesquisa técnica e decisões de design
|
||||||
|
├── data-model.md # Entidades, DTOs e esquema JSON
|
||||||
|
├── contracts/
|
||||||
|
│ └── cli_contract.md # Contrato de argumentos CLI e I/O streams
|
||||||
|
├── quickstart.md # Guia rápido de execução e validação
|
||||||
|
└── checklists/
|
||||||
|
├── requirements.md # Checklist de qualidade dos requisitos
|
||||||
|
└── readiness.md # Checklist de prontidão e completude
|
||||||
|
```
|
||||||
|
|
||||||
|
### Source Code
|
||||||
|
|
||||||
|
```text
|
||||||
|
scripts/
|
||||||
|
├── __init__.py # Identificador de pacote Python
|
||||||
|
└── extract_google_news.py # Script CLI principal e lógica de extração
|
||||||
|
|
||||||
|
tests/
|
||||||
|
├── fixtures/
|
||||||
|
│ └── google_news_sample.xml # Amostra de feed RSS para testes offline
|
||||||
|
└── test_extract_google_news.py# Testes unitários, integração e E2E ao vivo
|
||||||
|
```
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
# Quickstart: Google News Headlines Extractor
|
||||||
|
|
||||||
|
Guia rápido para execução e validação de ponta a ponta do extrator de notícias via linha de comando.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Pré-requisitos e Instalação
|
||||||
|
|
||||||
|
Instale as dependências necessárias no ambiente Python:
|
||||||
|
```bash
|
||||||
|
pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
Baixe os binários de browser stealth do Camoufox (executado uma única vez):
|
||||||
|
```bash
|
||||||
|
python -m camoufox fetch
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Cenários Práticos de Uso
|
||||||
|
|
||||||
|
### Cenário 1: River Plate — Argentina (Espanhol / 2 Páginas / Salvar em Arquivo)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "River Plate" -l es --locale AR -p 2 -o out/river_plate.json
|
||||||
|
```
|
||||||
|
* **Logs no terminal**:
|
||||||
|
```text
|
||||||
|
[INFO] 🔍 Consultando Google News: 'River Plate' (idioma: es, locale: AR, max_pages: 2)...
|
||||||
|
[INFO] 📥 Feed RSS recebido (162117 bytes).
|
||||||
|
[INFO] 📰 20 artigos extraídos do feed XML.
|
||||||
|
[INFO] 🔗 Decodificando 20 URLs do Google News para os portais reais...
|
||||||
|
[INFO] ✅ 20/20 URLs resolvidas com sucesso para os domínios de origem.
|
||||||
|
[INFO] 💾 Arquivo salvo com sucesso: 'out/river_plate.json' (20 notícias).
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Cenário 2: Cruzeiro — Brasil (Português / Formatado no Terminal)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py --query "Cruzeiro" --lang pt --locale BR --pretty
|
||||||
|
```
|
||||||
|
* Retorna JSON formatado com 10 manchetes e links diretos (*ge.globo.com, lance.com.br, gazetaesportiva.com*).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Cenário 3: Fórmula 1 — Reino Unido (Inglês)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py --query "Formula 1" --lang en --locale GB --pretty
|
||||||
|
```
|
||||||
|
* Retorna notícias de veículos britânicos (*BBC Sport, Sky Sports F1, Autosport*).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Cenário 4: Integração em Pipeline com `jq` (Modo Silencioso)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "inteligência artificial" -s | jq '.items[].url'
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Cenário 5: Extração Rápida com Links Brutos (Sem Resolução de URLs)
|
||||||
|
```bash
|
||||||
|
python scripts/extract_google_news.py -q "São Paulo" --no-resolve-urls --pretty
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Validação dos Testes Automatizados e Linter
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Executa todos os testes unitários e E2E ao vivo
|
||||||
|
pytest tests/test_extract_google_news.py -v
|
||||||
|
|
||||||
|
# Validação estática com Ruff e Mypy
|
||||||
|
ruff check scripts/extract_google_news.py tests/test_extract_google_news.py
|
||||||
|
mypy scripts/ tests/test_extract_google_news.py
|
||||||
|
```
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
# Research: Google News Headlines Extractor
|
||||||
|
|
||||||
|
## 1. Technical Decisions & Tradeoffs
|
||||||
|
|
||||||
|
### Decision 1: Motor de Requisição e Scraping com `foxcape` em Modo Headless
|
||||||
|
- **Decision**: Adotar o pacote `foxcape` com `FoxcapeConfig(headless=True, humanize=False)` como motor de requisição primário.
|
||||||
|
- **Rationale**: `foxcape` integra Camoufox e BeautifulSoup com evasões de fingerprinting TLS, runtime JS e headers avançados, impedindo bloqueios (429/403/Captchas) frequentes do Google News. A configuração `headless=True` garante que a execução ocorra 100% em segundo plano sem abrir janelas gráficas no sistema.
|
||||||
|
- **Alternatives Considered**:
|
||||||
|
- `curl_cffi` + `beautifulsoup4` manual: Boa alternativa, mas exige orquestração manual de impersonação de TLS e headers.
|
||||||
|
- `requests` padrão: Alto risco de bloqueio anti-bot pelo Google News.
|
||||||
|
- Foxcape padrão sem configuração (`headless=False`): Abre janela visual do Firefox indesejada em execuções CLI e servidores.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Decision 2: Endpoint RSS do Google News vs. Scraping de DOM
|
||||||
|
- **Decision**: Utilizar o endpoint oficial de busca RSS do Google News: `https://news.google.com/rss/search?q={query}&hl={hl}&gl={gl}&ceid={gl}:{hl}`.
|
||||||
|
- **Rationale**: Formato estruturado em XML padrão, com carregamento rápido e direto de todos os metadados necessários (`title`, `link`, `pubDate`, `description`), sem necessidade de lidar com seletores CSS voláteis da interface web renderizada.
|
||||||
|
- **Alternatives Considered**:
|
||||||
|
- Scraping direto da interface HTML do Google News (`news.google.com/search`): Classes CSS ofuscadas e alteradas frequentemente pelo Google, quebrando facilmente a extração.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Decision 3: Mapeamento de Idioma e Locale (`hl`, `gl`, `ceid`)
|
||||||
|
- **Decision**: Tabela de mapeamento determinística com fallback dinâmico.
|
||||||
|
- `pt` → `hl=pt-BR`, `gl=BR`, `ceid=BR:pt-BR`
|
||||||
|
- `es` → `hl=es-419`, `gl=AR`, `ceid=AR:es-419`
|
||||||
|
- `en` → `hl=en-US`, `gl=US`, `ceid=US:en-US`
|
||||||
|
- `de` → `hl=de`, `gl=DE`, `ceid=DE:de`
|
||||||
|
- `fr` → `hl=fr`, `gl=FR`, `ceid=FR:fr`
|
||||||
|
- `it` → `hl=it`, `gl=IT`, `ceid=IT:it`
|
||||||
|
- Customizado: se fornecido `--locale MX`, sobrescreve o `gl` e ajusta `ceid={gl}:{hl}`.
|
||||||
|
- **Rationale**: Garante notícias contextualmente adequadas por país sem que o usuário precise memorizar os códigos técnicos internos do Google News.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Decision 4: Resolução de URLs do Google News via `googlenewsdecoder`
|
||||||
|
- **Decision**: Resolver automaticamente as URLs intermediárias (`news.google.com/rss/articles/CBMi...`) para os links originais dos veículos de imprensa em lote com `concurrent.futures.ThreadPoolExecutor(max_workers=5)`.
|
||||||
|
- **Rationale**: Os links gerados pelo Google News contêm tokens RPC intermediários que dificultam a leitura e ingestão direta. A decodificação em lote resolve até 50 URLs em menos de 1 segundo sem sobrecarga.
|
||||||
|
- **Alternatives Considered**:
|
||||||
|
- Resolução via Playwright headless para cada link: Muito lenta para listas de 20 a 50 notícias (demora 30 a 60 segundos).
|
||||||
|
- Manter apenas a URL do Google News: Prejudica o usuário e sistemas downstream que precisam do domínio e link real do portal de notícias.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Decision 5: Logging em Tempo Real no `stderr` e Segregação de Streams
|
||||||
|
- **Decision**: Enviar mensagens de status (`[INFO] ...`) para `sys.stderr` e reservar `sys.stdout` exclusivamente para o JSON.
|
||||||
|
- **Rationale**: Permite que o operador acompanhe o progresso em tempo real no terminal (`Consultando...`, `Decodificando URLs...`, `Arquivo salvo...`) sem quebrar a interoperabilidade com ferramentas de pipe como `jq` ou redirecionamentos de arquivo.
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
# Feature Specification: Google News Headlines Extractor
|
||||||
|
|
||||||
|
**Feature Branch**: `002-google-news-extractor`
|
||||||
|
**Created**: 2026-08-20
|
||||||
|
**Status**: Implemented & Validated
|
||||||
|
**Input**: User description: "Extrator de manchetes do Google News de acordo com assunto, idioma, locale, com Foxcape headless, resolução de URLs reais e logging"
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Clarifications
|
||||||
|
|
||||||
|
### Session 2026-08-20
|
||||||
|
- Q: Como o extrator de notícias do Google News deve ser disponibilizado e consumido dentro do projeto? → A: Apenas Script CLI autônomo para execução direta via linha de comando no terminal (`scripts/extract_google_news.py`).
|
||||||
|
- Q: Qual biblioteca/mecanismo de requisição e raspagem deve ser utilizado no script CLI? → A: Biblioteca `foxcape` (https://pypi.org/project/foxcape/) em modo `headless=True` com proteção anti-bot e fingerprinting stealth.
|
||||||
|
- Q: Como lidar com as URLs intermediárias do Google News? → A: Resolução e decodificação automática das URLs intermediárias (`https://news.google.com/rss/articles/...`) para as URLs reais e finais dos portais de notícias via `googlenewsdecoder`.
|
||||||
|
- Q: Como acompanhar o progresso de extração no terminal? → A: Logs informativos em tempo real enviados exclusivamente para `sys.stderr`, mantendo `sys.stdout` limpo para pipes JSON e suportando a flag `-s / --silent`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## User Scenarios & Testing *(mandatory)*
|
||||||
|
|
||||||
|
### User Story 1 - Extração Básica de Notícias com URLs Finais Resolvidas (Priority: P1) 🌟 MVP
|
||||||
|
|
||||||
|
Como analista ou operador no terminal, quero fornecer um termo de busca e um idioma de interesse via linha de comando para obter rapidamente uma lista estruturada de manchetes recentes com as URLs finais reais dos veículos de imprensa (ex: *Olé, TyC Sports, ge, ESPN*).
|
||||||
|
|
||||||
|
**Why this priority**: É o valor fundamental do produto. Sem a capacidade de buscar, obter manchetes e fornecer os links diretos dos veículos, o extrator não cumpre sua função de pesquisa e ingestão.
|
||||||
|
|
||||||
|
**Independent Test**: Executar `python scripts/extract_google_news.py -q "River Plate" -l es --locale AR -p 1` e verificar se o JSON retornado contém notícias com URLs apontando para os domínios finais (`tycsports.com`, `ole.com.ar`, etc.).
|
||||||
|
|
||||||
|
**Acceptance Scenarios**:
|
||||||
|
1. **Given** um termo de busca válido e idioma, **When** o script for executado com o motor `foxcape` em modo headless, **Then** o sistema extrai o feed RSS e decodifica as URLs de cada artigo para os sites de origem.
|
||||||
|
2. **Given** um termo de busca sem notícias correspondentes, **When** o script for executado, **Then** o sistema retorna uma coleção vazia com código de saída 0.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2)
|
||||||
|
|
||||||
|
Como usuário que monitora notícias em mercados específicos, quero definir a região/país geográfica (*locale*) além do idioma (por exemplo, espanhol da Argentina `--locale AR` vs. México `--locale MX`, ou inglês do Reino Unido `--locale GB` vs. Estados Unidos `--locale US`) para receber manchetes contextualmente relevantes àquele território.
|
||||||
|
|
||||||
|
**Why this priority**: Garante relevância e precisão geográfica para análises de mídia multinacionais e segmentadas.
|
||||||
|
|
||||||
|
**Independent Test**: Executar o script com idioma "es" e locale "AR" e verificar que portais e manchetes são da Argentina.
|
||||||
|
|
||||||
|
**Acceptance Scenarios**:
|
||||||
|
1. **Given** um termo de busca, idioma "es" e país/região "AR", **When** o script CLI for executado, **Then** as manchetes retornadas priorizam a edição e veículos argentinos.
|
||||||
|
2. **Given** um idioma informado sem país explícito (ex: "pt"), **When** o script for solicitado, **Then** o sistema aplica o mapeamento padrão correspondente (ex: Brasil / pt-BR).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### User Story 3 - Paginação, Exportação em Arquivo e Feedback Visual (Priority: P3)
|
||||||
|
|
||||||
|
Como operador de automação de dados, quero parametrizar o número de páginas de resultados (`--max-pages 1..10`), exportar direto para arquivo (`--output`) e acompanhar o progresso no terminal com mensagens descritivas de log.
|
||||||
|
|
||||||
|
**Why this priority**: Permite flexibilidade de uso em rotinas batch, pipelines de ETL e depuração interativa.
|
||||||
|
|
||||||
|
**Independent Test**: Executar com `-p 2 -o out/resultado.json` e verificar criação do arquivo com diretórios pais automáticos e logs em `stderr`.
|
||||||
|
|
||||||
|
**Acceptance Scenarios**:
|
||||||
|
1. **Given** uma execução CLI com `-p 2 -o out/teste.json`, **When** o processo é executado, **Then** o terminal exibe logs em `stderr` (`[INFO] 🔍 Consultando...`, `[INFO] 🔗 Decodificando...`, `[INFO] 💾 Arquivo salvo...`) e grava o JSON final com 20 itens.
|
||||||
|
2. **Given** o uso da flag `-s` ou `--silent`, **When** o script é executado, **Then** os logs em `stderr` são suprimidos.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- **Termo de busca vazio ou composto apenas por espaços**: O script rejeita a solicitação com mensagem de erro em `stderr` e código de saída 1.
|
||||||
|
- **Falha de conectividade ou bloqueio**: O Foxcape executa com evasões anti-bot em modo headless, com captura de exceções e emissão de erro em `stderr` com código de saída 2.
|
||||||
|
- **Resumo da notícia idêntico ao título**: Subtítulo redundante vira `None`.
|
||||||
|
- **Caminho de saída em pasta inexistente**: O script cria os diretórios pais automaticamente antes de salvar.
|
||||||
|
- **Falha na decodificação de URL específica**: Mantém a URL original como fallback gracioso sem abortar a execução dos demais itens.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Requirements *(mandatory)*
|
||||||
|
|
||||||
|
### Functional Requirements
|
||||||
|
|
||||||
|
- **FR-001**: O sistema DEVE ser disponibilizado como script CLI autônomo em `scripts/extract_google_news.py`.
|
||||||
|
- **FR-002**: O script DEVE utilizar a biblioteca `foxcape` com `FoxcapeConfig(headless=True)` como motor primário de requisição com proteção anti-bot.
|
||||||
|
- **FR-003**: O CLI DEVE aceitar argumento obrigatório para o termo de busca (`-q, --query, --keyword`).
|
||||||
|
- **FR-004**: O CLI DEVE suportar argumentos opcionais para código de idioma (`-l, --lang, --language`, padrão `pt`) e região/país (`--locale, --country`).
|
||||||
|
- **FR-005**: O sistema DEVE aplicar mapeamento padrão de região quando apenas o idioma for informado (`pt` → BR, `es` → AR, `en` → US, `de` → DE, `fr` → FR, `it` → IT).
|
||||||
|
- **FR-006**: O CLI DEVE permitir configurar a paginação lógica (`-p, --max-pages`, de 1 a 10 páginas / 10 a 100 itens).
|
||||||
|
- **FR-007**: O sistema DEVE extrair: título, URL, data de publicação, subtítulo higienizado de HTML e número da página.
|
||||||
|
- **FR-008**: O sistema DEVE resolver e decodificar automaticamente as URLs intermediárias do Google News para as URLs finais dos portais de notícias em paralelo.
|
||||||
|
- **FR-009**: O CLI DEVE emitir logs informativos de progresso em `sys.stderr` e suportar a flag `-s, --silent` para supressão.
|
||||||
|
- **FR-010**: O CLI DEVE suportar a flag `--no-resolve-urls` para obter as URLs brutas do feed RSS quando desejado.
|
||||||
|
- **FR-011**: O CLI DEVE suportar gravação em arquivo via `-o, --output` com criação de diretórios pais e formatação legível com `--pretty`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Success Criteria *(mandatory)*
|
||||||
|
|
||||||
|
- **SC-001**: Extração executada com sucesso utilizando Foxcape headless e alta taxa de entrega.
|
||||||
|
- **SC-002**: 100% das URLs de notícias decodificadas para os portais reais dos veículos quando a resolução de URLs estiver ativa.
|
||||||
|
- **SC-003**: 100% dos resumos/subtítulos livres de marcações HTML.
|
||||||
|
- **SC-004**: Formato JSON no `stdout` 100% compatível com utilitários como `jq` e pipelines shell.
|
||||||
|
- **SC-005**: 100% dos testes unitários e E2E aprovados no pytest.
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
# Implementation Tasks: Google News Headlines Extractor
|
||||||
|
|
||||||
|
**Feature**: Google News Headlines Extractor
|
||||||
|
**Branch**: `002-google-news-extractor` | **Date**: 2026-08-20
|
||||||
|
**Spec**: [spec.md](./spec.md) | **Plan**: [plan.md](./plan.md) | **Contracts**: [cli_contract.md](./contracts/cli_contract.md)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 1: Setup (Shared Infrastructure)
|
||||||
|
|
||||||
|
**Purpose**: Instalar dependências necessárias e configurar fixtures de teste.
|
||||||
|
|
||||||
|
- [X] T001 Adicionar `foxcape`, `beautifulsoup4`, `googlenewsdecoder` e `selectolax` ao requirements.txt e validar instalação
|
||||||
|
- [X] T002 [P] Criar fixture XML de exemplo em tests/fixtures/google_news_sample.xml para testes offline
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 2: Foundational (Blocking Prerequisites)
|
||||||
|
|
||||||
|
**Purpose**: Estruturas de dados base e utilitários de localização que sustentam todas as histórias de usuário.
|
||||||
|
|
||||||
|
- [X] T003 Definir as dataclasses `SearchQuery`, `NewsArticle` e `ExtractionResult` em scripts/extract_google_news.py
|
||||||
|
- [X] T004 [P] Implementar utilitário de mapeamento de idioma e país `get_hl_gl_ceid` em scripts/extract_google_news.py
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 3: User Story 1 - Extração Básica de Notícias por Assunto e Idioma (Priority: P1) 🌟 MVP
|
||||||
|
|
||||||
|
**Goal**: Permitir a extração de notícias de um termo e idioma via CLI, retornando JSON formatado com títulos, links e datas de publicação.
|
||||||
|
|
||||||
|
**Independent Test**: Executar `python scripts/extract_google_news.py --query "tecnologia" --lang pt` e verificar retorno de JSON válido no `stdout` contendo itens com título e URL.
|
||||||
|
|
||||||
|
### Tests for User Story 1 🧪
|
||||||
|
- [X] T005 [P] [US1] Criar testes unitários para parsing XML do RSS e limpeza de tags HTML em tests/test_extract_google_news.py
|
||||||
|
|
||||||
|
### Implementation for User Story 1
|
||||||
|
- [X] T006 [US1] Implementar função `parse_google_news_rss` com limpeza de tags HTML e deduplicação de subtítulos em scripts/extract_google_news.py
|
||||||
|
- [X] T007 [US1] Implementar função `extract_google_news` utilizando a biblioteca `foxcape` em modo headless como motor primário em scripts/extract_google_news.py
|
||||||
|
- [X] T008 [US1] Implementar interface CLI básica com `argparse` emitindo JSON para `stdout` em scripts/extract_google_news.py
|
||||||
|
|
||||||
|
**Checkpoint**: User Story 1 (MVP) 100% funcional e testável de forma independente.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 4: User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2)
|
||||||
|
|
||||||
|
**Goal**: Permitir direcionamento de notícias por região/país com a flag `--locale` (ex: `es` com `MX` vs `ES`).
|
||||||
|
|
||||||
|
**Independent Test**: Executar `python scripts/extract_google_news.py --query "futebol" --lang es --locale MX` e verificar que a query foi montada com os parâmetros de edição regional correspondentes.
|
||||||
|
|
||||||
|
### Tests for User Story 2 🧪
|
||||||
|
- [X] T009 [P] [US2] Adicionar testes unitários para a flag `--locale` e resolução de `ceid` em tests/test_extract_google_news.py
|
||||||
|
|
||||||
|
### Implementation for User Story 2
|
||||||
|
- [X] T010 [US2] Integrar suporte à flag `--locale` / `--country` no CLI e fluxo de requisição em scripts/extract_google_news.py
|
||||||
|
|
||||||
|
**Checkpoint**: User Stories 1 e 2 funcionais e testáveis de forma independente.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 5: User Story 3 - Paginação e Controle de Volume de Resultados (Priority: P3)
|
||||||
|
|
||||||
|
**Goal**: Permitir configuração de quantidade de páginas/itens (`--max-pages` de 1 a 10) e salvamento em arquivo (`--output`).
|
||||||
|
|
||||||
|
**Independent Test**: Executar com `--max-pages 2 --output out/test_news.json` e verificar arquivo gerado com até 20 itens divididos em páginas 1 e 2.
|
||||||
|
|
||||||
|
### Tests for User Story 3 🧪
|
||||||
|
- [X] T011 [P] [US3] Adicionar testes unitários para limite de páginas e exportação em arquivo em tests/test_extract_google_news.py
|
||||||
|
|
||||||
|
### Implementation for User Story 3
|
||||||
|
- [X] T012 [US3] Implementar fatiamento de páginas lógicas (`--max-pages 1..10`) e flag `--output` para gravação de arquivo com criação automática de diretórios pais em scripts/extract_google_news.py
|
||||||
|
|
||||||
|
**Checkpoint**: Todas as histórias de usuário funcionais e integradas.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 6: URL Resolution, Headless Configuration & Real-Time Logging
|
||||||
|
|
||||||
|
**Purpose**: Resolver URLs reais dos veículos, suprimir janelas visuais de navegador e emitir feedback de progresso no terminal.
|
||||||
|
|
||||||
|
- [X] T013 Implementar `resolve_article_url` e `resolve_articles_urls` com `googlenewsdecoder` e pool concorrente em scripts/extract_google_news.py
|
||||||
|
- [X] T014 Configurar `FoxcapeConfig(headless=True)` garantindo raspagem 100% em background sem interface gráfica em scripts/extract_google_news.py
|
||||||
|
- [X] T015 Adicionar logging em tempo real em `sys.stderr` e flag `-s / --silent` para supressão em scripts/extract_google_news.py
|
||||||
|
- [X] T016 Adicionar flag `--no-resolve-urls` para permitir acesso às URLs brutas do feed quando desejado em scripts/extract_google_news.py
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 7: Verification & E2E Testing
|
||||||
|
|
||||||
|
**Purpose**: Testes automatizados de ponta a ponta e auditoria de tipagem e estilo.
|
||||||
|
|
||||||
|
- [X] T017 Criar testes unitários para o resolvedor de URLs em tests/test_extract_google_news.py
|
||||||
|
- [X] T018 Criar testes E2E ao vivo (`test_e2e_resolve_real_google_news_url`, `test_e2e_extract_google_news_live_pipeline`, `test_e2e_cli_live_file_output`) em tests/test_extract_google_news.py
|
||||||
|
- [X] T019 Validar 100% de conformidade com `ruff check`, `ruff format` e `mypy` estrito
|
||||||
|
- [X] T020 Executar suite completa do repositório garantindo 0 regressões
|
||||||
+42
@@ -0,0 +1,42 @@
|
|||||||
|
<?xml version="1.0" encoding="UTF-8"?>
|
||||||
|
<rss version="2.0" xmlns:media="http://search.yahoo.com/mrss/">
|
||||||
|
<channel>
|
||||||
|
<generator>NFE/5.0</generator>
|
||||||
|
<title>"inteligencia artificial" - Google Notícias</title>
|
||||||
|
<link>https://news.google.com/search?q=inteligencia+artificial&hl=pt-BR&gl=BR&ceid=BR:pt-BR</link>
|
||||||
|
<language>pt-BR</language>
|
||||||
|
<webMaster>news-webmaster@google.com</webMaster>
|
||||||
|
<copyright>2026 Google Inc.</copyright>
|
||||||
|
<lastBuildDate>Thu, 20 Aug 2026 13:00:00 GMT</lastBuildDate>
|
||||||
|
<description>Google Notícias</description>
|
||||||
|
<item>
|
||||||
|
<title>Empresas aceleram adoção de inteligência artificial no Brasil - InfoMoney</title>
|
||||||
|
<link>https://news.google.com/rss/articles/CBMiYWh0dHBzOi8vd3d3LmluZm9tb25leS5jb20uYnIvbmVnb2Npb3MvYWRvcGNhby1kZS1pYS1uby1icmFzaWwtY3Jlc2NlLTQwLXF1YWRyby1lc3BlY2lhbC1zZXRvci8</link>
|
||||||
|
<guid isPermaLink="false">CBMiYWh0dHBzOi8vd3d3LmluZm9tb25leS5jb20uYnIvbmVnb2Npb3MvYWRvcGNhby1kZS1pYS1uby1icmFzaWwtY3Jlc2NlLTQwLXF1YWRyby1lc3BlY2lhbC1zZXRvci8</guid>
|
||||||
|
<pubDate>Thu, 20 Aug 2026 10:30:00 GMT</pubDate>
|
||||||
|
<description><ol><li><a href="https://news.google.com/rss/articles/CBMiYWh0dHBzOi8vd3d3LmluZm9tb25leS5jb20uYnIvbmVnb2Npb3MvYWRvcGNhby1kZS1pYS1uby1icmFzaWwtY3Jlc2NlLTQwLXF1YWRyby1lc3BlY2lhbC1zZXRvci8" target="_blank">Empresas aceleram adoção de inteligência artificial no Brasil</a>&nbsp;&nbsp;<font color="#6f6f6f">InfoMoney</font><p>Levantamento mostra crescimento expressivo na automação e ganhos de produtividade empresarial.</p></li></ol></description>
|
||||||
|
<source url="https://www.infomoney.com.br">InfoMoney</source>
|
||||||
|
</item>
|
||||||
|
<item>
|
||||||
|
<title>Nova IA generativa promete revolucionar diagnósticos médicos - G1</title>
|
||||||
|
<link>https://news.google.com/rss/articles/CBMiX2h0dHBzOi8vZzEuZ2xvYm8uY29tL3RlY25vbG9naWEvbm90aWNpYS8yMDI2LzA4LzIwL25vdmEtaWEtZ2VuZXJhdGl2YS1kaWFnbm9zdGljb3MtbWVkaWNvcy5naHRtbA</link>
|
||||||
|
<guid isPermaLink="false">CBMiX2h0dHBzOi8vZzEuZ2xvYm8uY29tL3RlY25vbG9naWEvbm90aWNpYS8yMDI2LzA4LzIwL25vdmEtaWEtZ2VuZXJhdGl2YS1kaWFnbm9zdGljb3MtbWVkaWNvcy5naHRtbA</guid>
|
||||||
|
<pubDate>Thu, 20 Aug 2026 09:15:00 GMT</pubDate>
|
||||||
|
<description><p>Pesquisadores anunciam modelo capaz de detectar patologias precoces com alta acurácia.</p></description>
|
||||||
|
<source url="https://g1.globo.com">G1</source>
|
||||||
|
</item>
|
||||||
|
<item>
|
||||||
|
<title>IA e o futuro do trabalho: desafios e regulamentações - Folha de S.Paulo</title>
|
||||||
|
<link>https://news.google.com/rss/articles/CBMiWWh0dHBzOi8vd3d3MS5mb2xoYS51b2wuY29tLmJyL21lcmNhZG8vMjAyNi8wOC9pYS1lLW8tZnV0dXJvLWRvLXRyYWJhbGhvLWRlc2FmaW9zLnNodG1s</link>
|
||||||
|
<guid isPermaLink="false">CBMiWWh0dHBzOi8vd3d3MS5mb2xoYS51b2wuY29tLmJyL21lcmNhZG8vMjAyNi8wOC9pYS1lLW8tZnV0dXJvLWRvLXRyYWJhbGhvLWRlc2FmaW9zLnNodG1s</guid>
|
||||||
|
<pubDate>Thu, 20 Aug 2026 08:00:00 GMT</pubDate>
|
||||||
|
<description>IA e o futuro do trabalho: desafios e regulamentações - Folha de S.Paulo</description>
|
||||||
|
<source url="https://www1.folha.uol.com.br">Folha de S.Paulo</source>
|
||||||
|
</item>
|
||||||
|
<item>
|
||||||
|
<title>Item sem link</title>
|
||||||
|
<pubDate>Thu, 20 Aug 2026 07:00:00 GMT</pubDate>
|
||||||
|
<description>Este item não deve ser retornado pois falta link.</description>
|
||||||
|
</item>
|
||||||
|
</channel>
|
||||||
|
</rss>
|
||||||
@@ -0,0 +1,355 @@
|
|||||||
|
"""
|
||||||
|
Testes unitários e de integração para o Extrator de Manchetes do Google News.
|
||||||
|
|
||||||
|
Cobre validação de entrada, mapeamento de idiomas/locales, parsing e
|
||||||
|
higienização de XML/HTML, orquestração, decodificação de URLs e contrato de execução CLI.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from scripts.extract_google_news import (
|
||||||
|
ExtractionResult,
|
||||||
|
NewsArticle,
|
||||||
|
SearchQuery,
|
||||||
|
extract_google_news,
|
||||||
|
get_hl_gl_ceid,
|
||||||
|
main,
|
||||||
|
parse_google_news_rss,
|
||||||
|
resolve_article_url,
|
||||||
|
resolve_articles_urls,
|
||||||
|
)
|
||||||
|
|
||||||
|
FIXTURE_PATH = Path(__file__).parent / "fixtures" / "google_news_sample.xml"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def sample_rss_xml() -> str:
|
||||||
|
"""Fixture que fornece o conteúdo do XML de exemplo para testes offline."""
|
||||||
|
return FIXTURE_PATH.read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_hl_gl_ceid_default_mappings():
|
||||||
|
"""Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)."""
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("pt")
|
||||||
|
assert hl == "pt-BR"
|
||||||
|
assert gl == "BR"
|
||||||
|
assert ceid == "BR:pt-BR"
|
||||||
|
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("en")
|
||||||
|
assert hl == "en-US"
|
||||||
|
assert gl == "US"
|
||||||
|
assert ceid == "US:en-US"
|
||||||
|
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("es")
|
||||||
|
assert hl == "es-419"
|
||||||
|
assert gl == "AR"
|
||||||
|
assert ceid == "AR:es-419"
|
||||||
|
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("de")
|
||||||
|
assert hl == "de"
|
||||||
|
assert gl == "DE"
|
||||||
|
assert ceid == "DE:de"
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_hl_gl_ceid_with_custom_locale():
|
||||||
|
"""Valida a sobrescrita geográfica quando o argumento locale é especificado."""
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("es", locale="MX")
|
||||||
|
assert hl == "es-419"
|
||||||
|
assert gl == "MX"
|
||||||
|
assert ceid == "MX:es-419"
|
||||||
|
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("en", locale="GB")
|
||||||
|
assert hl == "en-GB"
|
||||||
|
assert gl == "GB"
|
||||||
|
assert ceid == "GB:en-GB"
|
||||||
|
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("es", locale="ES")
|
||||||
|
assert hl == "es"
|
||||||
|
assert gl == "ES"
|
||||||
|
assert ceid == "ES:es"
|
||||||
|
|
||||||
|
|
||||||
|
def test_get_hl_gl_ceid_dynamic_fallback():
|
||||||
|
"""Valida fallback dinâmico para idiomas regionais não listados explicitamente."""
|
||||||
|
hl, gl, ceid = get_hl_gl_ceid("ja_jp")
|
||||||
|
assert hl == "ja-JP"
|
||||||
|
assert gl == "JP"
|
||||||
|
assert ceid == "JP:ja-JP"
|
||||||
|
|
||||||
|
|
||||||
|
def test_search_query_validation():
|
||||||
|
"""Valida as regras de negócio e limites de SearchQuery."""
|
||||||
|
# Instanciação válida
|
||||||
|
q = SearchQuery(keyword="inteligencia artificial", language="pt", max_pages=1)
|
||||||
|
assert q.clean_keyword == "inteligencia artificial"
|
||||||
|
assert q.clean_language == "pt"
|
||||||
|
assert q.clean_locale is None
|
||||||
|
assert q.max_pages == 1
|
||||||
|
|
||||||
|
# Palavra-chave vazia ou apenas espaços deve lançar ValueError
|
||||||
|
with pytest.raises(ValueError, match="palavra-chave"):
|
||||||
|
SearchQuery(keyword=" ", language="pt")
|
||||||
|
|
||||||
|
# Idioma com menos de 2 caracteres deve lançar ValueError
|
||||||
|
with pytest.raises(ValueError, match="idioma"):
|
||||||
|
SearchQuery(keyword="test", language="p")
|
||||||
|
|
||||||
|
# Intervalo de páginas fora de 1..10 deve lançar ValueError
|
||||||
|
with pytest.raises(ValueError, match="páginas"):
|
||||||
|
SearchQuery(keyword="test", language="pt", max_pages=0)
|
||||||
|
|
||||||
|
with pytest.raises(ValueError, match="páginas"):
|
||||||
|
SearchQuery(keyword="test", language="pt", max_pages=11)
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_google_news_rss_with_fixture(sample_rss_xml: str):
|
||||||
|
"""Valida o parsing do feed RSS, higienização de tags HTML e deduplicação."""
|
||||||
|
articles = parse_google_news_rss(sample_rss_xml, max_pages=1)
|
||||||
|
|
||||||
|
# 4 itens no fixture, mas 1 não possui link -> exatamente 3 válidos
|
||||||
|
assert len(articles) == 3
|
||||||
|
|
||||||
|
# Artigo 1: InfoMoney com HTML no description que deve ser limpo
|
||||||
|
art1 = articles[0]
|
||||||
|
assert "InfoMoney" in art1.titulo
|
||||||
|
assert art1.url.startswith("https://news.google.com/rss/articles/")
|
||||||
|
assert art1.quando_publicado == "Thu, 20 Aug 2026 10:30:00 GMT"
|
||||||
|
assert art1.pagina == 1
|
||||||
|
assert "<" not in (art1.subtitulo or "")
|
||||||
|
assert ">" not in (art1.subtitulo or "")
|
||||||
|
assert "crescimento expressivo" in (art1.subtitulo or "")
|
||||||
|
|
||||||
|
# Artigo 2: G1 com parágrafos limpos
|
||||||
|
art2 = articles[1]
|
||||||
|
assert "G1" in art2.titulo
|
||||||
|
assert "<p>" not in (art2.subtitulo or "")
|
||||||
|
|
||||||
|
# Artigo 3: Folha com descrição redundante/igual ao título -> subtitulo deve ser None
|
||||||
|
art3 = articles[2]
|
||||||
|
assert art3.subtitulo is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_resolve_article_url_fallback():
|
||||||
|
"""Valida fallback gracioso de URL quando não é link do Google News ou em erro."""
|
||||||
|
direct_url = "https://www.globo.com/noticia/123"
|
||||||
|
assert resolve_article_url(direct_url) == direct_url
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"scripts.extract_google_news.gnewsdecoder", return_value={"status": False}
|
||||||
|
):
|
||||||
|
gn_url = "https://news.google.com/rss/articles/fake_token"
|
||||||
|
assert resolve_article_url(gn_url) == gn_url
|
||||||
|
|
||||||
|
|
||||||
|
def test_resolve_article_url_success():
|
||||||
|
"""Valida resolução bem-sucedida de URL do Google News para o portal destino."""
|
||||||
|
gn_url = "https://news.google.com/rss/articles/valid_token"
|
||||||
|
dest_url = "https://infomoney.com.br/mercados/artigo-ia"
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"scripts.extract_google_news.gnewsdecoder",
|
||||||
|
return_value={"status": True, "decoded_url": dest_url},
|
||||||
|
):
|
||||||
|
resolved = resolve_article_url(gn_url)
|
||||||
|
assert resolved == dest_url
|
||||||
|
|
||||||
|
|
||||||
|
def test_resolve_articles_urls_batch():
|
||||||
|
"""Valida a resolução concorrente em lote de uma lista de NewsArticle."""
|
||||||
|
articles = [
|
||||||
|
NewsArticle(
|
||||||
|
titulo="Notícia 1",
|
||||||
|
url="https://news.google.com/rss/articles/1",
|
||||||
|
pagina=1,
|
||||||
|
),
|
||||||
|
NewsArticle(
|
||||||
|
titulo="Notícia 2",
|
||||||
|
url="https://news.google.com/rss/articles/2",
|
||||||
|
pagina=1,
|
||||||
|
),
|
||||||
|
]
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"scripts.extract_google_news.resolve_article_url",
|
||||||
|
side_effect=lambda u: f"https://destinofinal.com/{u.split('/')[-1]}",
|
||||||
|
):
|
||||||
|
resolved = resolve_articles_urls(articles)
|
||||||
|
assert len(resolved) == 2
|
||||||
|
assert resolved[0].url == "https://destinofinal.com/1"
|
||||||
|
assert resolved[1].url == "https://destinofinal.com/2"
|
||||||
|
|
||||||
|
|
||||||
|
def test_extract_google_news_orchestration_mocked(sample_rss_xml: str):
|
||||||
|
"""Valida a consolidação do ExtractionResult a partir da busca mockada com URLs resolvidas."""
|
||||||
|
query = SearchQuery(
|
||||||
|
keyword="inteligência artificial", language="pt", locale="BR", max_pages=1
|
||||||
|
)
|
||||||
|
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news._fetch_rss_content",
|
||||||
|
return_value=sample_rss_xml,
|
||||||
|
),
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news.resolve_article_url",
|
||||||
|
side_effect=lambda u: f"https://resolved.com/{u[-5:]}",
|
||||||
|
),
|
||||||
|
):
|
||||||
|
result = extract_google_news(query, resolve_urls=True)
|
||||||
|
|
||||||
|
assert isinstance(result, ExtractionResult)
|
||||||
|
assert result.query == "inteligência artificial"
|
||||||
|
assert result.language == "pt"
|
||||||
|
assert result.locale == "BR"
|
||||||
|
assert result.total_paginas == 1
|
||||||
|
assert result.total_itens == 3
|
||||||
|
assert len(result.items) == 3
|
||||||
|
assert result.scraped_at is not None
|
||||||
|
assert result.items[0].url.startswith("https://resolved.com/")
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_execution_stdout(sample_rss_xml: str, capsys: pytest.CaptureFixture[str]):
|
||||||
|
"""Valida execução padrão do CLI com saída JSON no stdout."""
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news._fetch_rss_content",
|
||||||
|
return_value=sample_rss_xml,
|
||||||
|
),
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u
|
||||||
|
),
|
||||||
|
):
|
||||||
|
exit_code = main(
|
||||||
|
["--query", "inteligencia artificial", "--lang", "pt", "--pretty"]
|
||||||
|
)
|
||||||
|
assert exit_code == 0
|
||||||
|
|
||||||
|
captured = capsys.readouterr()
|
||||||
|
data = json.loads(captured.out)
|
||||||
|
assert data["query"] == "inteligencia artificial"
|
||||||
|
assert data["total_itens"] == 3
|
||||||
|
assert len(data["items"]) == 3
|
||||||
|
# Validar indentação presente por causa de --pretty
|
||||||
|
assert "\n " in captured.out
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_execution_file_output(sample_rss_xml: str, tmp_path: Path):
|
||||||
|
"""Valida gravação em arquivo com criação automática de diretórios pais."""
|
||||||
|
out_file = tmp_path / "sub_dir" / "news_out.json"
|
||||||
|
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news._fetch_rss_content",
|
||||||
|
return_value=sample_rss_xml,
|
||||||
|
),
|
||||||
|
patch(
|
||||||
|
"scripts.extract_google_news.resolve_article_url", side_effect=lambda u: u
|
||||||
|
),
|
||||||
|
):
|
||||||
|
exit_code = main(["-q", "IA", "-p", "1", "-o", str(out_file)])
|
||||||
|
assert exit_code == 0
|
||||||
|
|
||||||
|
assert out_file.exists()
|
||||||
|
data = json.loads(out_file.read_text(encoding="utf-8"))
|
||||||
|
assert data["total_itens"] == 3
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_empty_query_error(capsys: pytest.CaptureFixture[str]):
|
||||||
|
"""Valida tratamento de erro e código de saída 1 para parâmetro vazio."""
|
||||||
|
exit_code = main(["--query", " "])
|
||||||
|
assert exit_code == 1
|
||||||
|
|
||||||
|
captured = capsys.readouterr()
|
||||||
|
assert "Erro de validação" in captured.err
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_network_error_handling(capsys: pytest.CaptureFixture[str]):
|
||||||
|
"""Valida tratamento de erro e código de saída 2 para falhas de rede."""
|
||||||
|
with patch(
|
||||||
|
"scripts.extract_google_news._fetch_rss_content",
|
||||||
|
side_effect=RuntimeError("Connection refused"),
|
||||||
|
):
|
||||||
|
exit_code = main(["--query", "IA"])
|
||||||
|
assert exit_code == 2
|
||||||
|
|
||||||
|
captured = capsys.readouterr()
|
||||||
|
assert "Erro na extração" in captured.err
|
||||||
|
assert "Connection refused" in captured.err
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Testes E2E (End-to-End) com resolução real de rede e validação de URLs finais
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_e2e_resolve_real_google_news_url():
|
||||||
|
"""Valida E2E que o decodificador resolve uma URL real do Google News para o veículo de imprensa."""
|
||||||
|
# URL real de artigo extraída do Google News RSS
|
||||||
|
sample_gn_url = (
|
||||||
|
"https://news.google.com/rss/articles/"
|
||||||
|
"CBMi6wFBVV95cUxQUS0tMGxacUJycDlCWXpWSGR3T0hfR1E0Q0txWjdqek9LWjNZTUo1WEx3"
|
||||||
|
"UEw3Skt1eW1Wa2VYUUE2ajNYMmpEckhsdGFXUGtmQktYX2JWa1hmclhEZEVIa3hNODhpVXNO"
|
||||||
|
"UGN4cDhmWmJMczBEUVFDNS1aX3EzRHh6VmQ3cVY0ZnZmaW9YWDZJSE9UNFJ2dXpyNFlHaVVs"
|
||||||
|
"VlY0V1FiZ0tzZ3FpRVhUYnhPbmdnOFRveW5oOVB3WDAzS3c0eWFjMDBZSERwNmRkRk1MRHZF"
|
||||||
|
"UE92UE9GMmpRcFZ5cUU2Ym1NeDdQU2tn?oc=5"
|
||||||
|
)
|
||||||
|
|
||||||
|
resolved_url = resolve_article_url(sample_gn_url)
|
||||||
|
|
||||||
|
# Não deve mais ser URL do Google News
|
||||||
|
assert "news.google.com" not in resolved_url
|
||||||
|
# Deve ser uma URL absoluta http/https apontando para o portal real (TyC Sports)
|
||||||
|
assert resolved_url.startswith("http")
|
||||||
|
assert "tycsports.com" in resolved_url
|
||||||
|
|
||||||
|
|
||||||
|
def test_e2e_extract_google_news_live_pipeline():
|
||||||
|
"""Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo."""
|
||||||
|
query = SearchQuery(keyword="tecnologia", language="pt", locale="BR", max_pages=1)
|
||||||
|
result = extract_google_news(query, resolve_urls=True)
|
||||||
|
|
||||||
|
assert isinstance(result, ExtractionResult)
|
||||||
|
assert result.total_itens > 0
|
||||||
|
assert len(result.items) == result.total_itens
|
||||||
|
|
||||||
|
for article in result.items:
|
||||||
|
assert article.titulo
|
||||||
|
assert article.url.startswith("http")
|
||||||
|
# Garante que as URLs foram decodificadas e não permanecem no formato intermediário
|
||||||
|
assert "news.google.com/rss/articles/" not in article.url
|
||||||
|
|
||||||
|
|
||||||
|
def test_e2e_cli_live_file_output(tmp_path: Path):
|
||||||
|
"""Valida E2E a execução do CLI com saída real em arquivo e URLs decodificadas."""
|
||||||
|
out_file = tmp_path / "e2e_result.json"
|
||||||
|
exit_code = main(
|
||||||
|
[
|
||||||
|
"--query",
|
||||||
|
"economia",
|
||||||
|
"--lang",
|
||||||
|
"pt",
|
||||||
|
"--max-pages",
|
||||||
|
"1",
|
||||||
|
"--output",
|
||||||
|
str(out_file),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert exit_code == 0
|
||||||
|
assert out_file.exists()
|
||||||
|
|
||||||
|
data = json.loads(out_file.read_text(encoding="utf-8"))
|
||||||
|
assert data["query"] == "economia"
|
||||||
|
assert data["total_itens"] > 0
|
||||||
|
assert len(data["items"]) > 0
|
||||||
|
|
||||||
|
first_item = data["items"][0]
|
||||||
|
assert first_item["titulo"]
|
||||||
|
assert first_item["url"].startswith("http")
|
||||||
|
assert "news.google.com/rss/articles/" not in first_item["url"]
|
||||||
Reference in New Issue
Block a user