feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution

- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping
- Integrate foxcape in headless mode as primary stealth anti-bot engine
- Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor
- Support language and regional locale mapping (-l, --lang, --locale)
- Implement real-time progress logging in stderr and --silent flag
- Add unit, integration, and live E2E tests in tests/test_extract_google_news.py
- Add full SpecKit documentation (specs/002-google-news-extractor/)
- Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
2026-08-20 11:50:16 -03:00
parent 67cc40f91a
commit 6e3d57619b
59 changed files with 16118 additions and 2160 deletions
@@ -0,0 +1,72 @@
# Data Model: Google News Headlines Extractor
## 1. Entidades de Domínio & DTOs
### 1.1 SearchQuery (Parâmetros da Busca)
Representa os parâmetros de entrada sanitizados e validados para a consulta ao feed do Google News.
| Atributo | Tipo | Obrigatório | Padrão | Validação / Regra |
| :--- | :--- | :--- | :--- | :--- |
| `keyword` | `str` | Sim | — | Não vazio, sem espaços em branco apenas. |
| `language` | `str` | Não | `"pt"` | Mínimo 2 caracteres, normalizado em minúsculas. |
| `locale` | `str | None` | Não | `None` | Código ISO alpha-2 de país ou inferido do idioma. |
| `max_pages` | `int` | Não | `1` | Intervalo entre `1` e `10` (correspondente a 10 até 100 itens). |
### 1.2 NewsArticle (Item de Notícia)
Representa uma notícia individual extraída do feed RSS.
| Atributo | Tipo | Descrição | Origem no Feed |
| :--- | :--- | :--- | :--- |
| `titulo` | `str` | Título da manchete | Nó `<title>` |
| `subtitulo` | `str | None` | Resumo textual limpo de tags HTML (ou `None` se ausente/idêntico ao título) | Nó `<description>` sanitizado |
| `quando_publicado` | `str | None` | Data original de publicação do feed (RFC 822) | Nó `<pubDate>` |
| `url` | `str` | Link de acesso à notícia | Nó `<link>` |
| `pagina` | `int` | Número da página lógica calculada | `(idx // 10) + 1` |
### 1.3 ExtractionResult (Saída Estruturada Consolidada)
Pacote consolidado emitido para consumo/stdout.
| Atributo | Tipo | Descrição |
| :--- | :--- | :--- |
| `query` | `str` | Palavra-chave/expressão pesquisada |
| `language` | `str` | Código do idioma utilizado na busca |
| `locale` | `str` | Código da região/país utilizado |
| `total_paginas` | `int` | Total de páginas lógicas solicitadas/extraídas |
| `total_itens` | `int` | Quantidade total de notícias retornadas na lista |
| `scraped_at` | `str` | Data/hora ISO 8601 UTC do momento da extração |
| `items` | `list[NewsArticle]` | Lista ordenada de notícias |
---
## 2. Esquema JSON de Saída
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "ExtractionResult",
"type": "object",
"required": ["query", "language", "locale", "total_paginas", "total_itens", "scraped_at", "items"],
"properties": {
"query": { "type": "string" },
"language": { "type": "string" },
"locale": { "type": "string" },
"total_paginas": { "type": "integer", "minimum": 1, "maximum": 10 },
"total_itens": { "type": "integer", "minimum": 0 },
"scraped_at": { "type": "string", "format": "date-time" },
"items": {
"type": "array",
"items": {
"type": "object",
"required": ["titulo", "url", "pagina"],
"properties": {
"titulo": { "type": "string" },
"subtitulo": { "type": ["string", "null"] },
"quando_publicado": { "type": ["string", "null"] },
"url": { "type": "string", "format": "uri" },
"pagina": { "type": "integer", "minimum": 1 }
}
}
}
}
}
```