- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
2.9 KiB
2.9 KiB
Data Model: Google News Headlines Extractor
1. Entidades de Domínio & DTOs
1.1 SearchQuery (Parâmetros da Busca)
Representa os parâmetros de entrada sanitizados e validados para a consulta ao feed do Google News.
| Atributo | Tipo | Obrigatório | Padrão | Validação / Regra |
|---|---|---|---|---|
keyword |
str |
Sim | — | Não vazio, sem espaços em branco apenas. |
language |
str |
Não | "pt" |
Mínimo 2 caracteres, normalizado em minúsculas. |
locale |
`str | None` | Não | None |
max_pages |
int |
Não | 1 |
Intervalo entre 1 e 10 (correspondente a 10 até 100 itens). |
1.2 NewsArticle (Item de Notícia)
Representa uma notícia individual extraída do feed RSS.
| Atributo | Tipo | Descrição | Origem no Feed |
|---|---|---|---|
titulo |
str |
Título da manchete | Nó <title> |
subtitulo |
`str | None` | Resumo textual limpo de tags HTML (ou None se ausente/idêntico ao título) |
quando_publicado |
`str | None` | Data original de publicação do feed (RFC 822) |
url |
str |
Link de acesso à notícia | Nó <link> |
pagina |
int |
Número da página lógica calculada | (idx // 10) + 1 |
1.3 ExtractionResult (Saída Estruturada Consolidada)
Pacote consolidado emitido para consumo/stdout.
| Atributo | Tipo | Descrição |
|---|---|---|
query |
str |
Palavra-chave/expressão pesquisada |
language |
str |
Código do idioma utilizado na busca |
locale |
str |
Código da região/país utilizado |
total_paginas |
int |
Total de páginas lógicas solicitadas/extraídas |
total_itens |
int |
Quantidade total de notícias retornadas na lista |
scraped_at |
str |
Data/hora ISO 8601 UTC do momento da extração |
items |
list[NewsArticle] |
Lista ordenada de notícias |
2. Esquema JSON de Saída
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "ExtractionResult",
"type": "object",
"required": ["query", "language", "locale", "total_paginas", "total_itens", "scraped_at", "items"],
"properties": {
"query": { "type": "string" },
"language": { "type": "string" },
"locale": { "type": "string" },
"total_paginas": { "type": "integer", "minimum": 1, "maximum": 10 },
"total_itens": { "type": "integer", "minimum": 0 },
"scraped_at": { "type": "string", "format": "date-time" },
"items": {
"type": "array",
"items": {
"type": "object",
"required": ["titulo", "url", "pagina"],
"properties": {
"titulo": { "type": "string" },
"subtitulo": { "type": ["string", "null"] },
"quando_publicado": { "type": ["string", "null"] },
"url": { "type": "string", "format": "uri" },
"pagina": { "type": "integer", "minimum": 1 }
}
}
}
}
}