- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
73 lines
2.9 KiB
Markdown
73 lines
2.9 KiB
Markdown
# Data Model: Google News Headlines Extractor
|
|
|
|
## 1. Entidades de Domínio & DTOs
|
|
|
|
### 1.1 SearchQuery (Parâmetros da Busca)
|
|
Representa os parâmetros de entrada sanitizados e validados para a consulta ao feed do Google News.
|
|
|
|
| Atributo | Tipo | Obrigatório | Padrão | Validação / Regra |
|
|
| :--- | :--- | :--- | :--- | :--- |
|
|
| `keyword` | `str` | Sim | — | Não vazio, sem espaços em branco apenas. |
|
|
| `language` | `str` | Não | `"pt"` | Mínimo 2 caracteres, normalizado em minúsculas. |
|
|
| `locale` | `str | None` | Não | `None` | Código ISO alpha-2 de país ou inferido do idioma. |
|
|
| `max_pages` | `int` | Não | `1` | Intervalo entre `1` e `10` (correspondente a 10 até 100 itens). |
|
|
|
|
### 1.2 NewsArticle (Item de Notícia)
|
|
Representa uma notícia individual extraída do feed RSS.
|
|
|
|
| Atributo | Tipo | Descrição | Origem no Feed |
|
|
| :--- | :--- | :--- | :--- |
|
|
| `titulo` | `str` | Título da manchete | Nó `<title>` |
|
|
| `subtitulo` | `str | None` | Resumo textual limpo de tags HTML (ou `None` se ausente/idêntico ao título) | Nó `<description>` sanitizado |
|
|
| `quando_publicado` | `str | None` | Data original de publicação do feed (RFC 822) | Nó `<pubDate>` |
|
|
| `url` | `str` | Link de acesso à notícia | Nó `<link>` |
|
|
| `pagina` | `int` | Número da página lógica calculada | `(idx // 10) + 1` |
|
|
|
|
### 1.3 ExtractionResult (Saída Estruturada Consolidada)
|
|
Pacote consolidado emitido para consumo/stdout.
|
|
|
|
| Atributo | Tipo | Descrição |
|
|
| :--- | :--- | :--- |
|
|
| `query` | `str` | Palavra-chave/expressão pesquisada |
|
|
| `language` | `str` | Código do idioma utilizado na busca |
|
|
| `locale` | `str` | Código da região/país utilizado |
|
|
| `total_paginas` | `int` | Total de páginas lógicas solicitadas/extraídas |
|
|
| `total_itens` | `int` | Quantidade total de notícias retornadas na lista |
|
|
| `scraped_at` | `str` | Data/hora ISO 8601 UTC do momento da extração |
|
|
| `items` | `list[NewsArticle]` | Lista ordenada de notícias |
|
|
|
|
---
|
|
|
|
## 2. Esquema JSON de Saída
|
|
|
|
```json
|
|
{
|
|
"$schema": "http://json-schema.org/draft-07/schema#",
|
|
"title": "ExtractionResult",
|
|
"type": "object",
|
|
"required": ["query", "language", "locale", "total_paginas", "total_itens", "scraped_at", "items"],
|
|
"properties": {
|
|
"query": { "type": "string" },
|
|
"language": { "type": "string" },
|
|
"locale": { "type": "string" },
|
|
"total_paginas": { "type": "integer", "minimum": 1, "maximum": 10 },
|
|
"total_itens": { "type": "integer", "minimum": 0 },
|
|
"scraped_at": { "type": "string", "format": "date-time" },
|
|
"items": {
|
|
"type": "array",
|
|
"items": {
|
|
"type": "object",
|
|
"required": ["titulo", "url", "pagina"],
|
|
"properties": {
|
|
"titulo": { "type": "string" },
|
|
"subtitulo": { "type": ["string", "null"] },
|
|
"quando_publicado": { "type": ["string", "null"] },
|
|
"url": { "type": "string", "format": "uri" },
|
|
"pagina": { "type": "integer", "minimum": 1 }
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
```
|