Files
TextNLPClassifierApp/specs/002-google-news-extractor/data-model.md
T
andreferraro 6e3d57619b feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping
- Integrate foxcape in headless mode as primary stealth anti-bot engine
- Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor
- Support language and regional locale mapping (-l, --lang, --locale)
- Implement real-time progress logging in stderr and --silent flag
- Add unit, integration, and live E2E tests in tests/test_extract_google_news.py
- Add full SpecKit documentation (specs/002-google-news-extractor/)
- Create comprehensive README.md covering both NLP Classifier and Google News Extractor
2026-08-20 11:50:16 -03:00

2.9 KiB

Data Model: Google News Headlines Extractor

1. Entidades de Domínio & DTOs

1.1 SearchQuery (Parâmetros da Busca)

Representa os parâmetros de entrada sanitizados e validados para a consulta ao feed do Google News.

Atributo Tipo Obrigatório Padrão Validação / Regra
keyword str Sim — Não vazio, sem espaços em branco apenas.
language str Não "pt" Mínimo 2 caracteres, normalizado em minúsculas.
locale `str None` Não None
max_pages int Não 1 Intervalo entre 1 e 10 (correspondente a 10 até 100 itens).

1.2 NewsArticle (Item de Notícia)

Representa uma notícia individual extraída do feed RSS.

Atributo Tipo Descrição Origem no Feed
titulo str Título da manchete Nó <title>
subtitulo `str None` Resumo textual limpo de tags HTML (ou None se ausente/idêntico ao título)
quando_publicado `str None` Data original de publicação do feed (RFC 822)
url str Link de acesso à notícia Nó <link>
pagina int Número da página lógica calculada (idx // 10) + 1

1.3 ExtractionResult (Saída Estruturada Consolidada)

Pacote consolidado emitido para consumo/stdout.

Atributo Tipo Descrição
query str Palavra-chave/expressão pesquisada
language str Código do idioma utilizado na busca
locale str Código da região/país utilizado
total_paginas int Total de páginas lógicas solicitadas/extraídas
total_itens int Quantidade total de notícias retornadas na lista
scraped_at str Data/hora ISO 8601 UTC do momento da extração
items list[NewsArticle] Lista ordenada de notícias

2. Esquema JSON de Saída

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "title": "ExtractionResult",
  "type": "object",
  "required": ["query", "language", "locale", "total_paginas", "total_itens", "scraped_at", "items"],
  "properties": {
    "query": { "type": "string" },
    "language": { "type": "string" },
    "locale": { "type": "string" },
    "total_paginas": { "type": "integer", "minimum": 1, "maximum": 10 },
    "total_itens": { "type": "integer", "minimum": 0 },
    "scraped_at": { "type": "string", "format": "date-time" },
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "required": ["titulo", "url", "pagina"],
        "properties": {
          "titulo": { "type": "string" },
          "subtitulo": { "type": ["string", "null"] },
          "quando_publicado": { "type": ["string", "null"] },
          "url": { "type": "string", "format": "uri" },
          "pagina": { "type": "integer", "minimum": 1 }
        }
      }
    }
  }
}