feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
@@ -0,0 +1,561 @@
|
||||
# Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||
|
||||
> Documento auto-contido. Pode ser copiado e colado em outra sessão sem memória:
|
||||
> ele contém toda a informação necessária para entender, reproduzir ou portar
|
||||
> o extrator de manchetes do Google News deste repositório (`GoogleNewsETL`).
|
||||
|
||||
---
|
||||
|
||||
## 1. Visão geral (arquitetura)
|
||||
|
||||
O extrator segue **Clean Architecture** em camadas:
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ application/ (casos de uso + DTOs) │
|
||||
│ extract_news_use_case.py → orquestra o fluxo │
|
||||
│ dtos/extract_news_dto.py → contratos de entrada/saída │
|
||||
├──────────────────────────────────────────────────────────────┤
|
||||
│ domain/ (entidades, value objects, portas, serviços) │
|
||||
│ entities/news_article.py → entidade NewsArticle │
|
||||
│ entities/search_query.py → value object SearchQuery │
|
||||
│ ports/news_extractor_port.py → interface NewsExtractorPort│
|
||||
│ ports/url_resolver_port.py → interface UrlResolverPort │
|
||||
│ services/rate_limiter_service.py → throttling │
|
||||
├──────────────────────────────────────────────────────────────┤
|
||||
│ infrastructure/ (adaptadores concretos) │
|
||||
│ adapters/google_news_extractor_adapter.py → ★ o extrator │
|
||||
│ adapters/playwright_url_resolver_adapter.py → resolve URLs│
|
||||
└──────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Princípio chave:** o caso de uso depende da *porta* (`NewsExtractorPort`), nunca do adaptador concreto. O `GoogleNewsExtractorAdapter` implementa essa porta. Isso permite trocar o mecanismo de raspagem (RSS, Playwright, etc.) sem tocar no domínio.
|
||||
|
||||
**Fluxo resumido:**
|
||||
|
||||
```
|
||||
InputDTO ──► ExtractNewsUseCase.execute() ──► SearchQuery (validação)
|
||||
│
|
||||
▼
|
||||
NewsExtractorPort.extract(query)
|
||||
│
|
||||
▼
|
||||
GoogleNewsExtractorAdapter.extract()
|
||||
├─ _fetch_rss() → parseia feed RSS
|
||||
└─ url_resolver.resolve_batch() → URLs finais
|
||||
│
|
||||
▼
|
||||
list[NewsArticle]
|
||||
│
|
||||
▼
|
||||
ExtractNewsOutputDTO (JSON)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Entrada
|
||||
|
||||
### 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`)
|
||||
|
||||
```python
|
||||
class ExtractNewsInputDTO(BaseModel):
|
||||
"""DTO de Entrada do Caso de Uso de Extração."""
|
||||
|
||||
keyword: str = Field(..., description="Palavra ou expressão de busca")
|
||||
language: str = Field(..., description="Código do idioma (ex: 'es', 'pt', 'en')")
|
||||
max_pages: int = Field(default=3, ge=1, le=10, description="Quantidade de páginas para extrair (1 a 10)")
|
||||
```
|
||||
|
||||
| Campo | Tipo | Obrigatório | Descrição |
|
||||
|------------|------|-------------|--------------------------------------------|
|
||||
| `keyword` | str | sim | Palavra/expressão de busca |
|
||||
| `language` | str | sim | Código do idioma (ex: `pt`, `en`, `es`) |
|
||||
| `max_pages`| int | não (def=3) | Páginas a extrair, entre **1 e 10** (10 itens/página) |
|
||||
|
||||
### 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`)
|
||||
|
||||
O use case **nunca usa o DTO cru**: converte-o em `SearchQuery`, que valida as regras de domínio no `__post_init__`:
|
||||
|
||||
```python
|
||||
@dataclass(frozen=True)
|
||||
class SearchQuery:
|
||||
"""Value Object representando os parâmetros validados de consulta."""
|
||||
|
||||
keyword: str
|
||||
language: str
|
||||
max_pages: int = 3
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not self.keyword or not self.keyword.strip():
|
||||
raise InvalidSearchQueryError("A palavra-chave não pode ser vazia.")
|
||||
|
||||
if not self.language or len(self.language.strip()) < 2:
|
||||
raise InvalidSearchQueryError("O idioma deve conter pelo menos 2 caracteres (ex: 'es', 'pt', 'en').")
|
||||
|
||||
if self.max_pages < 1 or self.max_pages > 10:
|
||||
raise InvalidSearchQueryError("O número máximo de páginas deve estar entre 1 e 10.")
|
||||
|
||||
@property
|
||||
def clean_keyword(self) -> str:
|
||||
return self.keyword.strip()
|
||||
|
||||
@property
|
||||
def clean_language(self) -> str:
|
||||
return self.language.strip().lower()
|
||||
```
|
||||
|
||||
- Campos congelados (`frozen=True`) → imutáveis.
|
||||
- `clean_keyword` / `clean_language` são os valores normalizados usados na busca.
|
||||
- Falha de validação lança `InvalidSearchQueryError` (exceção de domínio).
|
||||
|
||||
---
|
||||
|
||||
## 3. Processamento — passo a passo
|
||||
|
||||
### 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`)
|
||||
|
||||
É o orquestrador completo. Ele faz 3 coisas:
|
||||
|
||||
```python
|
||||
class ExtractNewsUseCase:
|
||||
def __init__(self, extractor: NewsExtractorPort | None = None) -> None:
|
||||
if extractor is None:
|
||||
from googlenews_etl.infrastructure.adapters.google_news_extractor_adapter import (
|
||||
GoogleNewsExtractorAdapter,
|
||||
)
|
||||
self.extractor = GoogleNewsExtractorAdapter()
|
||||
else:
|
||||
self.extractor = extractor
|
||||
|
||||
def execute(self, input_dto: ExtractNewsInputDTO) -> ExtractNewsOutputDTO:
|
||||
# 1. Validação de Domínio (Value Object)
|
||||
search_query = SearchQuery(
|
||||
keyword=input_dto.keyword,
|
||||
language=input_dto.language,
|
||||
max_pages=input_dto.max_pages,
|
||||
)
|
||||
|
||||
# 2. Execução da Extração via Porta (Desacoplada)
|
||||
articles = self.extractor.extract(search_query)
|
||||
|
||||
# 3. Mapeamento de Entidades de Domínio -> Output DTO
|
||||
article_dtos = [
|
||||
NewsArticleDTO(
|
||||
titulo=art.title,
|
||||
subtitulo=art.subtitle,
|
||||
quando_publicado=art.published_at,
|
||||
url=art.url,
|
||||
pagina=art.page,
|
||||
)
|
||||
for art in articles
|
||||
]
|
||||
|
||||
return ExtractNewsOutputDTO(
|
||||
query=search_query.clean_keyword,
|
||||
language=search_query.clean_language,
|
||||
total_paginas=search_query.max_pages,
|
||||
total_itens=len(article_dtos),
|
||||
scraped_at=datetime.now(UTC).isoformat(),
|
||||
items=article_dtos,
|
||||
)
|
||||
```
|
||||
|
||||
Notas importantes:
|
||||
- **Injeção de dependência com fallback:** se não receber um extrator, o use case instancia `GoogleNewsExtractorAdapter()` por padrão (import lazy dentro do `__init__`).
|
||||
- A extração passa pela **porta** `NewsExtractorPort.extract(query)` — desacoplamento da infraestrutura.
|
||||
- A saída é montada com `scraped_at` = timestamp UTC ISO do momento da raspagem.
|
||||
|
||||
### 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`)
|
||||
|
||||
```python
|
||||
class NewsExtractorPort(ABC):
|
||||
"""
|
||||
Porta (Interface) para o serviço de extração de notícias do Google News.
|
||||
Permite desacoplar totalmente o mecanismo de raspagem (HTTP, Playwright, RSS, RabbitMQ, etc.) do domínio.
|
||||
"""
|
||||
|
||||
@abstractmethod
|
||||
def extract(self, query: SearchQuery) -> list[NewsArticle]:
|
||||
"""Extrai as notícias correspondentes aos critérios de busca."""
|
||||
pass
|
||||
```
|
||||
|
||||
### 3.3 O adaptador concreto — o coração do extrator (`googlenews_etl/infrastructure/adapters/google_news_extractor_adapter.py`)
|
||||
|
||||
#### 3.3.1 Inicialização: sessão HTTP com impersonação de browser
|
||||
|
||||
```python
|
||||
class GoogleNewsExtractorAdapter(NewsExtractorPort):
|
||||
DEFAULT_HEADERS = {
|
||||
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
|
||||
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/132.0.0.0 Safari/537.36",
|
||||
}
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
rate_limiter: RateLimiterService | None = None,
|
||||
url_resolver: UrlResolverPort | None = None,
|
||||
impersonate: str = "chrome120",
|
||||
resolve_final_urls: bool = True,
|
||||
) -> None:
|
||||
self.impersonate = impersonate
|
||||
self.rate_limiter = rate_limiter or RateLimiterService(min_delay_seconds=0.5, max_delay_seconds=1.0)
|
||||
self.url_resolver = url_resolver or PlaywrightUrlResolverAdapter()
|
||||
self.resolve_final_urls = resolve_final_urls
|
||||
self.session = requests.Session(impersonate=self.impersonate)
|
||||
self.session.headers.update(self.DEFAULT_HEADERS)
|
||||
```
|
||||
|
||||
- Usa `curl_cffi` com `impersonate="chrome120"` — imita o TLS fingerprint do Chrome para evitar bloqueios.
|
||||
- `RateLimiterService` com delay aleatório 0.5–1.0s entre requisições (configurável).
|
||||
- `resolve_final_urls=True` por padrão → resolve as URLs intermediárias do Google para as URLs reais dos veículos.
|
||||
|
||||
#### 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`)
|
||||
|
||||
```python
|
||||
def _get_hl_gl(self, lang_raw: str) -> tuple[str, str]:
|
||||
"""Mapeia dinamicamente código de idioma e país para os parâmetros hl e gl do Google News."""
|
||||
lang_clean = lang_raw.lower().replace("-", "_")
|
||||
|
||||
locale_map = {
|
||||
"pt": ("pt-BR", "BR"),
|
||||
"pt_br": ("pt-BR", "BR"),
|
||||
"es": ("es-419", "AR"),
|
||||
"es_mx": ("es-419", "MX"),
|
||||
"es_es": ("es", "ES"),
|
||||
"en": ("en-US", "US"),
|
||||
"en_gb": ("en-GB", "GB"),
|
||||
"en_uk": ("en-GB", "GB"),
|
||||
"en_us": ("en-US", "US"),
|
||||
"de": ("de", "DE"),
|
||||
"de_de": ("de", "DE"),
|
||||
"it": ("it", "IT"),
|
||||
"it_it": ("it", "IT"),
|
||||
"fr": ("fr", "FR"),
|
||||
}
|
||||
|
||||
if lang_clean in locale_map:
|
||||
return locale_map[lang_clean]
|
||||
|
||||
parts = lang_clean.split("_")
|
||||
if len(parts) == 2:
|
||||
return (f"{parts[0]}-{parts[1].upper()}", parts[1].upper())
|
||||
|
||||
return (lang_clean, lang_clean.upper())
|
||||
```
|
||||
|
||||
- `hl` = idioma da interface, `gl` = país da região (ex: `pt` → `pt-BR`/`BR`).
|
||||
- Fallback genérico: `xx_yy` → `xx-YY`/`YY`; senão `xx`/`XX`.
|
||||
|
||||
#### 3.3.3 ★ Extração do RSS — o ponto fundamental (`_fetch_rss`)
|
||||
|
||||
**Este é o trecho que extrai de fato as manchetes com título, URL e data de publicação:**
|
||||
|
||||
```python
|
||||
def _fetch_rss(self, query: SearchQuery) -> list[NewsArticle]:
|
||||
encoded_query = urllib.parse.quote_plus(query.clean_keyword)
|
||||
hl, gl = self._get_hl_gl(query.clean_language)
|
||||
rss_url = f"https://news.google.com/rss/search?q={encoded_query}&hl={hl}&gl={gl}&ceid={gl}:{hl}"
|
||||
|
||||
response = self.session.get(rss_url)
|
||||
response.raise_for_status()
|
||||
|
||||
soup = BeautifulSoup(response.content, "xml")
|
||||
rss_items = soup.find_all("item")
|
||||
|
||||
articles: list[NewsArticle] = []
|
||||
max_allowed = query.max_pages * 10
|
||||
|
||||
for idx, item in enumerate(rss_items[:max_allowed]):
|
||||
page_number = (idx // 10) + 1
|
||||
title = item.find("title").text if item.find("title") else ""
|
||||
link = item.find("link").text if item.find("link") else ""
|
||||
pub_date = item.find("pubDate").text if item.find("pubDate") else ""
|
||||
desc_raw = item.find("description").text if item.find("description") else ""
|
||||
|
||||
desc_soup = BeautifulSoup(desc_raw, "html.parser")
|
||||
snippet = desc_soup.get_text(separator=" ", strip=True) if desc_raw else None
|
||||
|
||||
if title and link:
|
||||
articles.append(
|
||||
NewsArticle(
|
||||
title=title,
|
||||
subtitle=snippet if snippet != title else None,
|
||||
published_at=pub_date,
|
||||
url=link,
|
||||
page=page_number,
|
||||
)
|
||||
)
|
||||
|
||||
return articles
|
||||
```
|
||||
|
||||
**Pontos fundamentais deste trecho:**
|
||||
|
||||
| # | Mecanismo | Detalhe |
|
||||
|---|-----------|---------|
|
||||
| 1 | **URL do RSS** | `https://news.google.com/rss/search?q={keyword}&hl={hl}&gl={gl}&ceid={gl}:{hl}` — RSS oficial de busca do Google News |
|
||||
| 2 | **Parsing XML** | `BeautifulSoup(response.content, "xml")` + `find_all("item")` (formato RSS padrão) |
|
||||
| 3 | **Limite** | `max_pages * 10` itens — cada "página" do Google News = 10 itens |
|
||||
| 4 | **Nº da página** | `page_number = (idx // 10) + 1` — agrupa os itens em páginas de 10 |
|
||||
| 5 | **Campos extraídos por item** | `title`, `link`, `pubDate`, `description` (XML do RSS) |
|
||||
| 6 | **Limpeza do resumo** | `description` contém HTML — `BeautifulSoup(desc_raw, "html.parser")` + `get_text(separator=" ", strip=True)` remove tags |
|
||||
| 7 | **Filtro** | item só entra se tiver `title` **e** `link` |
|
||||
| 8 | **Dedupe de subtítulo** | `subtitle = snippet if snippet != title else None` — se o resumo for igual ao título, fica `None` |
|
||||
|
||||
#### 3.3.4 Orquestração com resolução de URLs (`extract`)
|
||||
|
||||
```python
|
||||
def extract(self, query: SearchQuery) -> list[NewsArticle]:
|
||||
# Busca direta das manchetes pelo RSS oficial do Google News
|
||||
all_articles = self._fetch_rss(query)
|
||||
|
||||
# Resolução paralela das URLs finais dos veículos se ativado
|
||||
if self.resolve_final_urls and all_articles:
|
||||
raw_urls = [a.url for a in all_articles]
|
||||
resolved_urls = self.url_resolver.resolve_batch(raw_urls)
|
||||
|
||||
resolved_articles: list[NewsArticle] = []
|
||||
for idx, article in enumerate(all_articles):
|
||||
new_url = resolved_urls[idx] if idx < len(resolved_urls) else article.url
|
||||
resolved_articles.append(dataclasses.replace(article, url=new_url))
|
||||
|
||||
all_articles = resolved_articles
|
||||
|
||||
return all_articles
|
||||
```
|
||||
|
||||
- As URLs do RSS do Google News são intermediárias (`news.google.com/rss/articles/CBMi...`).
|
||||
- `resolve_batch` resolve **em paralelo** (Playwright headless) para as URLs diretas dos veículos.
|
||||
- `dataclasses.replace(article, url=new_url)` preserva os demais campos.
|
||||
|
||||
### 3.4 Resolução de URLs (`googlenews_etl/infrastructure/adapters/playwright_url_resolver_adapter.py`)
|
||||
|
||||
```python
|
||||
class PlaywrightUrlResolverAdapter(UrlResolverPort):
|
||||
def __init__(self, timeout_ms: int = 6000, max_concurrent: int = 5) -> None:
|
||||
self.timeout_ms = timeout_ms
|
||||
self.max_concurrent = max_concurrent
|
||||
|
||||
async def _resolve_single_async(self, context, semaphore, url: str) -> str:
|
||||
if not url or "news.google.com/rss/articles/" not in url:
|
||||
return url
|
||||
|
||||
async with semaphore:
|
||||
page = await context.new_page()
|
||||
target_url = url
|
||||
|
||||
def handle_request(req):
|
||||
nonlocal target_url
|
||||
u = req.url
|
||||
if not any(
|
||||
x in u
|
||||
for x in [
|
||||
"google.", "gstatic.", "googleapis.", "googletagmanager.",
|
||||
"w3.org", "schema.org",
|
||||
]
|
||||
):
|
||||
if not target_url or target_url == url:
|
||||
if u.startswith("http"):
|
||||
target_url = u
|
||||
|
||||
page.on("request", handle_request)
|
||||
try:
|
||||
await page.goto(url, wait_until="commit", timeout=self.timeout_ms)
|
||||
for _ in range(12):
|
||||
await asyncio.sleep(0.25)
|
||||
if "google.com" not in page.url:
|
||||
target_url = page.url
|
||||
break
|
||||
if target_url and target_url != url:
|
||||
break
|
||||
except Exception:
|
||||
pass
|
||||
finally:
|
||||
await page.close()
|
||||
|
||||
return target_url or url
|
||||
```
|
||||
|
||||
- Abre cada URL em uma página headless e **captura o primeiro request não-Google** (o redirecionamento para o veículo).
|
||||
- `Semaphore(max_concurrent=5)` limita concorrência; timeout de 6s por URL.
|
||||
- Intercepta requisições (`page.on("request")`) filtrando domínios de Google/telemetria.
|
||||
|
||||
### 3.5 Rate limiter (`googlenews_etl/domain/services/rate_limiter_service.py`)
|
||||
|
||||
```python
|
||||
class RateLimiterService:
|
||||
def __init__(self, min_delay_seconds: float = 1.0, max_delay_seconds: float = 2.5) -> None:
|
||||
self.min_delay = min_delay_seconds
|
||||
self.max_delay = max_delay_seconds
|
||||
|
||||
def wait(self) -> float:
|
||||
"""Aplica uma pausa aleatória dentro dos limites configurados e retorna o tempo aguardado."""
|
||||
delay = random.uniform(self.min_delay, self.max_delay)
|
||||
time.sleep(delay)
|
||||
return delay
|
||||
```
|
||||
|
||||
- Delay **aleatório** entre requisições (evita padrão detectável/anti-bot).
|
||||
|
||||
---
|
||||
|
||||
## 4. Entidade de domínio da notícia (`googlenews_etl/domain/entities/news_article.py`)
|
||||
|
||||
```python
|
||||
@dataclass(frozen=True)
|
||||
class NewsArticle:
|
||||
"""Entidade do Domínio representando uma notícia extraída do Google News."""
|
||||
|
||||
title: str
|
||||
url: str
|
||||
page: int
|
||||
subtitle: str | None = None
|
||||
published_at: str | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not self.title or not self.title.strip():
|
||||
raise ValueError("O título da notícia não pode ser vazio.")
|
||||
if not self.url or not self.url.strip():
|
||||
raise ValueError("A URL da notícia não pode ser vazia.")
|
||||
```
|
||||
|
||||
| Campo | Tipo | Origem no RSS |
|
||||
|----------------|------------|----------------------------------|
|
||||
| `title` | str | `<title>` do `<item>` |
|
||||
| `url` | str | `<link>` do `<item>` (resolvida depois) |
|
||||
| `page` | int | Calculado: `(idx // 10) + 1` |
|
||||
| `subtitle` | str\|None | `<description>` limpo (ou `None`)|
|
||||
| `published_at` | str\|None | `<pubDate>` do `<item>` |
|
||||
|
||||
**Nota:** `published_at` é mantido como **string crua** do RSS (formato RFC 822, ex: `Thu, 20 Aug 2026 10:00:00 GMT`). Nenhum parse de data acontece no extrator.
|
||||
|
||||
---
|
||||
|
||||
## 5. Saída
|
||||
|
||||
### 5.1 DTOs de saída (`googlenews_etl/application/dtos/extract_news_dto.py`)
|
||||
|
||||
```python
|
||||
class NewsArticleDTO(BaseModel):
|
||||
"""DTO individual para representação de cada artigo."""
|
||||
|
||||
titulo: str
|
||||
subtitulo: str | None = None
|
||||
quando_publicado: str | None = None
|
||||
url: str
|
||||
pagina: int
|
||||
|
||||
|
||||
class ExtractNewsOutputDTO(BaseModel):
|
||||
"""DTO de Saída estruturado com o resultado consolidador do ETL."""
|
||||
|
||||
query: str
|
||||
language: str
|
||||
total_paginas: int
|
||||
total_itens: int
|
||||
scraped_at: str
|
||||
items: list[NewsArticleDTO]
|
||||
```
|
||||
|
||||
### 5.2 Exemplo de saída JSON
|
||||
|
||||
Entrada: `{"keyword": "inteligencia artificial", "language": "pt", "max_pages": 1}`
|
||||
|
||||
```json
|
||||
{
|
||||
"query": "inteligencia artificial",
|
||||
"language": "pt",
|
||||
"total_paginas": 1,
|
||||
"total_itens": 10,
|
||||
"scraped_at": "2026-08-20T14:32:10.482930+00:00",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Empresas aceleram adoção de inteligência artificial no Brasil",
|
||||
"subtitulo": "Levantamento mostra crescimento de 40% no uso de IA generativa...",
|
||||
"quando_publicado": "Thu, 20 Aug 2026 09:12:00 GMT",
|
||||
"url": "https://exemplo.com.br/noticia/123",
|
||||
"pagina": 1
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Resumo do fluxo completo (entrada → saída)
|
||||
|
||||
```
|
||||
1. ExtractNewsInputDTO(keyword, language, max_pages)
|
||||
│
|
||||
2. SearchQuery (valida: keyword não vazia, language ≥ 2 chars, 1 ≤ max_pages ≤ 10)
|
||||
│
|
||||
3. NewsExtractorPort.extract(query) ← porta (abstração)
|
||||
│
|
||||
4. GoogleNewsExtractorAdapter.extract(query)
|
||||
│ a) _get_hl_gl(language) → (hl, gl) ex: "pt" → ("pt-BR", "BR")
|
||||
│ b) _fetch_rss(query):
|
||||
│ URL: https://news.google.com/rss/search?q=...&hl=...&gl=...&ceid=...
|
||||
│ GET com curl_cffi (impersonate chrome120)
|
||||
│ BeautifulSoup XML → itens
|
||||
│ extrai title, link, pubDate, description (HTML limpo)
|
||||
│ limita a max_pages * 10 itens, pagina = (idx // 10) + 1
|
||||
│ c) resolve_batch(urls) via Playwright → URLs finais dos veículos
|
||||
│
|
||||
5. list[NewsArticle] (entidade imutável com validação title/url não vazios)
|
||||
│
|
||||
6. Mapeamento → list[NewsArticleDTO] (titulo, subtitulo, quando_publicado, url, pagina)
|
||||
│
|
||||
7. ExtractNewsOutputDTO (query, language, total_paginas, total_itens, scraped_at, items)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Dependências para rodar (o que NÃO está nos 4 arquivos centrais)
|
||||
|
||||
Para portar a extração a outro repositório, além dos 4 arquivos centrais
|
||||
(`extract_news_use_case.py`, `google_news_extractor_adapter.py`, `news_article.py`,
|
||||
`extract_news_dto.py`), é preciso:
|
||||
|
||||
| Componente | Arquivo | Papel |
|
||||
|------------|---------|-------|
|
||||
| `NewsExtractorPort` | `domain/ports/news_extractor_port.py` | Interface que o adaptador implementa |
|
||||
| `SearchQuery` | `domain/entities/search_query.py` | Value object validado usado na busca |
|
||||
| `UrlResolverPort` | `domain/ports/url_resolver_port.py` | Interface do resolvedor de URLs |
|
||||
| `PlaywrightUrlResolverAdapter` | `infrastructure/adapters/playwright_url_resolver_adapter.py` | Resolução das URLs intermediárias |
|
||||
| `RateLimiterService` | `domain/services/rate_limiter_service.py` | Throttling entre requisições |
|
||||
| `InvalidSearchQueryError` | `domain/exceptions/domain_exceptions.py` | Exceção de validação do `SearchQuery` |
|
||||
|
||||
### Dependências de bibliotecas (requirements)
|
||||
|
||||
- `curl_cffi` — sessão HTTP com impersonação de TLS do Chrome
|
||||
- `beautifulsoup4` — parsing XML/HTML do feed
|
||||
- `pydantic` — DTOs de entrada/saída
|
||||
- `playwright` — resolução de URLs (navegador headless)
|
||||
- Python ≥ 3.11 (uso de `X | None` em anotações de tipo)
|
||||
|
||||
### Comportamentos observáveis (honestidade)
|
||||
|
||||
- `published_at` é string crua do RSS (RFC 822), **não** parseada.
|
||||
- `subtitle` vem do `description` do item com HTML removido; vira `None` se igual ao título.
|
||||
- O item só é mantido se tiver `title` e `link` não vazios.
|
||||
- `resolve_final_urls` pode ser desligado (`False`) para pular a etapa Playwright.
|
||||
- O rate limiter existe, mas no `_fetch_rss` atual o delay é aplicado apenas por construção
|
||||
do `RateLimiterService` (o método `wait()` está disponível para chamadas sequenciais).
|
||||
|
||||
---
|
||||
|
||||
## 8. Como chamar (exemplo de uso)
|
||||
|
||||
```python
|
||||
from googlenews_etl.application.dtos.extract_news_dto import ExtractNewsInputDTO
|
||||
from googlenews_etl.application.use_cases.extract_news_use_case import ExtractNewsUseCase
|
||||
|
||||
dto_in = ExtractNewsInputDTO(keyword="inteligencia artificial", language="pt", max_pages=1)
|
||||
resultado = ExtractNewsUseCase().execute(dto_in)
|
||||
|
||||
print(resultado.total_itens) # ex: 10
|
||||
print(resultado.items[0].titulo) # título da primeira manchete
|
||||
print(resultado.items[0].url) # URL final resolvida
|
||||
print(resultado.items[0].quando_publicado) # data crua do RSS
|
||||
```
|
||||
Reference in New Issue
Block a user