feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution

- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping
- Integrate foxcape in headless mode as primary stealth anti-bot engine
- Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor
- Support language and regional locale mapping (-l, --lang, --locale)
- Implement real-time progress logging in stderr and --silent flag
- Add unit, integration, and live E2E tests in tests/test_extract_google_news.py
- Add full SpecKit documentation (specs/002-google-news-extractor/)
- Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
2026-08-20 11:50:16 -03:00
parent 67cc40f91a
commit 6e3d57619b
59 changed files with 16118 additions and 2160 deletions
+149 -38
View File
@@ -1,16 +1,16 @@
# Graph Report - TextNLPClassifierApp (2026-08-20)
## Corpus Check
- 133 files · ~54,759 words
- 148 files · ~67,357 words
- Verdict: corpus is large enough that graph structure adds value.
## Summary
- 595 nodes · 704 edges · 80 communities (43 shown, 37 thin omitted)
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 31 edges (avg confidence: 0.95)
- 802 nodes · 957 edges · 104 communities (66 shown, 38 thin omitted)
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
- Token cost: 0 input · 0 output
## Graph Freshness
- Built from commit: `d371b81a`
- Built from commit: `67cc40f9`
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost).
@@ -51,15 +51,15 @@
- Media Transcription
- Extraction Specification
- Graphify Workflows
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- main
- 1. Technical Decisions & Tradeoffs
- 1. Input Schemas
- 2. Basic CLI Usage Examples
- 2. Standard Streams & Exit Codes
- ECPSnapshot
- test_models.py
- ClassificationResult
- InherenceClassifier
- detect_language
- main
- test_models.py
- content_northvolt_de.md
- content_presal_pt.md
- content_tangential_es.md
@@ -91,35 +91,58 @@
- pt/tangential.md
- tests/__init__.py
- text-nlp-classifier
- get_hl_gl_ceid
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
- extract_google_news.py
- ExtractionResult
- test_extract_google_news.py
- Implementation Tasks: Google News Headlines Extractor
- Feature Specification: Google News Headlines Extractor
- 2. Cenários Práticos de Uso
- Implementation Plan: Google News Headlines Extractor
- scripts/__init__.py
- SearchQuery
- 1. Technical Decisions & Tradeoffs
- General Readiness Checklist: Google News Headlines Extractor
- 1. Entidades de Domínio & DTOs
- Specification Quality Checklist: Google News Headlines Extractor
- CLI Contract: Google News Headlines Extractor
- 🧠 TextNLPClassifierApp
- build_parser
- sample_rss_xml
- classifier.py
- ECPSnapshot
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- main
## God Nodes (most connected - your core abstractions)
1. `ECPSnapshot` - 31 edges
2. `InherenceClassifier` - 25 edges
3. `DecisionCategory` - 18 edges
4. `ClassificationResult` - 17 edges
5. `LocalEmbeddingsAdapter` - 14 edges
6. `LLMFallbackAdapter` - 14 edges
7. `detect_language()` - 14 edges
8. `main()` - 13 edges
5. `main()` - 14 edges
6. `LocalEmbeddingsAdapter` - 14 edges
7. `LLMFallbackAdapter` - 14 edges
8. `detect_language()` - 14 edges
9. `Tasks: [FEATURE NAME]` - 13 edges
10. `BaseNLPAdapter` - 12 edges
10. `SearchQuery` - 12 edges
## Surprising Connections (you probably didn't know these)
- `emit_error()` --uses--> `ErrorCode` [INFERRED]
classify.py → src/models.py
- `main()` --uses--> `ECPSnapshot` [INFERRED]
classify.py → src/models.py
- `main()` --uses--> `ErrorCode` [INFERRED]
classify.py → src/models.py
- `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
tests/test_models.py → src/models.py
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
tests/test_extract_google_news.py → scripts/extract_google_news.py
- `classifier()` --uses--> `InherenceClassifier` [INFERRED]
tests/test_benchmark_24.py → src/classifier.py
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
tests/test_models.py → src/models.py
## Import Cycles
- None detected.
## Communities (80 total, 37 thin omitted)
## Communities (104 total, 38 thin omitted)
### Community 0 - "Task Planning"
Cohesion: 0.07
@@ -241,9 +264,9 @@ Nodes (3): For --cluster-only, For --update (incremental re-extraction), graphif
Cohesion: 0.50
Nodes (3): Boundaries, Output, Scan
### Community 40 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
### Community 40 - "main"
Cohesion: 0.19
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
### Community 41 - "1. Technical Decisions & Tradeoffs"
Cohesion: 0.22
@@ -261,36 +284,124 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
Cohesion: 0.29
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
### Community 45 - "ECPSnapshot"
Cohesion: 0.05
Nodes (63): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+55 more)
### Community 45 - "ClassificationResult"
Cohesion: 0.10
Nodes (18): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., LocalEmbeddingsAdapter (+10 more)
### Community 46 - "test_models.py"
### Community 46 - "InherenceClassifier"
Cohesion: 0.12
Nodes (17): Any, emit_error(), ClassificationError, extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences. (+9 more)
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more)
### Community 47 - "detect_language"
Cohesion: 0.19
Nodes (16): detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization., Normalize text by converting to lowercase and stripping combining diacritical…, Tokenize text into lowercase alphanumeric words., Detect the ISO-639-1 language code of text among supported languages (pt, en,…, Unit tests for language detection and text normalization. (+8 more)
Cohesion: 0.14
Nodes (20): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization. (+12 more)
### Community 48 - "main"
### Community 48 - "test_models.py"
Cohesion: 0.21
Nodes (12): Classify inherence of content against an ECP snapshot., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+4 more)
### Community 80 - "get_hl_gl_ceid"
Cohesion: 0.25
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
Cohesion: 0.08
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
### Community 82 - "extract_google_news.py"
Cohesion: 0.20
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
### Community 83 - "ExtractionResult"
Cohesion: 0.29
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
### Community 84 - "test_extract_google_news.py"
Cohesion: 0.21
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
Cohesion: 0.14
Nodes (14): Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Tasks: Google News Headlines Extractor, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Extração Básica de Notícias por Assunto e Idioma (Priority: P1) 🌟 MVP, Phase 4: User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2) (+6 more)
### Community 86 - "Feature Specification: Google News Headlines Extractor"
Cohesion: 0.18
Nodes (11): Clarifications, Edge Cases, Feature Specification: Google News Headlines Extractor, Functional Requirements, Requirements *(mandatory)*, Session 2026-08-20, Success Criteria *(mandatory)*, User Scenarios & Testing *(mandatory)* (+3 more)
### Community 87 - "2. Cenários Práticos de Uso"
Cohesion: 0.20
Nodes (9): 1. Pré-requisitos e Instalação, 2. Cenários Práticos de Uso, 3. Validação dos Testes Automatizados e Linter, Cenário 1: River Plate — Argentina (Espanhol / 2 Páginas / Salvar em Arquivo), Cenário 2: Cruzeiro — Brasil (Português / Formatado no Terminal), Cenário 3: Fórmula 1 — Reino Unido (Inglês), Cenário 4: Integração em Pipeline com `jq` (Modo Silencioso), Cenário 5: Extração Rápida com Links Brutos (Sem Resolução de URLs) (+1 more)
### Community 88 - "Implementation Plan: Google News Headlines Extractor"
Cohesion: 0.29
Nodes (7): Architecture & Pipeline, Documentation (this feature), Implementation Plan: Google News Headlines Extractor, Project Structure, Source Code, Summary, Technical Context
### Community 90 - "SearchQuery"
Cohesion: 0.20
Nodes (6): Value Object com parâmetros de busca validados., SearchQuery, Valida a consolidação do ExtractionResult a partir da busca mockada com URLs…, Valida as regras de negócio e limites de SearchQuery., test_extract_google_news_orchestration_mocked(), test_search_query_validation()
### Community 91 - "1. Technical Decisions & Tradeoffs"
Cohesion: 0.25
Nodes (7): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Requisição e Scraping com `foxcape` em Modo Headless, Decision 2: Endpoint RSS do Google News vs. Scraping de DOM, Decision 3: Mapeamento de Idioma e Locale (`hl`, `gl`, `ceid`), Decision 4: Resolução de URLs do Google News via `googlenewsdecoder`, Decision 5: Logging em Tempo Real no `stderr` e Segregação de Streams, Research: Google News Headlines Extractor
### Community 92 - "General Readiness Checklist: Google News Headlines Extractor"
Cohesion: 0.29
Nodes (7): CLI Interface & Parameter Contracts, Data Sanitization & Article Extraction, Error Handling & Edge Cases, General Readiness Checklist: Google News Headlines Extractor, Non-Functional & Operational Readiness, Notes, Scraping Engine & Feed Mapping
### Community 93 - "1. Entidades de Domínio & DTOs"
Cohesion: 0.29
Nodes (6): 1.1 SearchQuery (Parâmetros da Busca), 1.2 NewsArticle (Item de Notícia), 1.3 ExtractionResult (Saída Estruturada Consolidada), 1. Entidades de Domínio & DTOs, 2. Esquema JSON de Saída, Data Model: Google News Headlines Extractor
### Community 94 - "Specification Quality Checklist: Google News Headlines Extractor"
Cohesion: 0.33
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Google News Headlines Extractor
### Community 95 - "CLI Contract: Google News Headlines Extractor"
Cohesion: 0.33
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
### Community 97 - "🧠 TextNLPClassifierApp"
Cohesion: 0.07
Nodes (26): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências, 3. Baixar Binários do Navegador Stealth (Camoufox), 3. Fórmula 1 (Inglaterra / Inglês) (+18 more)
### Community 98 - "build_parser"
Cohesion: 0.67
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
### Community 99 - "sample_rss_xml"
Cohesion: 0.67
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
### Community 100 - "classifier.py"
Cohesion: 0.23
Nodes (9): emit_error(), Enum, Core deterministic classification engine (Tier 1 core)., ClassificationError, ErrorCode, MatchedGraphEntity, Data models and validation schemas for Multilingual NLP Entity Inherence…, str (+1 more)
### Community 101 - "ECPSnapshot"
Cohesion: 0.20
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more)
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
### Community 103 - "main"
Cohesion: 0.31
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
## Knowledge Gaps
- **291 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+286 more)
- **381 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+376 more)
These have ≤1 connection - possible missing edges or undocumented components.
- **37 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions
_Questions this graph is uniquely positioned to answer:_
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `test_models.py`?**
_High betweenness centrality (0.015) - this node is a cross-community bridge._
- **Why does `InherenceClassifier` connect `ECPSnapshot` to `main`?**
_High betweenness centrality (0.008) - this node is a cross-community bridge._
- **Why does `detect_language()` connect `detect_language` to `ECPSnapshot`?**
_High betweenness centrality (0.007) - this node is a cross-community bridge._
- **Why does `main()` connect `main` to `main`, `classifier.py`, `ECPSnapshot`, `InherenceClassifier`?**
_High betweenness centrality (0.029) - this node is a cross-community bridge._
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `classifier.py`, `main`, `ClassificationResult`, `InherenceClassifier`, `test_models.py`?**
_High betweenness centrality (0.020) - this node is a cross-community bridge._
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
_High betweenness centrality (0.020) - this node is a cross-community bridge._
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**