feat(extractor): implement multi-engine article content extractor

- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability)
- Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing)
- Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts
- Passed ruff linting/formatting and mypy type checking cleanly
This commit is contained in:
2026-08-20 19:22:20 -03:00
parent 6e3d57619b
commit 6a45368cb0
85 changed files with 18345 additions and 3897 deletions
+124 -49
View File
@@ -1,16 +1,16 @@
# Graph Report - TextNLPClassifierApp (2026-08-20)
## Corpus Check
- 147 files · ~65,826 words
- 161 files · ~78,476 words
- Verdict: corpus is large enough that graph structure adds value.
## Summary
- 775 nodes · 931 edges · 99 communities (61 shown, 38 thin omitted)
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
- 1000 nodes · 1226 edges · 115 communities (77 shown, 38 thin omitted)
- Extraction: 97% EXTRACTED · 3% INFERRED · 0% AMBIGUOUS · INFERRED: 39 edges (avg confidence: 0.95)
- Token cost: 0 input · 0 output
## Graph Freshness
- Built from commit: `67cc40f9`
- Built from commit: `6e3d5761`
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost).
@@ -20,7 +20,7 @@
- SpecKit Utilities
- Graphify Commands
- speckit-analyze/SKILL.md
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
- Feature Specification Template
- Graphify Rules
- Implementation Planning
@@ -56,10 +56,10 @@
- 1. Input Schemas
- 2. Basic CLI Usage Examples
- 2. Standard Streams & Exit Codes
- ECPSnapshot
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
- LocalEmbeddingsAdapter
- InherenceClassifier
- detect_language
- test_models.py
- classifier.py
- content_northvolt_de.md
- content_presal_pt.md
- content_tangential_es.md
@@ -91,7 +91,7 @@
- pt/tangential.md
- tests/__init__.py
- text-nlp-classifier
- get_hl_gl_ceid
- test_extract_article_contents.py
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
- extract_google_news.py
- ExtractionResult
@@ -107,37 +107,52 @@
- 1. Entidades de Domínio & DTOs
- Specification Quality Checklist: Google News Headlines Extractor
- CLI Contract: Google News Headlines Extractor
- build_parser
- 🧠 TextNLPClassifierApp
- Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
- sample_rss_xml
- models.py
- ECPSnapshot
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- 4. Requisitos Funcionais (FR)
- Tasks: Article Content Multi-Engine Extractor
- ClassificationResult
- Implementation Plan: Article Content Multi-Engine Extractor
- 2. Cenários de Validação
- 1. Technical Decisions & Tradeoffs
- Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
- Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
- CLI Contract: Article Content Multi-Engine Extractor
- JSON Schema Contract: Article Content Multi-Engine Extractor
## God Nodes (most connected - your core abstractions)
1. `ECPSnapshot` - 31 edges
2. `InherenceClassifier` - 25 edges
3. `DecisionCategory` - 18 edges
4. `ClassificationResult` - 17 edges
5. `main()` - 14 edges
5. `process_batch()` - 15 edges
6. `LocalEmbeddingsAdapter` - 14 edges
7. `LLMFallbackAdapter` - 14 edges
8. `detect_language()` - 14 edges
9. `Tasks: [FEATURE NAME]` - 13 edges
10. `SearchQuery` - 12 edges
9. `main()` - 13 edges
10. `ArticleCrawler` - 13 edges
## Surprising Connections (you probably didn't know these)
- `main()` --uses--> `ECPSnapshot` [INFERRED]
classify.py → src/models.py
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
tests/test_extract_google_news.py → scripts/extract_google_news.py
- `classifier()` --uses--> `InherenceClassifier` [INFERRED]
tests/test_benchmark_24.py → src/classifier.py
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
tests/test_models.py → src/models.py
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
tests/test_models.py → src/models.py
- `test_ecp_snapshot_missing_required()` --uses--> `ECPSnapshot` [INFERRED]
tests/test_models.py → src/models.py
- `petrobras_ecp()` --uses--> `ECPSnapshot` [INFERRED]
tests/test_classifier.py → src/models.py
## Import Cycles
- None detected.
## Communities (99 total, 38 thin omitted)
## Communities (115 total, 38 thin omitted)
### Community 0 - "Task Planning"
Cohesion: 0.07
@@ -159,9 +174,9 @@ Nodes (24): For /graphify add and --watch, For /graphify query, For the commit h
Cohesion: 0.08
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
### Community 5 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
Cohesion: 0.05
Nodes (34): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+26 more)
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.15
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
### Community 6 - "Feature Specification Template"
Cohesion: 0.15
@@ -260,8 +275,8 @@ Cohesion: 0.50
Nodes (3): Boundaries, Output, Scan
### Community 40 - "main"
Cohesion: 0.19
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
Cohesion: 0.14
Nodes (17): ArgumentParser, CaptureFixture, build_parser(), main(), Cria e configura o parser de argumentos CLI., Ponto de entrada do script CLI., Path, Valida execução padrão do CLI com saída JSON no stdout. (+9 more)
### Community 41 - "1. Technical Decisions & Tradeoffs"
Cohesion: 0.22
@@ -279,41 +294,41 @@ Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2
Cohesion: 0.29
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
### Community 45 - "ECPSnapshot"
Cohesion: 0.06
Nodes (55): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+47 more)
### Community 45 - "LocalEmbeddingsAdapter"
Cohesion: 0.14
Nodes (8): LocalEmbeddingsAdapter, Optional adapter for local multilingual semantic vector embeddings., LLMFallbackAdapter, Optional adapter for LLM fallback boundary disambiguation., Unit tests for optional adapter interfaces (Tier 2 / Tier 3)., test_classifier_with_adapter_flags(), test_embeddings_adapter_interface(), test_llm_adapter_interface()
### Community 46 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.15
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
### Community 46 - "InherenceClassifier"
Cohesion: 0.12
Nodes (27): InherenceClassifier, Tier 1 Deterministic NLP Entity Inherence Classifier., DecisionCategory, RelatedEntity, Adversarial and robustness test suite for Multilingual NLP Entity Inherence…, Run CLI via subprocess with empty content and verify error code and exit code., Content about city/state governance of São Paulo against ECP for São Paulo FC., Run CLI via subprocess with missing target_name and verify error payload. (+19 more)
### Community 47 - "detect_language"
Cohesion: 0.14
Nodes (21): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+13 more)
### Community 48 - "test_models.py"
Cohesion: 0.09
Nodes (29): emit_error(), main(), parse_args(), Namespace, ClassificationError, ErrorCode, Any, extract_evidence_snippets() (+21 more)
### Community 48 - "classifier.py"
Cohesion: 0.24
Nodes (11): Core deterministic classification engine (Tier 1 core)., extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences., Extract relevant sentence excerpts from Markdown text that contain any of the…, strip_markdown() (+3 more)
### Community 80 - "get_hl_gl_ceid"
Cohesion: 0.25
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
### Community 80 - "test_extract_article_contents.py"
Cohesion: 0.06
Nodes (56): ArticleCrawler, extract_all_engines(), ExtractedArticle, ExtractionBatchReport, InputArticle, load_search_json(), log_info(), main() (+48 more)
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
Cohesion: 0.08
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
### Community 82 - "extract_google_news.py"
Cohesion: 0.20
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
Cohesion: 0.15
Nodes (18): extract_google_news(), _fetch_rss_content(), get_hl_gl_ceid(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Remove pontuação e espaços extras para comparação de redundância. (+10 more)
### Community 83 - "ExtractionResult"
Cohesion: 0.29
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
### Community 84 - "test_extract_google_news.py"
Cohesion: 0.21
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
Cohesion: 0.15
Nodes (15): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado. (+7 more)
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
Cohesion: 0.14
@@ -355,28 +370,88 @@ Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness,
Cohesion: 0.33
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
### Community 98 - "build_parser"
Cohesion: 0.67
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
### Community 97 - "🧠 TextNLPClassifierApp"
Cohesion: 0.06
Nodes (33): 1. 🧠 Classificador de Conteúdo e Inerência (NLP / LLM / ECP), 1. Clonar o Repositório e Criar Ambiente Virtual, 1. Extração Completa Automática, 1. River Plate (Argentina / Espanhol / 2 Páginas / Salvar em Arquivo), 2. Amostragem Rápida (Limit 2 Notícias), 2. Cruzeiro (Brasil / Português / Formatado no Terminal), 2. 📰 Extrator de Manchetes do Google News, 2. Instalar Dependências (+25 more)
### Community 98 - "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor"
Cohesion: 0.05
Nodes (34): 1. Requirement Completeness, 2. Requirement Clarity & Non-Ambiguity, 3. Requirement Consistency & Data Contracts, 4. Scenario & Edge Case Coverage, 5. Non-Functional & Operational Readiness, Extraction Pipeline Checklist: Article Content Multi-Engine Extractor, Notes, Content Quality (+26 more)
### Community 99 - "sample_rss_xml"
Cohesion: 0.67
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
### Community 100 - "models.py"
Cohesion: 0.15
Nodes (17): emit_error(), main(), parse_args(), Namespace, Enum, ClassificationError, ErrorCode, MatchedGraphEntity (+9 more)
### Community 101 - "ECPSnapshot"
Cohesion: 0.22
Nodes (10): parametrize, ECPSnapshot, Any, classifier(), fixture, Controlled 24-case benchmark suite for Multilingual NLP Entity Inherence…, test_benchmark_case(), test_ecp_snapshot_defaults() (+2 more)
### Community 102 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
### Community 103 - "4. Requisitos Funcionais (FR)"
Cohesion: 0.10
Nodes (20): 1.1 Objetivo do Produto, 1. Visão Geral e Contexto, 2. Personas e Casos de Uso, 3. Arquitetura e Fluxo do Sistema, 4. Requisitos Funcionais (FR), 5. Requisitos Não Funcionais (NFR), 6.1 Esquema do JSON de Entrada (`Input`), 6.2 Esquema do JSON Consolidado de Saída (`Output`) (+12 more)
### Community 104 - "Tasks: Article Content Multi-Engine Extractor"
Cohesion: 0.11
Nodes (18): Dependencies & Execution Order, Entrega Incremental, Implementation Strategy, Implementação da User Story 1, Implementação da User Story 2, Implementação da User Story 3, MVP First (User Story 1 Only), Oportunidades de Execução Paralela (+10 more)
### Community 105 - "ClassificationResult"
Cohesion: 0.16
Nodes (11): ABC, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms., Optionally refine an ambiguous classification result., Optional local vector embeddings adapter (Tier 2). Disabled by default.… (+3 more)
### Community 106 - "Implementation Plan: Article Content Multi-Engine Extractor"
Cohesion: 0.17
Nodes (11): Constitution Check, Documentation (this feature), Implementation Phases, Implementation Plan: Article Content Multi-Engine Extractor, Phase 0: Outline & Research *(Completed)*, Phase 1: Design & Contracts *(Completed)*, Phase 2: Tasks & Execution *(Next: `/speckit-tasks`)*, Project Structure (+3 more)
### Community 107 - "2. Cenários de Validação"
Cohesion: 0.22
Nodes (8): 1. Pré-requisitos, 2. Cenários de Validação, 3. Validação Automatizada de Testes, Cenário 1: Extração com Amostragem Rápida (Limit 2), Cenário 2: Caminho Customizado de Saída, Cenário 3: Modo Silencioso (`--silent`), Cenário 4: Resiliência contra URLs Inválidas, Quickstart & Validation Guide: Article Content Multi-Engine Extractor
### Community 108 - "1. Technical Decisions & Tradeoffs"
Cohesion: 0.22
Nodes (8): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Navegação e Renderização com `foxcape` em Sessão Única, Decision 2: Orquestração Tripla de Extração de Conteúdo (NLP & Web Scraping), Decision 3: Resiliência e Isolamento de Falhas por Camada, Decision 4: Herança Inteligente de Idioma para NLP, Decision 5: Gerenciamento de Memória e Descarte do Raw HTML, Decision 6: Segregação de Streams e Feedback Visual em `stderr`, Research: Article Content Multi-Engine Extractor
### Community 109 - "Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.25
Nodes (8): Complexity Tracking, Constitution Check, Documentation (this feature), Implementation Plan: Multilingual NLP Entity Inherence Classifier (POC), Project Structure, Source Code (repository root), Summary, Technical Context
### Community 110 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
Cohesion: 0.29
Nodes (7): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
### Community 111 - "Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier"
Cohesion: 0.33
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Multilingual NLP Entity Inherence Classifier
### Community 112 - "CLI Contract: Article Content Multi-Engine Extractor"
Cohesion: 0.33
Nodes (5): 1. Comando de Execução, 2. Argumentos e Flags, 3. Códigos de Saída (Exit Codes), 4. Comportamento de Streams (I/O), CLI Contract: Article Content Multi-Engine Extractor
### Community 114 - "JSON Schema Contract: Article Content Multi-Engine Extractor"
Cohesion: 0.50
Nodes (3): 1. Schema de Entrada (Input JSON), 2. Schema de Saída (Output JSON), JSON Schema Contract: Article Content Multi-Engine Extractor
## Knowledge Gaps
- **361 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+356 more)
- **465 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+460 more)
These have ≤1 connection - possible missing edges or undocumented components.
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions
_Questions this graph is uniquely positioned to answer:_
- **Why does `main()` connect `test_models.py` to `main`, `ECPSnapshot`?**
_High betweenness centrality (0.031) - this node is a cross-community bridge._
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `test_models.py`, `detect_language`?**
_High betweenness centrality (0.021) - this node is a cross-community bridge._
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
_High betweenness centrality (0.021) - this node is a cross-community bridge._
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `models.py`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `InherenceClassifier`, `detect_language`, `classifier.py`?**
_High betweenness centrality (0.005) - this node is a cross-community bridge._
- **Why does `InherenceClassifier` connect `InherenceClassifier` to `models.py`, `ECPSnapshot`, `ClassificationResult`, `LocalEmbeddingsAdapter`, `detect_language`, `classifier.py`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?**
_High betweenness centrality (0.003) - this node is a cross-community bridge._
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**