feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
@@ -4,7 +4,7 @@
|
||||
"2": "SpecKit Utilities",
|
||||
"3": "Graphify Commands",
|
||||
"4": "speckit-analyze/SKILL.md",
|
||||
"5": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"5": "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier",
|
||||
"6": "Feature Specification Template",
|
||||
"7": "Graphify Rules",
|
||||
"8": "Implementation Planning",
|
||||
@@ -39,15 +39,15 @@
|
||||
"37": "Plan Setup",
|
||||
"38": "Task Setup",
|
||||
"39": "Graphify Workflows",
|
||||
"40": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"40": "main",
|
||||
"41": "1. Technical Decisions & Tradeoffs",
|
||||
"42": "1. Input Schemas",
|
||||
"43": "2. Basic CLI Usage Examples",
|
||||
"44": "2. Standard Streams & Exit Codes",
|
||||
"45": "ECPSnapshot",
|
||||
"46": "classifier.py",
|
||||
"46": "Tasks: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
"47": "detect_language",
|
||||
"48": "main",
|
||||
"48": "test_models.py",
|
||||
"49": "content_northvolt_de.md",
|
||||
"50": "content_presal_pt.md",
|
||||
"51": "content_tangential_es.md",
|
||||
@@ -78,5 +78,24 @@
|
||||
"76": "pt/not_related.md",
|
||||
"77": "pt/tangential.md",
|
||||
"78": "tests/__init__.py",
|
||||
"79": "text-nlp-classifier"
|
||||
"79": "text-nlp-classifier",
|
||||
"80": "get_hl_gl_ceid",
|
||||
"81": "Extrator de Notícias do Google News — Guia Completo de Funcionamento",
|
||||
"82": "extract_google_news.py",
|
||||
"83": "ExtractionResult",
|
||||
"84": "test_extract_google_news.py",
|
||||
"85": "Implementation Tasks: Google News Headlines Extractor",
|
||||
"86": "Feature Specification: Google News Headlines Extractor",
|
||||
"87": "2. Cenários Práticos de Uso",
|
||||
"88": "Implementation Plan: Google News Headlines Extractor",
|
||||
"89": "scripts/__init__.py",
|
||||
"90": "SearchQuery",
|
||||
"91": "1. Technical Decisions & Tradeoffs",
|
||||
"92": "General Readiness Checklist: Google News Headlines Extractor",
|
||||
"93": "1. Entidades de Domínio & DTOs",
|
||||
"94": "Specification Quality Checklist: Google News Headlines Extractor",
|
||||
"95": "CLI Contract: Google News Headlines Extractor",
|
||||
"96": "readiness.md",
|
||||
"98": "build_parser",
|
||||
"99": "sample_rss_xml"
|
||||
}
|
||||
|
||||
@@ -1,16 +1,16 @@
|
||||
# Graph Report - TextNLPClassifierApp (2026-08-20)
|
||||
|
||||
## Corpus Check
|
||||
- 132 files · ~53,988 words
|
||||
- 147 files · ~65,826 words
|
||||
- Verdict: corpus is large enough that graph structure adds value.
|
||||
|
||||
## Summary
|
||||
- 579 nodes · 674 edges · 80 communities (43 shown, 37 thin omitted)
|
||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 28 edges (avg confidence: 0.95)
|
||||
- 775 nodes · 931 edges · 99 communities (61 shown, 38 thin omitted)
|
||||
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 36 edges (avg confidence: 0.94)
|
||||
- Token cost: 0 input · 0 output
|
||||
|
||||
## Graph Freshness
|
||||
- Built from commit: `d371b81a`
|
||||
- Built from commit: `67cc40f9`
|
||||
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
|
||||
- Run `graphify update .` after code changes (no API cost).
|
||||
|
||||
@@ -20,7 +20,7 @@
|
||||
- SpecKit Utilities
|
||||
- Graphify Commands
|
||||
- speckit-analyze/SKILL.md
|
||||
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier
|
||||
- Feature Specification Template
|
||||
- Graphify Rules
|
||||
- Implementation Planning
|
||||
@@ -51,15 +51,15 @@
|
||||
- Media Transcription
|
||||
- Extraction Specification
|
||||
- Graphify Workflows
|
||||
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- main
|
||||
- 1. Technical Decisions & Tradeoffs
|
||||
- 1. Input Schemas
|
||||
- 2. Basic CLI Usage Examples
|
||||
- 2. Standard Streams & Exit Codes
|
||||
- ECPSnapshot
|
||||
- classifier.py
|
||||
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
|
||||
- detect_language
|
||||
- main
|
||||
- test_models.py
|
||||
- content_northvolt_de.md
|
||||
- content_presal_pt.md
|
||||
- content_tangential_es.md
|
||||
@@ -91,35 +91,53 @@
|
||||
- pt/tangential.md
|
||||
- tests/__init__.py
|
||||
- text-nlp-classifier
|
||||
- get_hl_gl_ceid
|
||||
- Extrator de Notícias do Google News — Guia Completo de Funcionamento
|
||||
- extract_google_news.py
|
||||
- ExtractionResult
|
||||
- test_extract_google_news.py
|
||||
- Implementation Tasks: Google News Headlines Extractor
|
||||
- Feature Specification: Google News Headlines Extractor
|
||||
- 2. Cenários Práticos de Uso
|
||||
- Implementation Plan: Google News Headlines Extractor
|
||||
- scripts/__init__.py
|
||||
- SearchQuery
|
||||
- 1. Technical Decisions & Tradeoffs
|
||||
- General Readiness Checklist: Google News Headlines Extractor
|
||||
- 1. Entidades de Domínio & DTOs
|
||||
- Specification Quality Checklist: Google News Headlines Extractor
|
||||
- CLI Contract: Google News Headlines Extractor
|
||||
- build_parser
|
||||
- sample_rss_xml
|
||||
|
||||
## God Nodes (most connected - your core abstractions)
|
||||
1. `ECPSnapshot` - 27 edges
|
||||
2. `InherenceClassifier` - 21 edges
|
||||
3. `ClassificationResult` - 17 edges
|
||||
4. `LocalEmbeddingsAdapter` - 14 edges
|
||||
5. `LLMFallbackAdapter` - 14 edges
|
||||
6. `detect_language()` - 14 edges
|
||||
7. `DecisionCategory` - 14 edges
|
||||
8. `main()` - 13 edges
|
||||
1. `ECPSnapshot` - 31 edges
|
||||
2. `InherenceClassifier` - 25 edges
|
||||
3. `DecisionCategory` - 18 edges
|
||||
4. `ClassificationResult` - 17 edges
|
||||
5. `main()` - 14 edges
|
||||
6. `LocalEmbeddingsAdapter` - 14 edges
|
||||
7. `LLMFallbackAdapter` - 14 edges
|
||||
8. `detect_language()` - 14 edges
|
||||
9. `Tasks: [FEATURE NAME]` - 13 edges
|
||||
10. `BaseNLPAdapter` - 12 edges
|
||||
10. `SearchQuery` - 12 edges
|
||||
|
||||
## Surprising Connections (you probably didn't know these)
|
||||
- `main()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `main()` --uses--> `ErrorCode` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `test_extract_google_news_orchestration_mocked()` --uses--> `ExtractionResult` [INFERRED]
|
||||
tests/test_extract_google_news.py → scripts/extract_google_news.py
|
||||
- `test_classification_result_serialization()` --uses--> `DecisionCategory` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `petrobras_ecp()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_classifier.py → src/models.py
|
||||
- `emit_error()` --uses--> `ErrorCode` [INFERRED]
|
||||
classify.py → src/models.py
|
||||
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
- `test_ecp_snapshot_missing_required()` --uses--> `ECPSnapshot` [INFERRED]
|
||||
tests/test_models.py → src/models.py
|
||||
|
||||
## Import Cycles
|
||||
- None detected.
|
||||
|
||||
## Communities (80 total, 37 thin omitted)
|
||||
## Communities (99 total, 38 thin omitted)
|
||||
|
||||
### Community 0 - "Task Planning"
|
||||
Cohesion: 0.07
|
||||
@@ -141,9 +159,9 @@ Nodes (24): For /graphify add and --watch, For /graphify query, For the commit h
|
||||
Cohesion: 0.08
|
||||
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
|
||||
|
||||
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.06
|
||||
Nodes (33): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+25 more)
|
||||
### Community 5 - "POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier"
|
||||
Cohesion: 0.05
|
||||
Nodes (34): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+26 more)
|
||||
|
||||
### Community 6 - "Feature Specification Template"
|
||||
Cohesion: 0.15
|
||||
@@ -241,9 +259,9 @@ Nodes (3): For --cluster-only, For --update (incremental re-extraction), graphif
|
||||
Cohesion: 0.50
|
||||
Nodes (3): Boundaries, Output, Scan
|
||||
|
||||
### Community 40 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.14
|
||||
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
|
||||
### Community 40 - "main"
|
||||
Cohesion: 0.19
|
||||
Nodes (14): CaptureFixture, Path, main(), Ponto de entrada do script CLI., Valida execução padrão do CLI com saída JSON no stdout., Valida gravação em arquivo com criação automática de diretórios pais., Valida tratamento de erro e código de saída 1 para parâmetro vazio., Valida tratamento de erro e código de saída 2 para falhas de rede. (+6 more)
|
||||
|
||||
### Community 41 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.22
|
||||
@@ -262,40 +280,108 @@ Cohesion: 0.29
|
||||
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
|
||||
|
||||
### Community 45 - "ECPSnapshot"
|
||||
Cohesion: 0.07
|
||||
Nodes (26): ABC, Any, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+18 more)
|
||||
Cohesion: 0.06
|
||||
Nodes (55): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+47 more)
|
||||
|
||||
### Community 46 - "classifier.py"
|
||||
Cohesion: 0.09
|
||||
Nodes (39): emit_error(), Enum, count_phrase_occurrences(), InherenceClassifier, match_phrase_in_text(), Core deterministic classification engine (Tier 1 core)., Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text. (+31 more)
|
||||
### Community 46 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
|
||||
Cohesion: 0.15
|
||||
Nodes (13): Dependencies & Execution Order, Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Strategy, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Core Tier 1 Deterministic Classification & CLI (Priority: P1) [MVP] (+5 more)
|
||||
|
||||
### Community 47 - "detect_language"
|
||||
Cohesion: 0.19
|
||||
Nodes (16): detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization., Normalize text by converting to lowercase and stripping combining diacritical…, Tokenize text into lowercase alphanumeric words., Detect the ISO-639-1 language code of text among supported languages (pt, en,…, Unit tests for language detection and text normalization. (+8 more)
|
||||
Cohesion: 0.14
|
||||
Nodes (21): count_phrase_occurrences(), match_phrase_in_text(), Check if a normalized phrase appears in normalized text with word boundary…, Count occurrences of a phrase in text., Classify inherence of content against an ECP snapshot., detect_language(), extract_words(), normalize_text() (+13 more)
|
||||
|
||||
### Community 48 - "main"
|
||||
Cohesion: 0.31
|
||||
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
|
||||
### Community 48 - "test_models.py"
|
||||
Cohesion: 0.09
|
||||
Nodes (29): emit_error(), main(), parse_args(), Namespace, ClassificationError, ErrorCode, Any, extract_evidence_snippets() (+21 more)
|
||||
|
||||
### Community 80 - "get_hl_gl_ceid"
|
||||
Cohesion: 0.25
|
||||
Nodes (8): get_hl_gl_ceid(), Mapeia idioma e locale para os parâmetros hl, gl e ceid do Google News., Valida o mapeamento padrão de idiomas para pares (hl, gl, ceid)., Valida a sobrescrita geográfica quando o argumento locale é especificado., Valida fallback dinâmico para idiomas regionais não listados explicitamente., test_get_hl_gl_ceid_default_mappings(), test_get_hl_gl_ceid_dynamic_fallback(), test_get_hl_gl_ceid_with_custom_locale()
|
||||
|
||||
### Community 81 - "Extrator de Notícias do Google News — Guia Completo de Funcionamento"
|
||||
Cohesion: 0.08
|
||||
Nodes (24): 1. Visão geral (arquitetura), 2.1 DTO de entrada (`googlenews_etl/application/dtos/extract_news_dto.py`), 2.2 Value Object de validação (`googlenews_etl/domain/entities/search_query.py`), 2. Entrada, 3.1 O caso de uso (`googlenews_etl/application/use_cases/extract_news_use_case.py`), 3.2 A porta (`googlenews_etl/domain/ports/news_extractor_port.py`), 3.3.1 Inicialização: sessão HTTP com impersonação de browser, 3.3.2 Mapeamento idioma → parâmetros `hl`/`gl` (`_get_hl_gl`) (+16 more)
|
||||
|
||||
### Community 82 - "extract_google_news.py"
|
||||
Cohesion: 0.20
|
||||
Nodes (14): extract_google_news(), _fetch_rss_content(), NewsArticle, _normalize_text_for_comparison(), parse_google_news_rss(), Remove pontuação e espaços extras para comparação de redundância., Parseia o XML do RSS do Google News e extrai os itens estruturados., Resolve em paralelo as URLs intermediárias do Google News para os links finais… (+6 more)
|
||||
|
||||
### Community 83 - "ExtractionResult"
|
||||
Cohesion: 0.29
|
||||
Nodes (5): ExtractionResult, Any, Resultado consolidado da extração., Valida E2E o fluxo completo de busca, parsing e resolução de URLs reais ao vivo., test_e2e_extract_google_news_live_pipeline()
|
||||
|
||||
### Community 84 - "test_extract_google_news.py"
|
||||
Cohesion: 0.21
|
||||
Nodes (11): Resolve a URL intermediária do Google News para a URL real do veículo., resolve_article_url(), Testes unitários e de integração para o Extrator de Manchetes do Google News.…, Valida o parsing do feed RSS, higienização de tags HTML e deduplicação., Valida fallback gracioso de URL quando não é link do Google News ou em erro., Valida resolução bem-sucedida de URL do Google News para o portal destino., Valida E2E que o decodificador resolve uma URL real do Google News para o…, test_e2e_resolve_real_google_news_url() (+3 more)
|
||||
|
||||
### Community 85 - "Implementation Tasks: Google News Headlines Extractor"
|
||||
Cohesion: 0.14
|
||||
Nodes (14): Implementation for User Story 1, Implementation for User Story 2, Implementation for User Story 3, Implementation Tasks: Google News Headlines Extractor, Phase 1: Setup (Shared Infrastructure), Phase 2: Foundational (Blocking Prerequisites), Phase 3: User Story 1 - Extração Básica de Notícias por Assunto e Idioma (Priority: P1) 🌟 MVP, Phase 4: User Story 2 - Filtragem Regional e Edição Geográfica (Priority: P2) (+6 more)
|
||||
|
||||
### Community 86 - "Feature Specification: Google News Headlines Extractor"
|
||||
Cohesion: 0.18
|
||||
Nodes (11): Clarifications, Edge Cases, Feature Specification: Google News Headlines Extractor, Functional Requirements, Requirements *(mandatory)*, Session 2026-08-20, Success Criteria *(mandatory)*, User Scenarios & Testing *(mandatory)* (+3 more)
|
||||
|
||||
### Community 87 - "2. Cenários Práticos de Uso"
|
||||
Cohesion: 0.20
|
||||
Nodes (9): 1. Pré-requisitos e Instalação, 2. Cenários Práticos de Uso, 3. Validação dos Testes Automatizados e Linter, Cenário 1: River Plate — Argentina (Espanhol / 2 Páginas / Salvar em Arquivo), Cenário 2: Cruzeiro — Brasil (Português / Formatado no Terminal), Cenário 3: Fórmula 1 — Reino Unido (Inglês), Cenário 4: Integração em Pipeline com `jq` (Modo Silencioso), Cenário 5: Extração Rápida com Links Brutos (Sem Resolução de URLs) (+1 more)
|
||||
|
||||
### Community 88 - "Implementation Plan: Google News Headlines Extractor"
|
||||
Cohesion: 0.29
|
||||
Nodes (7): Architecture & Pipeline, Documentation (this feature), Implementation Plan: Google News Headlines Extractor, Project Structure, Source Code, Summary, Technical Context
|
||||
|
||||
### Community 90 - "SearchQuery"
|
||||
Cohesion: 0.20
|
||||
Nodes (6): Value Object com parâmetros de busca validados., SearchQuery, Valida a consolidação do ExtractionResult a partir da busca mockada com URLs…, Valida as regras de negócio e limites de SearchQuery., test_extract_google_news_orchestration_mocked(), test_search_query_validation()
|
||||
|
||||
### Community 91 - "1. Technical Decisions & Tradeoffs"
|
||||
Cohesion: 0.25
|
||||
Nodes (7): 1. Technical Decisions & Tradeoffs, Decision 1: Motor de Requisição e Scraping com `foxcape` em Modo Headless, Decision 2: Endpoint RSS do Google News vs. Scraping de DOM, Decision 3: Mapeamento de Idioma e Locale (`hl`, `gl`, `ceid`), Decision 4: Resolução de URLs do Google News via `googlenewsdecoder`, Decision 5: Logging em Tempo Real no `stderr` e Segregação de Streams, Research: Google News Headlines Extractor
|
||||
|
||||
### Community 92 - "General Readiness Checklist: Google News Headlines Extractor"
|
||||
Cohesion: 0.29
|
||||
Nodes (7): CLI Interface & Parameter Contracts, Data Sanitization & Article Extraction, Error Handling & Edge Cases, General Readiness Checklist: Google News Headlines Extractor, Non-Functional & Operational Readiness, Notes, Scraping Engine & Feed Mapping
|
||||
|
||||
### Community 93 - "1. Entidades de Domínio & DTOs"
|
||||
Cohesion: 0.29
|
||||
Nodes (6): 1.1 SearchQuery (Parâmetros da Busca), 1.2 NewsArticle (Item de Notícia), 1.3 ExtractionResult (Saída Estruturada Consolidada), 1. Entidades de Domínio & DTOs, 2. Esquema JSON de Saída, Data Model: Google News Headlines Extractor
|
||||
|
||||
### Community 94 - "Specification Quality Checklist: Google News Headlines Extractor"
|
||||
Cohesion: 0.33
|
||||
Nodes (5): Content Quality, Feature Readiness, Notes, Requirement Completeness, Specification Quality Checklist: Google News Headlines Extractor
|
||||
|
||||
### Community 95 - "CLI Contract: Google News Headlines Extractor"
|
||||
Cohesion: 0.33
|
||||
Nodes (6): 1. Comando e Argumentos, 2. Códigos de Saída (Exit Codes), 3. Protocolo de Streams (Stdout / Stderr), Argumentos de Linha de Comando, CLI Contract: Google News Headlines Extractor, Sintaxe
|
||||
|
||||
### Community 98 - "build_parser"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): ArgumentParser, build_parser(), Cria e configura o parser de argumentos CLI.
|
||||
|
||||
### Community 99 - "sample_rss_xml"
|
||||
Cohesion: 0.67
|
||||
Nodes (3): fixture, Fixture que fornece o conteúdo do XML de exemplo para testes offline., sample_rss_xml()
|
||||
|
||||
## Knowledge Gaps
|
||||
- **291 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+286 more)
|
||||
- **361 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+356 more)
|
||||
These have ≤1 connection - possible missing edges or undocumented components.
|
||||
- **37 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
- **38 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
|
||||
|
||||
## Suggested Questions
|
||||
_Questions this graph is uniquely positioned to answer:_
|
||||
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `classifier.py`?**
|
||||
_High betweenness centrality (0.011) - this node is a cross-community bridge._
|
||||
- **Why does `detect_language()` connect `detect_language` to `classifier.py`?**
|
||||
_High betweenness centrality (0.007) - this node is a cross-community bridge._
|
||||
- **Why does `InherenceClassifier` connect `classifier.py` to `main`, `ECPSnapshot`?**
|
||||
_High betweenness centrality (0.006) - this node is a cross-community bridge._
|
||||
- **Why does `main()` connect `test_models.py` to `main`, `ECPSnapshot`?**
|
||||
_High betweenness centrality (0.031) - this node is a cross-community bridge._
|
||||
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `test_models.py`, `detect_language`?**
|
||||
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||
- **Why does `main()` connect `main` to `extract_google_news.py`, `test_extract_google_news.py`, `SearchQuery`, `build_parser`?**
|
||||
_High betweenness centrality (0.021) - this node is a cross-community bridge._
|
||||
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
|
||||
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
|
||||
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?**
|
||||
_`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
|
||||
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
|
||||
- **Are the 3 inferred relationships involving `LocalEmbeddingsAdapter` (e.g. with `ClassificationResult` and `ECPSnapshot`) actually correct?**
|
||||
_`LocalEmbeddingsAdapter` has 3 INFERRED edges - model-reasoned connections that need verification._
|
||||
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._
|
||||
+6111
-912
File diff suppressed because it is too large
Load Diff
@@ -294,15 +294,15 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"classify.py": {
|
||||
"mtime": 1787195933.7872643,
|
||||
"seen": 1787196463.095295,
|
||||
"ast_hash": "5fa90ecfb87bdf15ee99165a3be7c63b",
|
||||
"mtime": 1787197110.3753626,
|
||||
"seen": 1787197159.9828603,
|
||||
"ast_hash": "d09a35a5e25d42f6ecc537d5b67bef8e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"pyproject.toml": {
|
||||
"mtime": 1787195827.3305018,
|
||||
"seen": 1787196463.0953014,
|
||||
"ast_hash": "f2503e96d08f5a4e41b13352e81709c6",
|
||||
"mtime": 1787234982.768605,
|
||||
"seen": 1787235020.6850271,
|
||||
"ast_hash": "26f5f979658067bf447f375f044a8f21",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/__init__.py": {
|
||||
@@ -336,9 +336,9 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/classifier.py": {
|
||||
"mtime": 1787196413.3774571,
|
||||
"seen": 1787196463.095319,
|
||||
"ast_hash": "d7bab620ae29badf0d3779fab44552bc",
|
||||
"mtime": 1787197085.0250447,
|
||||
"seen": 1787197159.985785,
|
||||
"ast_hash": "a2e70f968109d213fdab0ae81ac819ff",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"src/language.py": {
|
||||
@@ -420,9 +420,9 @@
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"requirements.txt": {
|
||||
"mtime": 1787195817.4782183,
|
||||
"seen": 1787196463.1020837,
|
||||
"ast_hash": "5612073a1e7034d7c25765379baa5da4",
|
||||
"mtime": 1787236111.876247,
|
||||
"seen": 1787236262.0105975,
|
||||
"ast_hash": "9fb7e3f8ecb04ff30d650c772f509557",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/fixtures/benchmark_24/de/contextual.md": {
|
||||
@@ -568,5 +568,89 @@
|
||||
"seen": 1787196463.1031141,
|
||||
"ast_hash": "e05ab20a5190cfb5b9d41da64d5273cf",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_adversarial.py": {
|
||||
"mtime": 1787197141.2340336,
|
||||
"seen": 1787197159.9881344,
|
||||
"ast_hash": "5e8daf517ee32c261617d96264ef0473",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"scripts/extract_google_news.py": {
|
||||
"mtime": 1787236453.5231855,
|
||||
"seen": 1787236509.1147037,
|
||||
"ast_hash": "4216de490a87378d21dab707a3166679",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"tests/test_extract_google_news.py": {
|
||||
"mtime": 1787236301.513448,
|
||||
"seen": 1787236353.5582018,
|
||||
"ast_hash": "1fc9a108a8abfecf1d37c8642a885e9e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/googlenews_extractor_guia_completo.md": {
|
||||
"mtime": 1787231605.0830677,
|
||||
"seen": 1787234772.8112028,
|
||||
"ast_hash": "59db8e652000ba6087a00289a0ce0959",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/checklists/readiness.md": {
|
||||
"mtime": 1787234305.932687,
|
||||
"seen": 1787234772.8123577,
|
||||
"ast_hash": "a1200a6959d41e256a73403dd5c4e1f6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/checklists/requirements.md": {
|
||||
"mtime": 1787232386.6763985,
|
||||
"seen": 1787234772.812359,
|
||||
"ast_hash": "39cd86af76c9f74baa2984418870b677",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/contracts/cli_contract.md": {
|
||||
"mtime": 1787236744.3139226,
|
||||
"seen": 1787236817.4979818,
|
||||
"ast_hash": "76a61ad31df989244edc5e739081bf7b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/data-model.md": {
|
||||
"mtime": 1787234040.8134317,
|
||||
"seen": 1787234772.8123617,
|
||||
"ast_hash": "e8380c98d2f6418f60ac8a511ecdb637",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/plan.md": {
|
||||
"mtime": 1787236725.9825225,
|
||||
"seen": 1787236817.4983613,
|
||||
"ast_hash": "a787c974414ae6c173d4c39e76113c31",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/quickstart.md": {
|
||||
"mtime": 1787236764.0206513,
|
||||
"seen": 1787236817.4983652,
|
||||
"ast_hash": "ada22d34a0fa8341f88136cf42f055a9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/research.md": {
|
||||
"mtime": 1787236785.0466182,
|
||||
"seen": 1787236817.498368,
|
||||
"ast_hash": "60837e6c7463c4413911ceca0087374c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/spec.md": {
|
||||
"mtime": 1787236709.0783155,
|
||||
"seen": 1787236817.4983711,
|
||||
"ast_hash": "dfd963f7ca351c93e09f74af9b11d19e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"specs/002-google-news-extractor/tasks.md": {
|
||||
"mtime": 1787236802.4808047,
|
||||
"seen": 1787236817.4983742,
|
||||
"ast_hash": "427f24026deaa7c9aca95d99e7587ae6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"scripts/__init__.py": {
|
||||
"mtime": 1787234973.5782337,
|
||||
"seen": 1787235020.6850297,
|
||||
"ast_hash": "627a6c953b250083ebe76ee74d605bb5",
|
||||
"semantic_hash": ""
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user