Files
andreferraro 6e3d57619b feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping
- Integrate foxcape in headless mode as primary stealth anti-bot engine
- Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor
- Support language and regional locale mapping (-l, --lang, --locale)
- Implement real-time progress logging in stderr and --silent flag
- Add unit, integration, and live E2E tests in tests/test_extract_google_news.py
- Add full SpecKit documentation (specs/002-google-news-extractor/)
- Create comprehensive README.md covering both NLP Classifier and Google News Extractor
2026-08-20 11:50:16 -03:00

5.1 KiB

General Readiness Checklist: Google News Headlines Extractor

Purpose: Validate the completeness, clarity, consistency, and testability of requirements for the standalone Google News CLI extractor before implementation. Created: 2026-08-20 Validated: 2026-08-20 (Audited against guide, spec, research, plan, data-model, and cli_contract) Feature: spec.md | plan.md | cli_contract.md | data-model.md | research.md

Note: This custom checklist is generated by the /speckit-checklist command based on feature context and requirements. Review Ownership: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item [x] only when the reviewer determines the requirements-quality criterion is satisfied. Marker Semantics: [x] means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.

CLI Interface & Parameter Contracts

  • CHK001 - Are all required CLI arguments (e.g., --query / --keyword) and their aliases explicitly specified? [Completeness, Spec §FR-003, Contract §1]
  • CHK002 - Are default values clearly defined for optional parameters (--lang, --locale, --max-pages)? [Clarity, Spec §FR-003, Contract §1]
  • CHK003 - Is the allowed range for pagination (1 to 10 pages) explicitly bounded and unambiguous? [Clarity, Spec §FR-006, DataModel §1.1]
  • CHK004 - Are exit codes defined for each distinct execution outcome (success, validation failure, scraping/network error)? [Completeness, Contract §2]
  • CHK005 - Are stream separation requirements (clean JSON on stdout, diagnostics/errors on stderr) strictly documented? [Consistency, Spec §FR-009, Contract §3]

Scraping Engine & Feed Mapping

  • CHK006 - Is the integration role of the foxcape library and its anti-bot evasion responsibilities clearly documented? [Completeness, Spec §FR-002, Plan §Summary]
  • CHK007 - Is the RSS search URL template and its encoding rules (q, hl, gl, ceid) precisely defined? [Clarity, Research §Decision 2, Research §Decision 3]
  • CHK008 - Is the locale inference fallback rule (when --locale is omitted) explicitly specified across standard language codes? [Consistency, Spec §FR-005, Research §Decision 3]
  • CHK009 - Is the override behavior when a custom --locale is passed alongside a different language documented? [Completeness, Spec §User Story 2, Research §Decision 3]

Data Sanitization & Article Extraction

  • CHK010 - Are the target extraction fields (titulo, subtitulo, quando_publicado, url, pagina) mapped to specific RSS XML tags? [Completeness, Spec §FR-007, DataModel §1.2]
  • CHK011 - Is the HTML tag stripping requirement for description/snippets specified with unambiguous criteria? [Clarity, Spec §FR-008, SC-003]
  • CHK012 - Is the deduplication rule for subtitles identical to titles clearly documented with expected null/empty behavior? [Consistency, Spec §Edge Cases, DataModel §1.2]
  • CHK013 - Are filtering rules specified for discarding incomplete items missing title or link? [Coverage, Spec §Edge Cases, SC-002]
  • CHK014 - Is the consolidated JSON schema defined with all mandatory metadata fields (query, language, locale, total_itens, scraped_at)? [Completeness, Spec §FR-009, DataModel §2]

Error Handling & Edge Cases

  • CHK015 - Are validation requirements specified for empty or whitespace-only search keywords? [Coverage, Spec §Edge Cases, Spec §FR-010]
  • CHK016 - Is the behavior for queries yielding zero matching news articles specified without raising unhandled errors? [Coverage, Spec §User Story 1, SC-004]
  • CHK017 - Are network interruption or upstream HTTP block error handling requirements documented? [Coverage, Spec §Edge Cases, Contract §2]
  • CHK018 - Are character encoding and URL parameter escaping requirements documented for special characters and accents? [Clarity, Spec §Edge Cases]

Non-Functional & Operational Readiness

  • CHK019 - Is the performance target (< 2.0s for standard single-page extractions) quantified with measurable criteria? [Measurability, Spec §SC-001, Plan §Technical Context]
  • CHK020 - Are JSON output compatibility requirements with terminal streaming tools (e.g., jq, shell pipes) defined? [Completeness, Spec §SC-005, Contract §3]
  • CHK021 - Is the execution isolation assumption (pure RSS feed extraction without mandatory Playwright browser runtime) clearly stated? [Traceability, Spec §Assumptions, Research §Decision 1]

Notes

  • Mark items [x] only after review confirms the requirement-quality criterion is satisfied
  • Leave items unchecked when they still require clarification, correction, or reviewer evaluation
  • /speckit-implement reads checklist checkbox state as a gate and must not modify markers
  • checklists/requirements.md has a separate built-in lifecycle maintained by /speckit-specify and /speckit-clarify
  • Add comments or findings inline
  • Link to relevant resources or documentation
  • Items are numbered sequentially for easy reference