- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
5.1 KiB
General Readiness Checklist: Google News Headlines Extractor
Purpose: Validate the completeness, clarity, consistency, and testability of requirements for the standalone Google News CLI extractor before implementation. Created: 2026-08-20 Validated: 2026-08-20 (Audited against guide, spec, research, plan, data-model, and cli_contract) Feature: spec.md | plan.md | cli_contract.md | data-model.md | research.md
Note: This custom checklist is generated by the /speckit-checklist command based on feature context and requirements.
Review Ownership: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item [x] only when the reviewer determines the requirements-quality criterion is satisfied.
Marker Semantics: [x] means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
CLI Interface & Parameter Contracts
- CHK001 - Are all required CLI arguments (e.g.,
--query/--keyword) and their aliases explicitly specified? [Completeness, Spec §FR-003, Contract §1] - CHK002 - Are default values clearly defined for optional parameters (
--lang,--locale,--max-pages)? [Clarity, Spec §FR-003, Contract §1] - CHK003 - Is the allowed range for pagination (
1to10pages) explicitly bounded and unambiguous? [Clarity, Spec §FR-006, DataModel §1.1] - CHK004 - Are exit codes defined for each distinct execution outcome (success, validation failure, scraping/network error)? [Completeness, Contract §2]
- CHK005 - Are stream separation requirements (clean JSON on
stdout, diagnostics/errors onstderr) strictly documented? [Consistency, Spec §FR-009, Contract §3]
Scraping Engine & Feed Mapping
- CHK006 - Is the integration role of the
foxcapelibrary and its anti-bot evasion responsibilities clearly documented? [Completeness, Spec §FR-002, Plan §Summary] - CHK007 - Is the RSS search URL template and its encoding rules (
q,hl,gl,ceid) precisely defined? [Clarity, Research §Decision 2, Research §Decision 3] - CHK008 - Is the locale inference fallback rule (when
--localeis omitted) explicitly specified across standard language codes? [Consistency, Spec §FR-005, Research §Decision 3] - CHK009 - Is the override behavior when a custom
--localeis passed alongside a different language documented? [Completeness, Spec §User Story 2, Research §Decision 3]
Data Sanitization & Article Extraction
- CHK010 - Are the target extraction fields (
titulo,subtitulo,quando_publicado,url,pagina) mapped to specific RSS XML tags? [Completeness, Spec §FR-007, DataModel §1.2] - CHK011 - Is the HTML tag stripping requirement for description/snippets specified with unambiguous criteria? [Clarity, Spec §FR-008, SC-003]
- CHK012 - Is the deduplication rule for subtitles identical to titles clearly documented with expected null/empty behavior? [Consistency, Spec §Edge Cases, DataModel §1.2]
- CHK013 - Are filtering rules specified for discarding incomplete items missing title or link? [Coverage, Spec §Edge Cases, SC-002]
- CHK014 - Is the consolidated JSON schema defined with all mandatory metadata fields (
query,language,locale,total_itens,scraped_at)? [Completeness, Spec §FR-009, DataModel §2]
Error Handling & Edge Cases
- CHK015 - Are validation requirements specified for empty or whitespace-only search keywords? [Coverage, Spec §Edge Cases, Spec §FR-010]
- CHK016 - Is the behavior for queries yielding zero matching news articles specified without raising unhandled errors? [Coverage, Spec §User Story 1, SC-004]
- CHK017 - Are network interruption or upstream HTTP block error handling requirements documented? [Coverage, Spec §Edge Cases, Contract §2]
- CHK018 - Are character encoding and URL parameter escaping requirements documented for special characters and accents? [Clarity, Spec §Edge Cases]
Non-Functional & Operational Readiness
- CHK019 - Is the performance target (< 2.0s for standard single-page extractions) quantified with measurable criteria? [Measurability, Spec §SC-001, Plan §Technical Context]
- CHK020 - Are JSON output compatibility requirements with terminal streaming tools (e.g.,
jq, shell pipes) defined? [Completeness, Spec §SC-005, Contract §3] - CHK021 - Is the execution isolation assumption (pure RSS feed extraction without mandatory Playwright browser runtime) clearly stated? [Traceability, Spec §Assumptions, Research §Decision 1]
Notes
- Mark items
[x]only after review confirms the requirement-quality criterion is satisfied - Leave items unchecked when they still require clarification, correction, or reviewer evaluation
/speckit-implementreads checklist checkbox state as a gate and must not modify markerschecklists/requirements.mdhas a separate built-in lifecycle maintained by/speckit-specifyand/speckit-clarify- Add comments or findings inline
- Link to relevant resources or documentation
- Items are numbered sequentially for easy reference