feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping - Integrate foxcape in headless mode as primary stealth anti-bot engine - Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor - Support language and regional locale mapping (-l, --lang, --locale) - Implement real-time progress logging in stderr and --silent flag - Add unit, integration, and live E2E tests in tests/test_extract_google_news.py - Add full SpecKit documentation (specs/002-google-news-extractor/) - Create comprehensive README.md covering both NLP Classifier and Google News Extractor
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
# General Readiness Checklist: Google News Headlines Extractor
|
||||
|
||||
**Purpose**: Validate the completeness, clarity, consistency, and testability of requirements for the standalone Google News CLI extractor before implementation.
|
||||
**Created**: 2026-08-20
|
||||
**Validated**: 2026-08-20 (Audited against guide, spec, research, plan, data-model, and cli_contract)
|
||||
**Feature**: [spec.md](../spec.md) | [plan.md](../plan.md) | [cli_contract.md](../contracts/cli_contract.md) | [data-model.md](../data-model.md) | [research.md](../research.md)
|
||||
|
||||
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
|
||||
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
|
||||
|
||||
## CLI Interface & Parameter Contracts
|
||||
|
||||
- [x] CHK001 - Are all required CLI arguments (e.g., `--query` / `--keyword`) and their aliases explicitly specified? [Completeness, Spec §FR-003, Contract §1]
|
||||
- [x] CHK002 - Are default values clearly defined for optional parameters (`--lang`, `--locale`, `--max-pages`)? [Clarity, Spec §FR-003, Contract §1]
|
||||
- [x] CHK003 - Is the allowed range for pagination (`1` to `10` pages) explicitly bounded and unambiguous? [Clarity, Spec §FR-006, DataModel §1.1]
|
||||
- [x] CHK004 - Are exit codes defined for each distinct execution outcome (success, validation failure, scraping/network error)? [Completeness, Contract §2]
|
||||
- [x] CHK005 - Are stream separation requirements (clean JSON on `stdout`, diagnostics/errors on `stderr`) strictly documented? [Consistency, Spec §FR-009, Contract §3]
|
||||
|
||||
## Scraping Engine & Feed Mapping
|
||||
|
||||
- [x] CHK006 - Is the integration role of the `foxcape` library and its anti-bot evasion responsibilities clearly documented? [Completeness, Spec §FR-002, Plan §Summary]
|
||||
- [x] CHK007 - Is the RSS search URL template and its encoding rules (`q`, `hl`, `gl`, `ceid`) precisely defined? [Clarity, Research §Decision 2, Research §Decision 3]
|
||||
- [x] CHK008 - Is the locale inference fallback rule (when `--locale` is omitted) explicitly specified across standard language codes? [Consistency, Spec §FR-005, Research §Decision 3]
|
||||
- [x] CHK009 - Is the override behavior when a custom `--locale` is passed alongside a different language documented? [Completeness, Spec §User Story 2, Research §Decision 3]
|
||||
|
||||
## Data Sanitization & Article Extraction
|
||||
|
||||
- [x] CHK010 - Are the target extraction fields (`titulo`, `subtitulo`, `quando_publicado`, `url`, `pagina`) mapped to specific RSS XML tags? [Completeness, Spec §FR-007, DataModel §1.2]
|
||||
- [x] CHK011 - Is the HTML tag stripping requirement for description/snippets specified with unambiguous criteria? [Clarity, Spec §FR-008, SC-003]
|
||||
- [x] CHK012 - Is the deduplication rule for subtitles identical to titles clearly documented with expected null/empty behavior? [Consistency, Spec §Edge Cases, DataModel §1.2]
|
||||
- [x] CHK013 - Are filtering rules specified for discarding incomplete items missing title or link? [Coverage, Spec §Edge Cases, SC-002]
|
||||
- [x] CHK014 - Is the consolidated JSON schema defined with all mandatory metadata fields (`query`, `language`, `locale`, `total_itens`, `scraped_at`)? [Completeness, Spec §FR-009, DataModel §2]
|
||||
|
||||
## Error Handling & Edge Cases
|
||||
|
||||
- [x] CHK015 - Are validation requirements specified for empty or whitespace-only search keywords? [Coverage, Spec §Edge Cases, Spec §FR-010]
|
||||
- [x] CHK016 - Is the behavior for queries yielding zero matching news articles specified without raising unhandled errors? [Coverage, Spec §User Story 1, SC-004]
|
||||
- [x] CHK017 - Are network interruption or upstream HTTP block error handling requirements documented? [Coverage, Spec §Edge Cases, Contract §2]
|
||||
- [x] CHK018 - Are character encoding and URL parameter escaping requirements documented for special characters and accents? [Clarity, Spec §Edge Cases]
|
||||
|
||||
## Non-Functional & Operational Readiness
|
||||
|
||||
- [x] CHK019 - Is the performance target (< 2.0s for standard single-page extractions) quantified with measurable criteria? [Measurability, Spec §SC-001, Plan §Technical Context]
|
||||
- [x] CHK020 - Are JSON output compatibility requirements with terminal streaming tools (e.g., `jq`, shell pipes) defined? [Completeness, Spec §SC-005, Contract §3]
|
||||
- [x] CHK021 - Is the execution isolation assumption (pure RSS feed extraction without mandatory Playwright browser runtime) clearly stated? [Traceability, Spec §Assumptions, Research §Decision 1]
|
||||
|
||||
## Notes
|
||||
|
||||
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied
|
||||
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation
|
||||
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers
|
||||
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`
|
||||
- Add comments or findings inline
|
||||
- Link to relevant resources or documentation
|
||||
- Items are numbered sequentially for easy reference
|
||||
Reference in New Issue
Block a user