Files
TextNLPClassifierApp/specs/002-google-news-extractor/checklists/readiness.md
T
andreferraro 6e3d57619b feat(extractor): add Google News headlines extractor with Foxcape headless and URL resolution
- Add standalone CLI script scripts/extract_google_news.py for Google News RSS scraping
- Integrate foxcape in headless mode as primary stealth anti-bot engine
- Implement parallel article URL resolution using googlenewsdecoder and ThreadPoolExecutor
- Support language and regional locale mapping (-l, --lang, --locale)
- Implement real-time progress logging in stderr and --silent flag
- Add unit, integration, and live E2E tests in tests/test_extract_google_news.py
- Add full SpecKit documentation (specs/002-google-news-extractor/)
- Create comprehensive README.md covering both NLP Classifier and Google News Extractor
2026-08-20 11:50:16 -03:00

57 lines
5.1 KiB
Markdown

# General Readiness Checklist: Google News Headlines Extractor
**Purpose**: Validate the completeness, clarity, consistency, and testability of requirements for the standalone Google News CLI extractor before implementation.
**Created**: 2026-08-20
**Validated**: 2026-08-20 (Audited against guide, spec, research, plan, data-model, and cli_contract)
**Feature**: [spec.md](../spec.md) | [plan.md](../plan.md) | [cli_contract.md](../contracts/cli_contract.md) | [data-model.md](../data-model.md) | [research.md](../research.md)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
## CLI Interface & Parameter Contracts
- [x] CHK001 - Are all required CLI arguments (e.g., `--query` / `--keyword`) and their aliases explicitly specified? [Completeness, Spec §FR-003, Contract §1]
- [x] CHK002 - Are default values clearly defined for optional parameters (`--lang`, `--locale`, `--max-pages`)? [Clarity, Spec §FR-003, Contract §1]
- [x] CHK003 - Is the allowed range for pagination (`1` to `10` pages) explicitly bounded and unambiguous? [Clarity, Spec §FR-006, DataModel §1.1]
- [x] CHK004 - Are exit codes defined for each distinct execution outcome (success, validation failure, scraping/network error)? [Completeness, Contract §2]
- [x] CHK005 - Are stream separation requirements (clean JSON on `stdout`, diagnostics/errors on `stderr`) strictly documented? [Consistency, Spec §FR-009, Contract §3]
## Scraping Engine & Feed Mapping
- [x] CHK006 - Is the integration role of the `foxcape` library and its anti-bot evasion responsibilities clearly documented? [Completeness, Spec §FR-002, Plan §Summary]
- [x] CHK007 - Is the RSS search URL template and its encoding rules (`q`, `hl`, `gl`, `ceid`) precisely defined? [Clarity, Research §Decision 2, Research §Decision 3]
- [x] CHK008 - Is the locale inference fallback rule (when `--locale` is omitted) explicitly specified across standard language codes? [Consistency, Spec §FR-005, Research §Decision 3]
- [x] CHK009 - Is the override behavior when a custom `--locale` is passed alongside a different language documented? [Completeness, Spec §User Story 2, Research §Decision 3]
## Data Sanitization & Article Extraction
- [x] CHK010 - Are the target extraction fields (`titulo`, `subtitulo`, `quando_publicado`, `url`, `pagina`) mapped to specific RSS XML tags? [Completeness, Spec §FR-007, DataModel §1.2]
- [x] CHK011 - Is the HTML tag stripping requirement for description/snippets specified with unambiguous criteria? [Clarity, Spec §FR-008, SC-003]
- [x] CHK012 - Is the deduplication rule for subtitles identical to titles clearly documented with expected null/empty behavior? [Consistency, Spec §Edge Cases, DataModel §1.2]
- [x] CHK013 - Are filtering rules specified for discarding incomplete items missing title or link? [Coverage, Spec §Edge Cases, SC-002]
- [x] CHK014 - Is the consolidated JSON schema defined with all mandatory metadata fields (`query`, `language`, `locale`, `total_itens`, `scraped_at`)? [Completeness, Spec §FR-009, DataModel §2]
## Error Handling & Edge Cases
- [x] CHK015 - Are validation requirements specified for empty or whitespace-only search keywords? [Coverage, Spec §Edge Cases, Spec §FR-010]
- [x] CHK016 - Is the behavior for queries yielding zero matching news articles specified without raising unhandled errors? [Coverage, Spec §User Story 1, SC-004]
- [x] CHK017 - Are network interruption or upstream HTTP block error handling requirements documented? [Coverage, Spec §Edge Cases, Contract §2]
- [x] CHK018 - Are character encoding and URL parameter escaping requirements documented for special characters and accents? [Clarity, Spec §Edge Cases]
## Non-Functional & Operational Readiness
- [x] CHK019 - Is the performance target (< 2.0s for standard single-page extractions) quantified with measurable criteria? [Measurability, Spec §SC-001, Plan §Technical Context]
- [x] CHK020 - Are JSON output compatibility requirements with terminal streaming tools (e.g., `jq`, shell pipes) defined? [Completeness, Spec §SC-005, Contract §3]
- [x] CHK021 - Is the execution isolation assumption (pure RSS feed extraction without mandatory Playwright browser runtime) clearly stated? [Traceability, Spec §Assumptions, Research §Decision 1]
## Notes
- Mark items `[x]` only after review confirms the requirement-quality criterion is satisfied
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers
- `checklists/requirements.md` has a separate built-in lifecycle maintained by `/speckit-specify` and `/speckit-clarify`
- Add comments or findings inline
- Link to relevant resources or documentation
- Items are numbered sequentially for easy reference