- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability) - Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing) - Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts - Passed ruff linting/formatting and mypy type checking cleanly
4.6 KiB
Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
Purpose: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
Created: 2026-08-20
Feature: spec.md | data-model.md | contracts/
Note: This custom checklist is generated by the /speckit-checklist command based on feature context and requirements.
Review Ownership: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item [x] only when the reviewer determines the requirements-quality criterion is satisfied.
Marker Semantics: [x] means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
1. Requirement Completeness
- CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
- CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
- CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
- CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
- CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
- CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]
2. Requirement Clarity & Non-Ambiguity
- CHK007 Is the default naming convention for output JSON (
<input_stem>_extracted.json) unambiguously specified when--outputis omitted? [Clarity, Spec §FR-007, Clarifications] - CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
- CHK009 Is the mechanism for inheriting the
languageparameter with fallback to"en"clearly defined? [Clarity, Spec §FR-005, Clarifications] - CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]
3. Requirement Consistency & Data Contracts
- CHK011 Do entity field names in
data-model.mdalign consistently with the schemas incontracts/json-schema.md? [Consistency, DataModel §1.5, Contract §JSON] - CHK012 Are CLI argument definitions in
contracts/cli-contract.mdaligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI] - CHK013 Is stream segregation (
stderrfor progress logs,stdoutfor JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]
4. Scenario & Edge Case Coverage
- CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
- CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
- CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
- CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
- CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]
5. Non-Functional & Operational Readiness
- CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
- CHK020 Are execution sampling requirements via
--limittestable within 30 seconds? [Measurability, Spec §SC-004]
Notes
- All 20 items reviewed and validated against
docs/prd_extrator_artigos_nlp.mdand spec artifacts. - Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
- Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.