Files
andreferraro 6a45368cb0 feat(extractor): implement multi-engine article content extractor
- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability)
- Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing)
- Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts
- Passed ruff linting/formatting and mypy type checking cleanly
2026-08-20 19:22:20 -03:00

4.6 KiB
Raw Permalink Blame History

Extraction Pipeline Checklist: Article Content Multi-Engine Extractor

Purpose: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
Created: 2026-08-20
Feature: spec.md | data-model.md | contracts/

Note: This custom checklist is generated by the /speckit-checklist command based on feature context and requirements.
Review Ownership: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item [x] only when the reviewer determines the requirements-quality criterion is satisfied.
Marker Semantics: [x] means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.


1. Requirement Completeness

  • CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
  • CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
  • CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
  • CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
  • CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
  • CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]

2. Requirement Clarity & Non-Ambiguity

  • CHK007 Is the default naming convention for output JSON (<input_stem>_extracted.json) unambiguously specified when --output is omitted? [Clarity, Spec §FR-007, Clarifications]
  • CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
  • CHK009 Is the mechanism for inheriting the language parameter with fallback to "en" clearly defined? [Clarity, Spec §FR-005, Clarifications]
  • CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]

3. Requirement Consistency & Data Contracts

  • CHK011 Do entity field names in data-model.md align consistently with the schemas in contracts/json-schema.md? [Consistency, DataModel §1.5, Contract §JSON]
  • CHK012 Are CLI argument definitions in contracts/cli-contract.md aligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI]
  • CHK013 Is stream segregation (stderr for progress logs, stdout for JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]

4. Scenario & Edge Case Coverage

  • CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
  • CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
  • CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
  • CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
  • CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]

5. Non-Functional & Operational Readiness

  • CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
  • CHK020 Are execution sampling requirements via --limit testable within 30 seconds? [Measurability, Spec §SC-004]

Notes

  • All 20 items reviewed and validated against docs/prd_extrator_artigos_nlp.md and spec artifacts.
  • Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
  • Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.