Files
andreferraro 6a45368cb0 feat(extractor): implement multi-engine article content extractor
- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability)
- Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing)
- Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts
- Passed ruff linting/formatting and mypy type checking cleanly
2026-08-20 19:22:20 -03:00

63 lines
4.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
**Purpose**: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md) | [data-model.md](../data-model.md) | [contracts/](../contracts/)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
---
## 1. Requirement Completeness
- [x] CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
- [x] CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
- [x] CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
- [x] CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
- [x] CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
- [x] CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]
---
## 2. Requirement Clarity & Non-Ambiguity
- [x] CHK007 Is the default naming convention for output JSON (`<input_stem>_extracted.json`) unambiguously specified when `--output` is omitted? [Clarity, Spec §FR-007, Clarifications]
- [x] CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
- [x] CHK009 Is the mechanism for inheriting the `language` parameter with fallback to `"en"` clearly defined? [Clarity, Spec §FR-005, Clarifications]
- [x] CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]
---
## 3. Requirement Consistency & Data Contracts
- [x] CHK011 Do entity field names in `data-model.md` align consistently with the schemas in `contracts/json-schema.md`? [Consistency, DataModel §1.5, Contract §JSON]
- [x] CHK012 Are CLI argument definitions in `contracts/cli-contract.md` aligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI]
- [x] CHK013 Is stream segregation (`stderr` for progress logs, `stdout` for JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]
---
## 4. Scenario & Edge Case Coverage
- [x] CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
- [x] CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
- [x] CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
- [x] CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
- [x] CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]
---
## 5. Non-Functional & Operational Readiness
- [x] CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
- [x] CHK020 Are execution sampling requirements via `--limit` testable within 30 seconds? [Measurability, Spec §SC-004]
---
## Notes
- All 20 items reviewed and validated against `docs/prd_extrator_artigos_nlp.md` and spec artifacts.
- Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
- Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.