feat(extractor): implement multi-engine article content extractor

- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability)
- Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing)
- Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts
- Passed ruff linting/formatting and mypy type checking cleanly
This commit is contained in:
2026-08-20 19:22:20 -03:00
parent 6e3d57619b
commit 6a45368cb0
85 changed files with 18345 additions and 3897 deletions
@@ -0,0 +1,62 @@
# Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
**Purpose**: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
**Created**: 2026-08-20
**Feature**: [spec.md](../spec.md) | [data-model.md](../data-model.md) | [contracts/](../contracts/)
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
---
## 1. Requirement Completeness
- [x] CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
- [x] CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
- [x] CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
- [x] CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
- [x] CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
- [x] CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]
---
## 2. Requirement Clarity & Non-Ambiguity
- [x] CHK007 Is the default naming convention for output JSON (`<input_stem>_extracted.json`) unambiguously specified when `--output` is omitted? [Clarity, Spec §FR-007, Clarifications]
- [x] CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
- [x] CHK009 Is the mechanism for inheriting the `language` parameter with fallback to `"en"` clearly defined? [Clarity, Spec §FR-005, Clarifications]
- [x] CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]
---
## 3. Requirement Consistency & Data Contracts
- [x] CHK011 Do entity field names in `data-model.md` align consistently with the schemas in `contracts/json-schema.md`? [Consistency, DataModel §1.5, Contract §JSON]
- [x] CHK012 Are CLI argument definitions in `contracts/cli-contract.md` aligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI]
- [x] CHK013 Is stream segregation (`stderr` for progress logs, `stdout` for JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]
---
## 4. Scenario & Edge Case Coverage
- [x] CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
- [x] CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
- [x] CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
- [x] CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
- [x] CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]
---
## 5. Non-Functional & Operational Readiness
- [x] CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
- [x] CHK020 Are execution sampling requirements via `--limit` testable within 30 seconds? [Measurability, Spec §SC-004]
---
## Notes
- All 20 items reviewed and validated against `docs/prd_extrator_artigos_nlp.md` and spec artifacts.
- Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
- Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.