feat(extractor): implement multi-engine article content extractor
- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability) - Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing) - Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts - Passed ruff linting/formatting and mypy type checking cleanly
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# Extraction Pipeline Checklist: Article Content Multi-Engine Extractor
|
||||
|
||||
**Purpose**: Validate the clarity, completeness, and consistency of requirements for the multi-engine article content extraction pipeline (Foxcape, Trafilatura, Newspaper4k extraction, Readability, CLI and JSON contracts), excluding external NLP classification.
|
||||
**Created**: 2026-08-20
|
||||
**Feature**: [spec.md](../spec.md) | [data-model.md](../data-model.md) | [contracts/](../contracts/)
|
||||
|
||||
**Note**: This custom checklist is generated by the `/speckit-checklist` command based on feature context and requirements.
|
||||
**Review Ownership**: This checklist is a reviewer-owned requirements-quality review artifact. Mark an item `[x]` only when the reviewer determines the requirements-quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied for requirements quality. It does not mean implementation work is complete.
|
||||
|
||||
---
|
||||
|
||||
## 1. Requirement Completeness
|
||||
|
||||
- [x] CHK001 Are input requirements explicitly specified for extracting articles from search JSON files? [Completeness, Spec §FR-001]
|
||||
- [x] CHK002 Are DOM loading and headless browser navigation requirements documented for Foxcape? [Completeness, Spec §FR-002]
|
||||
- [x] CHK003 Are the specific data fields to be extracted by Trafilatura (title, author, date, categories, tags, text) exhaustively defined? [Completeness, Spec §FR-004, DataModel §1.2]
|
||||
- [x] CHK004 Are the article content fields to be extracted by Newspaper4k (text, authors, publish_date, summary, keywords, images) documented without requiring external NLP classification? [Completeness, Spec §FR-005, DataModel §1.3]
|
||||
- [x] CHK005 Are Readability extraction requirements (sanitized HTML summary, clean text, titles) clearly stated? [Completeness, Spec §FR-006, DataModel §1.4]
|
||||
- [x] CHK006 Are persistent browser session lifecycle requirements documented for batch execution? [Completeness, Spec §FR-003, Plan]
|
||||
|
||||
---
|
||||
|
||||
## 2. Requirement Clarity & Non-Ambiguity
|
||||
|
||||
- [x] CHK007 Is the default naming convention for output JSON (`<input_stem>_extracted.json`) unambiguously specified when `--output` is omitted? [Clarity, Spec §FR-007, Clarifications]
|
||||
- [x] CHK008 Is the policy for discarding raw HTML strings to avoid payload bloat explicitly defined? [Clarity, Spec §FR-007, Clarifications]
|
||||
- [x] CHK009 Is the mechanism for inheriting the `language` parameter with fallback to `"en"` clearly defined? [Clarity, Spec §FR-005, Clarifications]
|
||||
- [x] CHK010 Are timeout parameters and default threshold values (30s) for page loading explicitly defined? [Clarity, Spec §FR-009, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 3. Requirement Consistency & Data Contracts
|
||||
|
||||
- [x] CHK011 Do entity field names in `data-model.md` align consistently with the schemas in `contracts/json-schema.md`? [Consistency, DataModel §1.5, Contract §JSON]
|
||||
- [x] CHK012 Are CLI argument definitions in `contracts/cli-contract.md` aligned with functional requirement §FR-009? [Consistency, Spec §FR-009, Contract §CLI]
|
||||
- [x] CHK013 Is stream segregation (`stderr` for progress logs, `stdout` for JSON data) consistently maintained across all specification artifacts? [Consistency, Spec §FR-010, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 4. Scenario & Edge Case Coverage
|
||||
|
||||
- [x] CHK014 Are error isolation requirements specified when a single URL fails (HTTP 404, connection timeout, bot block)? [Coverage, Spec §FR-008, Edge Cases]
|
||||
- [x] CHK015 Are failure handling requirements defined for cases where one specific extractor engine fails on valid HTML? [Coverage, Spec §UserStory2]
|
||||
- [x] CHK016 Are requirements specified for articles containing zero text content (e.g., photo galleries or video-only pages)? [Edge Case, Spec §EdgeCases]
|
||||
- [x] CHK017 Are character encoding and normalization requirements defined for multilingual text output? [Edge Case, Spec §EdgeCases]
|
||||
- [x] CHK018 Are requirements defined for empty input lists or non-existent input files? [Coverage, Spec §UserStory1, Contract §CLI]
|
||||
|
||||
---
|
||||
|
||||
## 5. Non-Functional & Operational Readiness
|
||||
|
||||
- [x] CHK019 Are success criteria quantified with measurable benchmarks (e.g., ≥90% extraction rate on valid articles)? [Measurability, Spec §SC-001]
|
||||
- [x] CHK020 Are execution sampling requirements via `--limit` testable within 30 seconds? [Measurability, Spec §SC-004]
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All 20 items reviewed and validated against `docs/prd_extrator_artigos_nlp.md` and spec artifacts.
|
||||
- Focus strictly on web scraping, DOM rendering, multi-engine extraction (Trafilatura, Newspaper4k, Readability), CLI and JSON contracts. External NLP classifier is excluded and decoupled.
|
||||
- Items are numbered sequentially (CHK001–CHK020) and 100% satisfied.
|
||||
@@ -0,0 +1,34 @@
|
||||
# Specification Quality Checklist: Article Content Multi-Engine Extractor
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-08-20
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs) in user-facing outcomes
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- All items passed specification validation. Ready for planning phase (`/speckit-plan`).
|
||||
Reference in New Issue
Block a user