Files
TextNLPClassifierApp/requirements.txt
T
andreferraro 6a45368cb0 feat(extractor): implement multi-engine article content extractor
- Added scripts/extract_article_contents.py for batch scraping with stealth Foxcape and triple extraction (Trafilatura, Newspaper4k, Readability)
- Created unit, integration, and E2E test suite in tests/test_extract_article_contents.py (90/90 passing)
- Updated specs/003-article-content-extractor and README.md with usage documentation and CLI contracts
- Passed ruff linting/formatting and mypy type checking cleanly
2026-08-20 19:22:20 -03:00

14 lines
313 B
Plaintext

pytest>=7.0.0
foxcape>=0.1.1
beautifulsoup4>=4.12.0
googlenewsdecoder>=0.1.7
selectolax>=0.3.27
trafilatura>=1.8.0
newspaper4k>=0.9.3.1
readability-lxml>=0.8.1
lxml>=4.9.0
# Optional Tier 2 / Tier 3 dependencies (not required for POC core execution)
# sentence-transformers>=2.2.0
# httpx>=0.24.0
# openai>=1.0.0