feat(media-routing): implement 007 media article routing, runtime architecture diagram and update graphify knowledge graph
This commit is contained in:
@@ -1,50 +1,61 @@
|
||||
# [PROJECT_NAME] Constitution
|
||||
<!-- Example: Spec Constitution, TaskFlow Constitution, etc. -->
|
||||
<!--
|
||||
Sync Impact Report
|
||||
Version change: Initial Template -> 1.0.0
|
||||
Modified principles: Initialized all Core Principles from template placeholders
|
||||
Added sections:
|
||||
- Core Principles:
|
||||
- I. Modularity & CLI-First Interoperability
|
||||
- II. Determinism, Atomic Operations & Data Integrity
|
||||
- III. Multi-Engine Consensus & Fault-Tolerant Fallback
|
||||
- IV. Test-First & Empirical Validation
|
||||
- V. Observability, Structured Logging & Traceability
|
||||
- Technical, Security & Environmental Constraints
|
||||
- Development Workflow & Quality Gates
|
||||
- Governance & Amendment Protocol
|
||||
Removed sections: None
|
||||
Follow-up TODOs: None
|
||||
-->
|
||||
|
||||
# TextNLPClassifierApp Constitution
|
||||
|
||||
## Core Principles
|
||||
|
||||
### [PRINCIPLE_1_NAME]
|
||||
<!-- Example: I. Library-First -->
|
||||
[PRINCIPLE_1_DESCRIPTION]
|
||||
<!-- Example: Every feature starts as a standalone library; Libraries must be self-contained, independently testable, documented; Clear purpose required - no organizational-only libraries -->
|
||||
### I. Modularity & CLI-First Interoperability
|
||||
Every capability (crawling, multi-engine extraction, consensus selection, markdown conversion, NLP/LLM classification, and consolidation runtime) MUST be implemented as a modular, decoupled component under `src/` or `scripts/`. Every core component MUST expose a standard CLI interface supporting both human-readable logging and structured JSON input/output over standard streams, with explicit exit codes for automated pipeline orchestration.
|
||||
|
||||
### [PRINCIPLE_2_NAME]
|
||||
<!-- Example: II. CLI Interface -->
|
||||
[PRINCIPLE_2_DESCRIPTION]
|
||||
<!-- Example: Every library exposes functionality via CLI; Text in/out protocol: stdin/args → stdout, errors → stderr; Support JSON + human-readable formats -->
|
||||
### II. Determinism, Atomic Operations & Data Integrity
|
||||
Text processing, shingle extraction, $F_1$ consensus calculation, metadata matrix resolution, and formatting transformations MUST be strictly deterministic, reproducible, and idempotent. All filesystem writes MUST execute via atomic transactional write patterns (temporary file staging followed by atomic rename) to eliminate corrupted state. Source data and multi-extractor payloads MUST be preserved non-destructively.
|
||||
|
||||
### [PRINCIPLE_3_NAME]
|
||||
<!-- Example: III. Test-First (NON-NEGOTIABLE) -->
|
||||
[PRINCIPLE_3_DESCRIPTION]
|
||||
<!-- Example: TDD mandatory: Tests written → User approved → Tests fail → Then implement; Red-Green-Refactor cycle strictly enforced -->
|
||||
### III. Multi-Engine Consensus & Fault-Tolerant Fallback
|
||||
Data extraction and parsing MUST NOT rely on single points of failure. The architecture enforces multi-engine extraction (Trafilatura, Newspaper4k, Readability) coupled with automated consensus scoring ($F_1$ 5-token shingles) and strict fallback heuristics. Edge cases (paywalls, empty bodies, anti-bot challenges) MUST be handled gracefully with explicit diagnostics.
|
||||
|
||||
### [PRINCIPLE_4_NAME]
|
||||
<!-- Example: IV. Integration Testing -->
|
||||
[PRINCIPLE_4_DESCRIPTION]
|
||||
<!-- Example: Focus areas requiring integration tests: New library contract tests, Contract changes, Inter-service communication, Shared schemas -->
|
||||
### IV. Test-First & Empirical Validation (TDD & Regressions)
|
||||
Test-Driven Development (TDD) and empirical test suites are mandatory. Unit tests, integration contracts, and regression suites MUST cover parser algorithms, CLI flags, exit codes, and error conditions before deployment. A 100% passing test baseline MUST be maintained (`pytest`), accompanied by static linting (`ruff`) and static type validation (`mypy`/`pyright`).
|
||||
|
||||
### [PRINCIPLE_5_NAME]
|
||||
<!-- Example: V. Observability, VI. Versioning & Breaking Changes, VII. Simplicity -->
|
||||
[PRINCIPLE_5_DESCRIPTION]
|
||||
<!-- Example: Text I/O ensures debuggability; Structured logging required; Or: MAJOR.MINOR.BUILD format; Or: Start simple, YAGNI principles -->
|
||||
### V. Observability, Structured Logging & Traceability
|
||||
All pipeline phases (crawling, article selection, consolidation, LLM/NLP classification) MUST emit structured, contextual telemetry. Operations MUST log execution metrics, engine scores, token consumption, and decision paths. Tracing integrations (e.g., Langfuse) MUST provide full visibility into prompt performance, latency, cost, and classification rationale.
|
||||
|
||||
## [SECTION_2_NAME]
|
||||
<!-- Example: Additional Constraints, Security Requirements, Performance Standards, etc. -->
|
||||
## Technical, Security & Environmental Constraints
|
||||
|
||||
[SECTION_2_CONTENT]
|
||||
<!-- Example: Technology stack requirements, compliance standards, deployment policies, etc. -->
|
||||
- **Language & Runtime**: Python `>=3.10` with strict type annotations across all modules.
|
||||
- **Code Quality**: Linting and formatting governed by `ruff` (100-character line limit) and type safety verified via `mypy`.
|
||||
- **Security & Secrets**: Zero hardcoded credentials or API keys; configuration MUST be loaded from environment variables (`.env`) with schemas documented in `.env.example`.
|
||||
- **Stealth & Web Automation**: Browser automation (Foxcape/Camoufox) MUST enforce appropriate delays, backoff, and fingerprint management without exceeding target service rate limits.
|
||||
|
||||
## [SECTION_3_NAME]
|
||||
<!-- Example: Development Workflow, Review Process, Quality Gates, etc. -->
|
||||
## Development Workflow & Quality Gates
|
||||
|
||||
[SECTION_3_CONTENT]
|
||||
<!-- Example: Code review requirements, testing gates, deployment approval process, etc. -->
|
||||
- **Specification First**: Features and structural changes MUST be documented in specifications (`specs/` or `docs/`) with clear functional requirements and architecture decision records (ADRs).
|
||||
- **Mandatory Quality Gates**: Every code contribution MUST pass:
|
||||
1. `ruff check` and `ruff format --check` (clean formatting and linting).
|
||||
2. `mypy` / `pyright` (no type check violations).
|
||||
3. `pytest` (full test suite execution with 100% passing rate).
|
||||
- **Commit Standards**: Atomic commits using Conventional Commits convention (`feat:`, `fix:`, `docs:`, `test:`, `refactor:`, `chore:`).
|
||||
|
||||
## Governance
|
||||
<!-- Example: Constitution supersedes all other practices; Amendments require documentation, approval, migration plan -->
|
||||
|
||||
[GOVERNANCE_RULES]
|
||||
<!-- Example: All PRs/reviews must verify compliance; Complexity must be justified; Use [GUIDANCE_FILE] for runtime development guidance -->
|
||||
This Constitution represents the supreme architectural and development policy for `TextNLPClassifierApp`. All development, automated subagents, and pull requests MUST comply with the principles and quality gates established herein. Amendments require explicit documentation of rationale, team review, and formal version increments:
|
||||
- **MAJOR**: Incompatible principle removals or foundational architectural shifts.
|
||||
- **MINOR**: Addition of new principles, governance rules, or expanded standards.
|
||||
- **PATCH**: Wording improvements, clarifications, and non-semantic corrections.
|
||||
|
||||
**Version**: [CONSTITUTION_VERSION] | **Ratified**: [RATIFICATION_DATE] | **Last Amended**: [LAST_AMENDED_DATE]
|
||||
<!-- Example: Version: 2.1.1 | Ratified: 2025-06-13 | Last Amended: 2025-07-16 -->
|
||||
**Version**: 1.0.0 | **Ratified**: 2026-08-24 | **Last Amended**: 2026-08-24
|
||||
|
||||
@@ -15,39 +15,13 @@
|
||||
- [Visão Geral](#-visão-geral)
|
||||
- [Instalação e Setup](#-instalação-e-setup)
|
||||
- [1. Classificador de Conteúdo e Inerência (NLP / LLM / ECP)](#1--classificador-de-conteúdo-e-inerência-nlp--llm--ecp)
|
||||
- [O que é e Como Funciona](#o-que-é-e-como-funciona)
|
||||
- [Arquitetura de Classificação em 3 Tiers](#arquitetura-de-classificação-em-3-tiers)
|
||||
- [Categorias de Decisão](#categorias-de-decisão)
|
||||
- [Formato do ECP Snapshot e Markdown](#formato-do-ecp-snapshot-e-markdown)
|
||||
- [Exemplos de Uso CLI](#exemplos-de-uso-cli)
|
||||
- [2. Extrator de Manchetes do Google News](#2--extrator-de-manchetes-do-google-news)
|
||||
- [O que é e Como Funciona](#o-que-é-e-como-funciona-1)
|
||||
- [Diferenciais Técnicos](#diferenciais-técnicos)
|
||||
- [Argumentos e Flags de Linha de Comando](#argumentos-e-flags-de-linha-de-comando)
|
||||
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso)
|
||||
- [3. Extrator e Parser Multimotor de Artigos](#3--extrator-e-parser-multimotor-de-artigos)
|
||||
- [Visão Geral e Tríplice Extração](#visão-geral-e-tríplice-extração)
|
||||
- [Argumentos e Flags CLI](#argumentos-e-flags-cli)
|
||||
- [Exemplos de Uso](#exemplos-de-uso)
|
||||
- [4. Seletor Determinístico de Conteúdo de Artigos](#4--seletor-determinístico-de-conteúdo-de-artigos)
|
||||
- [Visão Geral e Algoritmo de Consenso ($F_1$)](#visão-geral-e-algoritmo-de-consenso-f_1)
|
||||
- [Pipeline de Normalização e Shingles](#pipeline-de-normalização-e-shingles)
|
||||
- [Critérios de Desempate Técnico e Resiliência](#critérios-de-desempate-técnico-e-resiliência)
|
||||
- [Argumentos e Flags CLI](#argumentos-e-flags-cli-1)
|
||||
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso-1)
|
||||
- [5. Conversor de Artigo JSON para Markdown](#5--conversor-de-artigo-json-para-markdown)
|
||||
- [Visão Geral e Estrutura do Documento](#visão-geral-e-estrutura-do-documento)
|
||||
- [Isolamento Estrito de Extratores e Fallback](#isolamento-estrito-de-extratores-e-fallback)
|
||||
- [Matriz Determinística de Metadados](#matriz-determinística-de-metadados)
|
||||
- [Sanitização Editorial e Deduplicação](#sanitização-editorial-e-deduplicação)
|
||||
- [Argumentos e Flags CLI](#argumentos-e-flags-cli-2)
|
||||
- [Exemplos Práticos de Uso](#exemplos-práticos-de-uso-2)
|
||||
- [6. Runtime de Consolidação e Higienização de Artigos (006-article-consolidation-runtime)](#6--runtime-de-consolidação-e-higienização-de-artigos-006-article-consolidation-runtime)
|
||||
- [Visão Geral e Arquitetura](#visão-geral-e-arquitetura-do-runtime)
|
||||
- [Configuração de Ambiente (.env)](#configuração-de-ambiente-env)
|
||||
- [Comandos e Utilitários CLI](#comandos-e-utilitários-cli)
|
||||
- [O que Esperar do Resultado (Artefatos Gerados)](#o-que-esperar-do-resultado-artefatos-gerados)
|
||||
- [Contrato de Códigos de Saída (Exit Codes)](#contrato-de-códigos-de-saída-exit-codes)
|
||||
- [7. Classificação e Roteamento de Mídia de Artigos (007-media-article-routing)](#7--classificação-e-roteamento-de-mídia-de-artigos-007-media-article-routing)
|
||||
- [8. Arquitetura do Sistema e Grafo de Conhecimento (Archify & Graphify)](#8--arquitetura-do-sistema-e-grafo-de-conhecimento-archify--graphify)
|
||||
- [Estrutura do Projeto](#-estrutura-do-projeto)
|
||||
- [Testes e Qualidade de Código](#-testes-e-qualidade-de-código)
|
||||
- [Licença](#-licença)
|
||||
@@ -60,9 +34,11 @@ O **TextNLPClassifierApp** reúne um ecossistema completo de ferramentas de enge
|
||||
|
||||
1. **`classify.py`**: Motor de classificação semântica e contextual que determina o grau de aderência e inerência de um documento Markdown em relação a uma entidade alvo descrita em um **ECP Snapshot (Entity Context Profile)**.
|
||||
2. **`scripts/extract_google_news.py`**: Extrator de notícias por palavra-chave, idioma e região geográfica utilizando navegação stealth **Foxcape** (headless), decodificação paralela de URLs para links reais e feedback em tempo real.
|
||||
3. **`scripts/extract_article_contents.py`**: Extrator e parser de artigos multimotor com navegação stealth Foxcape headless e extração simultânea via **Trafilatura**, **Newspaper4k** (NLP) e **Readability**, consolidando texto higienizado, autores, datas, imagens e resumos.
|
||||
3. **`scripts/extract_article_contents.py`**: Extrator e parser de artigos multimotor com navegação stealth Foxcape headless, extração simultânea via **Trafilatura**, **Newspaper4k** (NLP) e **Readability**, detecção de candidatos de mídia (`video`, `audio`, `gallery`) e classificação de layout.
|
||||
4. **`scripts/select_article_extractor.py`**: Motor determinístico de seleção de extratores que avalia as saídas dos três motores, aplica normalização em memória, calcula métricas de consenso de shingles (5-tokens) com pontuação $F_1$, desempata tecnicamente ($\le 0.03$) favorecendo menor concisão/ruído e enriquece os dados de forma não-destrutiva e atômica.
|
||||
5. **`scripts/convert_article_to_markdown.py`**: Conversor determinístico que recebe o JSON de um único artigo selecionado, isola estritamente o corpo do extrator vencedor (`trafilatura`, `newspaper4k` ou `readability`), resolve metadados editoriais por prioridade estrita, higieniza links/imagens/cabeçalhos e gera um documento Markdown (`.md`) padronizado com gravação atômica transacional.
|
||||
6. **`src/runtime/`**: Runtime de consolidação, higienização (LLM grounding), avaliação ECP e persistência transacional SQLite WAL.
|
||||
7. **Detecção e Roteamento de Mídia (`007-media-article-routing`)**: Análise de payloads editoriais para roteamento inteligente de artigos com mídia incorporada ou bypass direto para artigos puramente textuais.
|
||||
|
||||
---
|
||||
|
||||
@@ -636,6 +612,36 @@ O CLI segue estritamente os códigos de saída normativos:
|
||||
|
||||
---
|
||||
|
||||
## 7. 🎬 Classificação e Roteamento de Mídia de Artigos (007-media-article-routing)
|
||||
|
||||
### Visão Geral
|
||||
O subsistema de **Roteamento de Artigos de Mídia** (`specs/007-media-article-routing`) analisa o HTML e metadados extraídos para identificar conteúdos multimídia (vídeos incorporados do YouTube/Vimeo/DailyMotion, podcasts/áudio, galerias fotográficas e layouts orientados a mídia) e rotear adequadamente os fluxos de trabalho no pipeline:
|
||||
|
||||
- **Bypass Direto**: Artigos puramente textuais continuam no fluxo padrão de consolidação (`006-article-consolidation-runtime`) com custo e latência mínimos.
|
||||
- **Roteamento de Mídia**: Artigos com player de vídeo primário ou áudio são marcados com payload enriquecido de mídia (`video_url`, `embed_code`, `duration`, `media_type`), permitindo transcrição posterior via Whisper ou indexação multimídia especializada.
|
||||
|
||||
---
|
||||
|
||||
## 8. 🏛️ Arquitetura do Sistema e Grafo de Conhecimento (Archify & Graphify)
|
||||
|
||||
### 📐 Diagrama de Arquitetura Interativo (Archify)
|
||||
O repositório inclui um diagrama de arquitetura de tempo de execução formalmente validado e renderizado via **Archify**:
|
||||
|
||||
* **Diagrama Interativo**: [`docs/archify/runtime-architecture.html`](docs/archify/runtime-architecture.html)
|
||||
* **Especificação Declarativa**: [`docs/archify/runtime-architecture.json`](docs/archify/runtime-architecture.json)
|
||||
* **Relatório de Validação Visual**: [`docs/archify/runtime-architecture.visual-check.html`](docs/archify/runtime-architecture.visual-check.html)
|
||||
|
||||
O diagrama cobre os 11 componentes centrais da plataforma, o caminho primário ponta a ponta (desde a ingestão stealth até o classificador de inerência e persistência SQLite), 4 fronteiras de segurança (*Untrusted Public Web*, *Ingestion Sandbox*, *Trusted Core Processing Domain*, *External AI Provider Gateway*) e cards de detalhes contextuais.
|
||||
|
||||
### 🌐 Grafo de Conhecimento Navegável (Graphify)
|
||||
A base de código e seus artefatos estão indexados em um grafo de conhecimento completo gerado pelo **Graphify**:
|
||||
|
||||
* **Visualizador do Grafo**: [`graphify-out/graph.html`](graphify-out/graph.html) *(visualização agregada por comunidades navegável em qualquer navegador)*
|
||||
* **Relatório de Auditoria**: [`graphify-out/GRAPH_REPORT.md`](graphify-out/GRAPH_REPORT.md)
|
||||
* **Grafo Estruturado**: [`graphify-out/graph.json`](graphify-out/graph.json) *(GraphRAG / BFS / DFS)*
|
||||
|
||||
---
|
||||
|
||||
## 📁 Estrutura do Projeto
|
||||
|
||||
```text
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 178 KiB After Width: | Height: | Size: 178 KiB |
@@ -5,7 +5,7 @@
|
||||
"status": "pass",
|
||||
"visualReview": "pending",
|
||||
"artifact": {
|
||||
"path": "C:\\Users\\aferr\\Projects\\AFTech\\DunaMedia\\TextNLPClassifierApp\\docs\\runtime-architecture.html",
|
||||
"path": "C:\\Users\\aferr\\Projects\\AFTech\\DunaMedia\\TextNLPClassifierApp\\docs\\archify\\runtime-architecture.html",
|
||||
"sha256": "66284a667f9c3ec5afd8d275fdd83e5980425836788aec8d650c95cfffaa9d1f",
|
||||
"bytes": 693750
|
||||
},
|
||||
|
||||
@@ -0,0 +1,368 @@
|
||||
# ADR-001 — Classificar e rotear conteúdo predominantemente de mídia antes do multimotor
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
---
|
||||
|
||||
# Contexto
|
||||
|
||||
O pipeline atual obtém a página completamente carregada através do crawler e, posteriormente, executa múltiplos motores de extração textual.
|
||||
|
||||
Esse desenho é adequado para artigos cujo conteúdo principal é texto.
|
||||
|
||||
Entretanto, alguns veículos publicam páginas onde:
|
||||
|
||||
* existe apenas uma breve introdução textual;
|
||||
* o conteúdo principal é um vídeo, imagem, conjunto de imagens ou conteúdo incorporado.
|
||||
|
||||
Processar essas páginas com os motores textuais é desnecessário e pode gerar resultados pobres ou irrelevantes.
|
||||
|
||||
A responsabilidade de identificar esses casos deve ser adicionada sem alterar o crawler e sem transformar o pipeline em uma nova arquitetura.
|
||||
|
||||
---
|
||||
|
||||
# Decisão
|
||||
|
||||
Adicionar uma etapa de classificação imediatamente após o crawl e antes do multimotor.
|
||||
|
||||
Fluxo:
|
||||
|
||||
```text
|
||||
Crawler
|
||||
↓
|
||||
DOM completamente carregada
|
||||
↓
|
||||
Media Candidate Detection
|
||||
↓
|
||||
├── sem mídia candidata
|
||||
│ ↓
|
||||
│ multimotor atual
|
||||
│
|
||||
└── com mídia candidata
|
||||
↓
|
||||
Media Content Classifier
|
||||
↓
|
||||
├── text
|
||||
│ ↓
|
||||
│ multimotor atual
|
||||
│
|
||||
└── media
|
||||
↓
|
||||
JSON separado
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Detecção estrutural
|
||||
|
||||
Antes do LLM deve existir apenas uma análise estrutural simples da DOM.
|
||||
|
||||
Objetivo:
|
||||
|
||||
> verificar se existe algum elemento de mídia que justifique a classificação.
|
||||
|
||||
A análise deve operar sobre a DOM/HTML já carregada pelo crawler.
|
||||
|
||||
É permitido utilizar:
|
||||
|
||||
* parser HTML;
|
||||
* navegação por nós;
|
||||
* tags;
|
||||
* atributos estruturais;
|
||||
* relações entre elementos;
|
||||
* contagem de elementos.
|
||||
|
||||
É proibido utilizar:
|
||||
|
||||
* regex;
|
||||
* listas de palavras específicas por idioma;
|
||||
* heurísticas semânticas por idioma.
|
||||
|
||||
Essa etapa não decide se a publicação é `media`.
|
||||
|
||||
Ela decide apenas se há motivo para chamar o classificador.
|
||||
|
||||
---
|
||||
|
||||
# Conteúdo enviado ao modelo
|
||||
|
||||
O modelo deve receber uma representação compacta dos dados já existentes na página.
|
||||
|
||||
Devem ser utilizados somente dados necessários para a decisão, como:
|
||||
|
||||
```text
|
||||
título
|
||||
blocos textuais relevantes
|
||||
presença de imagem
|
||||
quantidade de imagens
|
||||
presença de vídeo
|
||||
presença de elementos incorporados
|
||||
```
|
||||
|
||||
Não enviar mídia binária.
|
||||
|
||||
Não realizar chamadas externas para compreender a mídia.
|
||||
|
||||
Não enviar HTML completo quando uma representação compacta da estrutura puder fornecer a mesma informação.
|
||||
|
||||
---
|
||||
|
||||
# Responsabilidade do classificador
|
||||
|
||||
O classificador responde uma única pergunta conceitual:
|
||||
|
||||
> O texto desta publicação possui conteúdo jornalístico substancial por si próprio ou funciona essencialmente como uma breve introdução, contextualização ou descrição da mídia presente?
|
||||
|
||||
Se o texto for substancial:
|
||||
|
||||
```json
|
||||
{
|
||||
"content_type": "text",
|
||||
"media_type": null
|
||||
}
|
||||
```
|
||||
|
||||
Se a mídia for o conteúdo principal:
|
||||
|
||||
```json
|
||||
{
|
||||
"content_type": "media",
|
||||
"media_type": "..."
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Tipos permitidos
|
||||
|
||||
```text
|
||||
video
|
||||
image
|
||||
images
|
||||
embed
|
||||
mixed
|
||||
```
|
||||
|
||||
Não criar subtipos adicionais.
|
||||
|
||||
---
|
||||
|
||||
# Regra para múltiplas imagens
|
||||
|
||||
Não é necessário identificar tecnicamente um componente carousel.
|
||||
|
||||
Se múltiplas imagens constituem o conteúdo principal, o resultado deve ser:
|
||||
|
||||
```text
|
||||
images
|
||||
```
|
||||
|
||||
Independentemente de serem apresentadas como:
|
||||
|
||||
* carousel;
|
||||
* slideshow;
|
||||
* galeria;
|
||||
* sequência vertical;
|
||||
* qualquer outra composição visual.
|
||||
|
||||
---
|
||||
|
||||
# Regra para conteúdo misto
|
||||
|
||||
Quando mais de uma categoria de mídia constituir o conteúdo principal:
|
||||
|
||||
```text
|
||||
mixed
|
||||
```
|
||||
|
||||
Não criar precedência artificial como:
|
||||
|
||||
```text
|
||||
video > image
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Artigos textuais
|
||||
|
||||
Um artigo continua sendo `text` mesmo contendo mídia quando existe conteúdo jornalístico textual substancial.
|
||||
|
||||
Portanto:
|
||||
|
||||
```text
|
||||
presença de mídia ≠ classificação media
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Artigos muito curtos
|
||||
|
||||
Uma publicação extremamente curta com uma imagem pode ser classificada como `media/image`.
|
||||
|
||||
Não é necessário tentar preservar esse conteúdo como artigo textual apenas porque existe algum texto.
|
||||
|
||||
---
|
||||
|
||||
# Posicionamento arquitetural
|
||||
|
||||
A nova etapa deve permanecer fora de:
|
||||
|
||||
```text
|
||||
Trafilatura
|
||||
Newspaper4k
|
||||
Readability
|
||||
```
|
||||
|
||||
Nenhum dos três motores é responsável pela classificação.
|
||||
|
||||
O método equivalente ao atual `extract_all_engines()` somente deve ser executado depois que o novo roteamento determinar:
|
||||
|
||||
```text
|
||||
content_type = text
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Saída física
|
||||
|
||||
Artigos textuais:
|
||||
|
||||
```text
|
||||
*_extracted.json
|
||||
```
|
||||
|
||||
Artigos predominantemente de mídia:
|
||||
|
||||
```text
|
||||
*_media.json
|
||||
```
|
||||
|
||||
A classificação de mídia não deve aparecer misturada aos artigos textuais processados com sucesso.
|
||||
|
||||
---
|
||||
|
||||
# Dados preservados para mídia
|
||||
|
||||
Cada registro de mídia deve preservar somente informações básicas necessárias à rastreabilidade:
|
||||
|
||||
```text
|
||||
input_meta
|
||||
crawled_url
|
||||
page_title
|
||||
http_status
|
||||
content_type
|
||||
media_type
|
||||
```
|
||||
|
||||
Não existe extração de mídia nesta feature.
|
||||
|
||||
---
|
||||
|
||||
# Alternativas consideradas
|
||||
|
||||
## Executar primeiro o multimotor e identificar mídia depois
|
||||
|
||||
Rejeitada.
|
||||
|
||||
Motivos:
|
||||
|
||||
* executa trabalho desnecessário;
|
||||
* mistura responsabilidades;
|
||||
* não evita custo de processamento;
|
||||
* mantém páginas inadequadas dentro do pipeline textual.
|
||||
|
||||
---
|
||||
|
||||
## Usar Newspaper4k para descobrir mídia
|
||||
|
||||
Rejeitada.
|
||||
|
||||
Embora Newspaper4k possa expor informações de imagens e vídeos, utilizá-lo significaria executar parte do multimotor justamente nos conteúdos que a nova feature pretende desviar antes do multimotor.
|
||||
|
||||
---
|
||||
|
||||
## Classificar tudo apenas com regras
|
||||
|
||||
Rejeitada.
|
||||
|
||||
A decisão:
|
||||
|
||||
```text
|
||||
texto jornalístico curto
|
||||
```
|
||||
|
||||
versus:
|
||||
|
||||
```text
|
||||
texto que apenas descreve uma mídia
|
||||
```
|
||||
|
||||
é semântica e precisa funcionar em até 10 idiomas.
|
||||
|
||||
Criar heurísticas específicas para resolver isso aumentaria código, manutenção e fragilidade.
|
||||
|
||||
---
|
||||
|
||||
## Utilizar regex
|
||||
|
||||
Rejeitada e proibida por requisito.
|
||||
|
||||
---
|
||||
|
||||
## Utilizar modelo multimodal
|
||||
|
||||
Rejeitada.
|
||||
|
||||
O sistema não precisa entender a mídia.
|
||||
|
||||
Precisa apenas determinar se o texto é o conteúdo principal ou se serve de introdução à mídia.
|
||||
|
||||
---
|
||||
|
||||
# Consequências positivas
|
||||
|
||||
* reduz processamento desnecessário;
|
||||
* mantém o multimotor focado em texto;
|
||||
* separa claramente conteúdos de naturezas diferentes;
|
||||
* mantém a implementação pequena;
|
||||
* evita regras linguísticas;
|
||||
* funciona de forma multilíngue;
|
||||
* não introduz nova infraestrutura;
|
||||
* permite evolução futura do pipeline de mídia de forma independente.
|
||||
|
||||
---
|
||||
|
||||
# Consequências aceitas
|
||||
|
||||
Alguns artigos textuais muito curtos acompanhados de imagem poderão ser classificados como `media/image`.
|
||||
|
||||
Esse comportamento é deliberado e aceito, pois conteúdos textuais extremamente pobres não são úteis para o objetivo do pipeline textual.
|
||||
|
||||
---
|
||||
|
||||
# Invariantes
|
||||
|
||||
A implementação deve sempre preservar:
|
||||
|
||||
```text
|
||||
MEDIA
|
||||
→ nunca executa multimotor
|
||||
|
||||
TEXT
|
||||
→ segue pipeline atual
|
||||
|
||||
classification_failed
|
||||
→ não assume TEXT
|
||||
→ não assume MEDIA
|
||||
```
|
||||
|
||||
E:
|
||||
|
||||
```text
|
||||
nenhuma regex
|
||||
nenhum download de mídia
|
||||
nenhuma análise multimodal
|
||||
nenhuma interação com carousel
|
||||
```
|
||||
@@ -0,0 +1,516 @@
|
||||
# ADR-002 — Utilizar Qwen3.5 2B local com fallback operacional Groq e OmniRoute
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
---
|
||||
|
||||
# Contexto
|
||||
|
||||
A classificação entre:
|
||||
|
||||
```text
|
||||
text
|
||||
```
|
||||
|
||||
e:
|
||||
|
||||
```text
|
||||
media
|
||||
```
|
||||
|
||||
exige uma pequena decisão semântica multilíngue.
|
||||
|
||||
O problema não exige um modelo grande.
|
||||
|
||||
A tarefa é extremamente restrita:
|
||||
|
||||
1. receber texto e informações estruturais;
|
||||
2. decidir se o texto é conteúdo jornalístico substancial ou apenas descrição/contextualização da mídia;
|
||||
3. quando for mídia, identificar uma entre cinco categorias.
|
||||
|
||||
O sistema precisa ser production-ready sem introduzir infraestrutura ou complexidade desnecessária.
|
||||
|
||||
---
|
||||
|
||||
# Decisão
|
||||
|
||||
Utilizar a seguinte cadeia sequencial:
|
||||
|
||||
```text
|
||||
Qwen3.5 2B / Ollama
|
||||
↓ falha operacional
|
||||
GPT-OSS 20B / Groq
|
||||
↓ falha operacional
|
||||
cgpt-web/gpt-5.5 / OmniRoute
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Provider primário
|
||||
|
||||
```text
|
||||
Qwen3.5 2B
|
||||
```
|
||||
|
||||
Runtime:
|
||||
|
||||
```text
|
||||
Ollama
|
||||
```
|
||||
|
||||
Motivos da decisão:
|
||||
|
||||
* execução local;
|
||||
* modelo pequeno;
|
||||
* suficiente para classificação restrita;
|
||||
* adequado ao cenário multilíngue;
|
||||
* elimina custo por chamada no caminho normal;
|
||||
* não depende de serviço externo durante operação normal.
|
||||
|
||||
---
|
||||
|
||||
# Primeiro fallback
|
||||
|
||||
```text
|
||||
GPT-OSS 20B
|
||||
```
|
||||
|
||||
Provider:
|
||||
|
||||
```text
|
||||
Groq
|
||||
```
|
||||
|
||||
O Groq somente é chamado quando o provider primário não consegue fornecer uma classificação tecnicamente utilizável.
|
||||
|
||||
---
|
||||
|
||||
# Segundo fallback
|
||||
|
||||
Modelo:
|
||||
|
||||
```text
|
||||
cgpt-web/gpt-5.5
|
||||
```
|
||||
|
||||
Provider:
|
||||
|
||||
```text
|
||||
OmniRoute
|
||||
```
|
||||
|
||||
O OmniRoute somente é utilizado quando:
|
||||
|
||||
```text
|
||||
Ollama falhou
|
||||
E
|
||||
Groq falhou
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Tipo de fallback
|
||||
|
||||
O fallback é exclusivamente operacional.
|
||||
|
||||
Exemplos de falha que justificam fallback:
|
||||
|
||||
```text
|
||||
conexão recusada
|
||||
timeout
|
||||
erro HTTP
|
||||
provider indisponível
|
||||
erro de execução
|
||||
resposta que não pode ser consumida
|
||||
resposta incompatível com o schema obrigatório
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# O que NÃO provoca fallback
|
||||
|
||||
Não executar fallback quando:
|
||||
|
||||
```text
|
||||
modelo retorna text
|
||||
modelo retorna media
|
||||
resultado parece improvável
|
||||
resultado parece ambíguo
|
||||
modelo parece estar inseguro
|
||||
outro modelo talvez respondesse diferente
|
||||
```
|
||||
|
||||
Não existe:
|
||||
|
||||
* score de confiança;
|
||||
* votação;
|
||||
* consenso;
|
||||
* LLM-as-a-judge;
|
||||
* segunda opinião;
|
||||
* comparação de respostas.
|
||||
|
||||
Uma resposta válida do provider atual encerra a cadeia.
|
||||
|
||||
---
|
||||
|
||||
# Execução sequencial
|
||||
|
||||
A cadeia deve ser executada sequencialmente.
|
||||
|
||||
```text
|
||||
try Ollama
|
||||
|
||||
se resposta válida:
|
||||
usar resposta
|
||||
finalizar
|
||||
|
||||
se falha operacional:
|
||||
try Groq
|
||||
|
||||
se resposta válida:
|
||||
usar resposta
|
||||
finalizar
|
||||
|
||||
se falha operacional:
|
||||
try OmniRoute
|
||||
```
|
||||
|
||||
Não executar providers em paralelo.
|
||||
|
||||
---
|
||||
|
||||
# Contrato do modelo
|
||||
|
||||
Todos os providers devem produzir o mesmo contrato lógico.
|
||||
|
||||
```json
|
||||
{
|
||||
"content_type": "text",
|
||||
"media_type": null
|
||||
}
|
||||
```
|
||||
|
||||
ou:
|
||||
|
||||
```json
|
||||
{
|
||||
"content_type": "media",
|
||||
"media_type": "video"
|
||||
}
|
||||
```
|
||||
|
||||
Valores permitidos:
|
||||
|
||||
```text
|
||||
content_type:
|
||||
- text
|
||||
- media
|
||||
```
|
||||
|
||||
```text
|
||||
media_type:
|
||||
- video
|
||||
- image
|
||||
- images
|
||||
- embed
|
||||
- mixed
|
||||
- null
|
||||
```
|
||||
|
||||
Regra:
|
||||
|
||||
```text
|
||||
text → media_type = null
|
||||
media → media_type obrigatório
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Structured output
|
||||
|
||||
Sempre que suportado pelo provider, a resposta deve utilizar JSON Schema/structured output nativo.
|
||||
|
||||
Independentemente do mecanismo do provider, a aplicação deve validar o resultado antes de aceitá-lo.
|
||||
|
||||
Não adicionar parsing tolerante complexo.
|
||||
|
||||
O contrato é pequeno e fechado.
|
||||
|
||||
Se a resposta não puder ser validada, aquela tentativa é considerada falha do provider e a cadeia avança.
|
||||
|
||||
---
|
||||
|
||||
# Prompt
|
||||
|
||||
O prompt deve ser único e pequeno.
|
||||
|
||||
Não criar prompts diferentes por idioma.
|
||||
|
||||
A instrução deve explicar apenas:
|
||||
|
||||
1. classificar `text` ou `media`;
|
||||
2. `media` significa que o texto é essencialmente uma introdução ou descrição da mídia;
|
||||
3. `text` significa que o artigo possui conteúdo jornalístico textual substancial;
|
||||
4. quando `media`, selecionar `video`, `image`, `images`, `embed` ou `mixed`;
|
||||
5. retornar exclusivamente o schema estabelecido.
|
||||
|
||||
Não solicitar:
|
||||
|
||||
* resumo;
|
||||
* justificativa;
|
||||
* reasoning;
|
||||
* confiança;
|
||||
* evidências;
|
||||
* tradução;
|
||||
* keywords.
|
||||
|
||||
---
|
||||
|
||||
# Configuração do modelo
|
||||
|
||||
A classificação deve utilizar comportamento determinístico sempre que o provider permitir.
|
||||
|
||||
Não habilitar reasoning desnecessário.
|
||||
|
||||
A configuração precisa privilegiar:
|
||||
|
||||
```text
|
||||
baixa variabilidade
|
||||
structured output
|
||||
baixa latência
|
||||
resposta curta
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Falha dos três providers
|
||||
|
||||
Se:
|
||||
|
||||
```text
|
||||
Ollama falhar
|
||||
Groq falhar
|
||||
OmniRoute falhar
|
||||
```
|
||||
|
||||
o sistema deve retornar internamente um estado de falha de classificação.
|
||||
|
||||
Exemplo lógico:
|
||||
|
||||
```json
|
||||
{
|
||||
"classification_status": "failed",
|
||||
"error_message": "media classification providers unavailable"
|
||||
}
|
||||
```
|
||||
|
||||
Não gerar automaticamente:
|
||||
|
||||
```text
|
||||
text
|
||||
```
|
||||
|
||||
Não gerar automaticamente:
|
||||
|
||||
```text
|
||||
media
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Destino da falha
|
||||
|
||||
O artigo com falha total:
|
||||
|
||||
* permanece registrado no JSON principal;
|
||||
* não passa pelo multimotor;
|
||||
* não entra no `*_media.json`;
|
||||
* não interrompe os outros artigos do lote.
|
||||
|
||||
Não criar:
|
||||
|
||||
```text
|
||||
*_failed.json
|
||||
```
|
||||
|
||||
Não criar:
|
||||
|
||||
* dead-letter queue;
|
||||
* banco para falhas;
|
||||
* worker de retry;
|
||||
* processo assíncrono de recuperação.
|
||||
|
||||
---
|
||||
|
||||
# Configuração e secrets
|
||||
|
||||
As informações específicas de cada provider devem ser externas ao código.
|
||||
|
||||
Incluem, quando aplicável:
|
||||
|
||||
```text
|
||||
endpoint
|
||||
model
|
||||
API key
|
||||
timeout
|
||||
```
|
||||
|
||||
Credenciais não podem aparecer:
|
||||
|
||||
* no código-fonte;
|
||||
* no JSON de saída;
|
||||
* nos logs;
|
||||
* nas mensagens de erro persistidas.
|
||||
|
||||
---
|
||||
|
||||
# Observabilidade
|
||||
|
||||
Para cada classificação deve ser possível identificar operacionalmente qual caminho foi utilizado:
|
||||
|
||||
```text
|
||||
ollama
|
||||
groq
|
||||
omniroute
|
||||
failed
|
||||
```
|
||||
|
||||
Devem existir logs para:
|
||||
|
||||
```text
|
||||
fallback Ollama → Groq
|
||||
fallback Groq → OmniRoute
|
||||
falha final
|
||||
```
|
||||
|
||||
O conteúdo integral do artigo não deve ser logado por padrão.
|
||||
|
||||
Não criar um novo sistema de observabilidade especificamente para esta feature.
|
||||
|
||||
---
|
||||
|
||||
# Testabilidade
|
||||
|
||||
A cadeia de providers deve ser testável sem depender dos serviços externos reais.
|
||||
|
||||
Os testes devem conseguir simular:
|
||||
|
||||
```text
|
||||
Ollama success
|
||||
Ollama fail + Groq success
|
||||
Ollama fail + Groq fail + OmniRoute success
|
||||
todos falham
|
||||
provider retorna schema inválido
|
||||
```
|
||||
|
||||
Não é necessário executar chamadas reais aos três providers na suíte normal de CI.
|
||||
|
||||
---
|
||||
|
||||
# Alternativas consideradas
|
||||
|
||||
## Apenas Ollama
|
||||
|
||||
Rejeitada.
|
||||
|
||||
O sistema é production-ready e precisa continuar funcionando caso o runtime local esteja indisponível.
|
||||
|
||||
---
|
||||
|
||||
## Ollama + Groq apenas
|
||||
|
||||
Rejeitada.
|
||||
|
||||
Foi definido um terceiro fallback já disponível através do OmniRoute.
|
||||
|
||||
---
|
||||
|
||||
## Chamar vários modelos e escolher maioria
|
||||
|
||||
Rejeitada.
|
||||
|
||||
Não existe requisito que justifique consenso.
|
||||
|
||||
Aumentaria:
|
||||
|
||||
* custo;
|
||||
* código;
|
||||
* latência;
|
||||
* pontos de falha.
|
||||
|
||||
---
|
||||
|
||||
## Fallback baseado em confidence
|
||||
|
||||
Rejeitada.
|
||||
|
||||
Não existe requisito de score de confiança e não deve existir interpretação adicional da resposta.
|
||||
|
||||
---
|
||||
|
||||
## Modelo grande como principal
|
||||
|
||||
Rejeitada.
|
||||
|
||||
A tarefa é pequena e fechada.
|
||||
|
||||
Qwen3.5 2B atende ao objetivo com custo operacional mínimo.
|
||||
|
||||
---
|
||||
|
||||
## Criar serviço separado de classificação
|
||||
|
||||
Rejeitada.
|
||||
|
||||
A classificação pertence ao fluxo atual e não exige um novo deployable.
|
||||
|
||||
---
|
||||
|
||||
# Consequências positivas
|
||||
|
||||
* custo marginal mínimo no caminho principal;
|
||||
* operação local normalmente;
|
||||
* alta disponibilidade através de dois fallbacks externos;
|
||||
* implementação pequena;
|
||||
* comportamento previsível;
|
||||
* nenhum acoplamento a framework de agentes;
|
||||
* fácil teste;
|
||||
* fácil troca futura de provider através da mesma interface lógica.
|
||||
|
||||
---
|
||||
|
||||
# Trade-off aceito
|
||||
|
||||
Em uma indisponibilidade simultânea de Ollama, Groq e OmniRoute, o artigo não será processado como texto nem mídia.
|
||||
|
||||
Essa é uma decisão intencional.
|
||||
|
||||
É preferível registrar explicitamente:
|
||||
|
||||
```text
|
||||
classification_failed
|
||||
```
|
||||
|
||||
a mascarar uma falha operacional tomando uma decisão que o sistema não conseguiu realizar.
|
||||
|
||||
---
|
||||
|
||||
# Invariantes
|
||||
|
||||
```text
|
||||
Ollama é sempre o primeiro provider.
|
||||
|
||||
Groq só é chamado após falha operacional do Ollama.
|
||||
|
||||
OmniRoute só é chamado após falha operacional de Ollama e Groq.
|
||||
|
||||
Resposta válida encerra a cadeia.
|
||||
|
||||
Nunca existe votação.
|
||||
|
||||
Nunca existe fallback por confiança.
|
||||
|
||||
Falha dos três nunca gera classificação presumida.
|
||||
```
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,516 +0,0 @@
|
||||
{
|
||||
"communities": {
|
||||
"0": [
|
||||
"specify_templates_tasks_template",
|
||||
"specify_templates_tasks_template_dependencies_execution_order",
|
||||
"specify_templates_tasks_template_format_id_p_story_description",
|
||||
"specify_templates_tasks_template_implementation_for_user_story_1",
|
||||
"specify_templates_tasks_template_implementation_for_user_story_2",
|
||||
"specify_templates_tasks_template_implementation_for_user_story_3",
|
||||
"specify_templates_tasks_template_implementation_strategy",
|
||||
"specify_templates_tasks_template_incremental_delivery",
|
||||
"specify_templates_tasks_template_mvp_first_user_story_1_only",
|
||||
"specify_templates_tasks_template_notes",
|
||||
"specify_templates_tasks_template_parallel_example_user_story_1",
|
||||
"specify_templates_tasks_template_parallel_opportunities",
|
||||
"specify_templates_tasks_template_parallel_team_strategy",
|
||||
"specify_templates_tasks_template_path_conventions",
|
||||
"specify_templates_tasks_template_phase_1_setup_shared_infrastructure",
|
||||
"specify_templates_tasks_template_phase_2_foundational_blocking_prerequisites",
|
||||
"specify_templates_tasks_template_phase_3_user_story_1_title_priority_p1_mvp",
|
||||
"specify_templates_tasks_template_phase_4_user_story_2_title_priority_p2",
|
||||
"specify_templates_tasks_template_phase_5_user_story_3_title_priority_p3",
|
||||
"specify_templates_tasks_template_phase_dependencies",
|
||||
"specify_templates_tasks_template_phase_n_polish_cross_cutting_concerns",
|
||||
"specify_templates_tasks_template_tasks_feature_name",
|
||||
"specify_templates_tasks_template_tests_for_user_story_1_optional_only_if_tests_requested",
|
||||
"specify_templates_tasks_template_tests_for_user_story_2_optional_only_if_tests_requested",
|
||||
"specify_templates_tasks_template_tests_for_user_story_3_optional_only_if_tests_requested",
|
||||
"specify_templates_tasks_template_user_story_dependencies",
|
||||
"specify_templates_tasks_template_within_each_user_story"
|
||||
],
|
||||
"1": [
|
||||
"agents_skills_speckit_converge_skill",
|
||||
"agents_skills_speckit_converge_skill_1_initialize_convergence_context",
|
||||
"agents_skills_speckit_converge_skill_2_load_artifacts_progressive_disclosure",
|
||||
"agents_skills_speckit_converge_skill_3_build_the_intent_inventory",
|
||||
"agents_skills_speckit_converge_skill_4_assess_the_codebase_and_classify_findings",
|
||||
"agents_skills_speckit_converge_skill_5_assign_severity",
|
||||
"agents_skills_speckit_converge_skill_6_present_the_in_session_findings_summary",
|
||||
"agents_skills_speckit_converge_skill_7_append_convergence_tasks_or_report_converged",
|
||||
"agents_skills_speckit_converge_skill_8_provide_next_actions_handoff",
|
||||
"agents_skills_speckit_converge_skill_9_check_for_extension_hooks",
|
||||
"agents_skills_speckit_converge_skill_convergence_findings",
|
||||
"agents_skills_speckit_converge_skill_execution_steps",
|
||||
"agents_skills_speckit_converge_skill_goal",
|
||||
"agents_skills_speckit_converge_skill_operating_constraints",
|
||||
"agents_skills_speckit_converge_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_converge_skill_user_input"
|
||||
],
|
||||
"2": [
|
||||
"specify_scripts_powershell_common",
|
||||
"specify_scripts_powershell_common_find_specifyroot",
|
||||
"specify_scripts_powershell_common_format_speckitcommand",
|
||||
"specify_scripts_powershell_common_get_currentbranch",
|
||||
"specify_scripts_powershell_common_get_featurepathsenv",
|
||||
"specify_scripts_powershell_common_get_invokeseparator",
|
||||
"specify_scripts_powershell_common_get_normalizedpriority",
|
||||
"specify_scripts_powershell_common_get_python3command",
|
||||
"specify_scripts_powershell_common_get_reporoot",
|
||||
"specify_scripts_powershell_common_get_sortedextensionids",
|
||||
"specify_scripts_powershell_common_resolve_specifyinitdir",
|
||||
"specify_scripts_powershell_common_resolve_template",
|
||||
"specify_scripts_powershell_common_resolve_templatecontent",
|
||||
"specify_scripts_powershell_common_save_featurejson",
|
||||
"specify_scripts_powershell_common_test_dirhasfiles",
|
||||
"specify_scripts_powershell_common_test_fileexists"
|
||||
],
|
||||
"3": [
|
||||
"agents_skills_graphify_skill",
|
||||
"agents_skills_graphify_skill_for_graphify_add_and_watch",
|
||||
"agents_skills_graphify_skill_for_graphify_query",
|
||||
"agents_skills_graphify_skill_for_the_commit_hook_and_native_claude_md_integration",
|
||||
"agents_skills_graphify_skill_for_update_and_cluster_only",
|
||||
"agents_skills_graphify_skill_graphify",
|
||||
"agents_skills_graphify_skill_honesty_rules",
|
||||
"agents_skills_graphify_skill_interpreter_guard_for_subcommands",
|
||||
"agents_skills_graphify_skill_part_a_structural_extraction_for_code_files",
|
||||
"agents_skills_graphify_skill_part_b_semantic_extraction_parallel_subagents",
|
||||
"agents_skills_graphify_skill_part_c_merge_ast_semantic_into_final_extraction",
|
||||
"agents_skills_graphify_skill_step_0_github_repos_and_multi_path_merge_only_if_a_url_or_several_paths",
|
||||
"agents_skills_graphify_skill_step_1_ensure_graphify_is_installed",
|
||||
"agents_skills_graphify_skill_step_2_5_video_and_audio_only_if_video_files_detected",
|
||||
"agents_skills_graphify_skill_step_2_detect_files",
|
||||
"agents_skills_graphify_skill_step_3_extract_entities_and_relationships",
|
||||
"agents_skills_graphify_skill_step_4_5_graph_health_check_read_only_integrity_gate",
|
||||
"agents_skills_graphify_skill_step_4_build_graph_cluster_analyze_generate_outputs",
|
||||
"agents_skills_graphify_skill_step_5_label_communities",
|
||||
"agents_skills_graphify_skill_step_6_generate_obsidian_vault_opt_in_html",
|
||||
"agents_skills_graphify_skill_step_9_save_manifest_update_cost_tracker_clean_up_and_report",
|
||||
"agents_skills_graphify_skill_steps_6b_8_wiki_neo4j_falkordb_svg_graphml_mcp_benchmark_only_on_their_flags",
|
||||
"agents_skills_graphify_skill_usage",
|
||||
"agents_skills_graphify_skill_what_graphify_is_for",
|
||||
"agents_skills_graphify_skill_what_you_must_do_when_invoked"
|
||||
],
|
||||
"4": [
|
||||
"agents_skills_speckit_analyze_skill",
|
||||
"agents_skills_speckit_analyze_skill_7_provide_next_actions",
|
||||
"agents_skills_speckit_analyze_skill_8_offer_remediation",
|
||||
"agents_skills_speckit_analyze_skill_9_check_for_extension_hooks",
|
||||
"agents_skills_speckit_analyze_skill_analysis_guidelines",
|
||||
"agents_skills_speckit_analyze_skill_context",
|
||||
"agents_skills_speckit_analyze_skill_context_efficiency",
|
||||
"agents_skills_speckit_analyze_skill_goal",
|
||||
"agents_skills_speckit_analyze_skill_operating_constraints",
|
||||
"agents_skills_speckit_analyze_skill_operating_principles",
|
||||
"agents_skills_speckit_analyze_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_analyze_skill_specification_analysis_report",
|
||||
"agents_skills_speckit_analyze_skill_user_input"
|
||||
],
|
||||
"5": [
|
||||
"agents_skills_speckit_analyze_skill_1_initialize_analysis_context",
|
||||
"agents_skills_speckit_analyze_skill_2_load_artifacts_progressive_disclosure",
|
||||
"agents_skills_speckit_analyze_skill_3_build_semantic_models",
|
||||
"agents_skills_speckit_analyze_skill_4_detection_passes_token_efficient_analysis",
|
||||
"agents_skills_speckit_analyze_skill_5_severity_assignment",
|
||||
"agents_skills_speckit_analyze_skill_6_produce_compact_analysis_report",
|
||||
"agents_skills_speckit_analyze_skill_a_duplication_detection",
|
||||
"agents_skills_speckit_analyze_skill_b_ambiguity_detection",
|
||||
"agents_skills_speckit_analyze_skill_c_underspecification",
|
||||
"agents_skills_speckit_analyze_skill_d_constitution_alignment",
|
||||
"agents_skills_speckit_analyze_skill_e_coverage_gaps",
|
||||
"agents_skills_speckit_analyze_skill_execution_steps",
|
||||
"agents_skills_speckit_analyze_skill_f_inconsistency"
|
||||
],
|
||||
"6": [
|
||||
"specify_templates_spec_template",
|
||||
"specify_templates_spec_template_assumptions",
|
||||
"specify_templates_spec_template_edge_cases",
|
||||
"specify_templates_spec_template_feature_specification_feature_name",
|
||||
"specify_templates_spec_template_functional_requirements",
|
||||
"specify_templates_spec_template_key_entities_include_if_feature_involves_data",
|
||||
"specify_templates_spec_template_measurable_outcomes",
|
||||
"specify_templates_spec_template_requirements_mandatory",
|
||||
"specify_templates_spec_template_success_criteria_mandatory",
|
||||
"specify_templates_spec_template_user_scenarios_testing_mandatory",
|
||||
"specify_templates_spec_template_user_story_1_brief_title_priority_p1",
|
||||
"specify_templates_spec_template_user_story_2_brief_title_priority_p2",
|
||||
"specify_templates_spec_template_user_story_3_brief_title_priority_p3"
|
||||
],
|
||||
"7": [
|
||||
"agents_rules_graphify",
|
||||
"agents_rules_graphify_graphify"
|
||||
],
|
||||
"8": [
|
||||
"agents_skills_speckit_plan_skill",
|
||||
"agents_skills_speckit_plan_skill_completion_report",
|
||||
"agents_skills_speckit_plan_skill_done_when",
|
||||
"agents_skills_speckit_plan_skill_key_rules",
|
||||
"agents_skills_speckit_plan_skill_mandatory_post_execution_hooks",
|
||||
"agents_skills_speckit_plan_skill_outline",
|
||||
"agents_skills_speckit_plan_skill_phase_0_outline_research",
|
||||
"agents_skills_speckit_plan_skill_phase_1_design_contracts",
|
||||
"agents_skills_speckit_plan_skill_phases",
|
||||
"agents_skills_speckit_plan_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_plan_skill_user_input"
|
||||
],
|
||||
"9": [
|
||||
"agents_skills_speckit_specify_skill",
|
||||
"agents_skills_speckit_specify_skill_completion_report",
|
||||
"agents_skills_speckit_specify_skill_done_when",
|
||||
"agents_skills_speckit_specify_skill_for_ai_generation",
|
||||
"agents_skills_speckit_specify_skill_mandatory_post_execution_hooks",
|
||||
"agents_skills_speckit_specify_skill_outline",
|
||||
"agents_skills_speckit_specify_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_specify_skill_quick_guidelines",
|
||||
"agents_skills_speckit_specify_skill_section_requirements",
|
||||
"agents_skills_speckit_specify_skill_success_criteria_guidelines",
|
||||
"agents_skills_speckit_specify_skill_user_input"
|
||||
],
|
||||
"10": [
|
||||
"agents_skills_speckit_tasks_skill",
|
||||
"agents_skills_speckit_tasks_skill_checklist_format_required",
|
||||
"agents_skills_speckit_tasks_skill_completion_report",
|
||||
"agents_skills_speckit_tasks_skill_done_when",
|
||||
"agents_skills_speckit_tasks_skill_mandatory_post_execution_hooks",
|
||||
"agents_skills_speckit_tasks_skill_outline",
|
||||
"agents_skills_speckit_tasks_skill_phase_structure",
|
||||
"agents_skills_speckit_tasks_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_tasks_skill_task_generation_rules",
|
||||
"agents_skills_speckit_tasks_skill_task_organization",
|
||||
"agents_skills_speckit_tasks_skill_user_input"
|
||||
],
|
||||
"11": [
|
||||
"specify_memory_constitution",
|
||||
"specify_memory_constitution_core_principles",
|
||||
"specify_memory_constitution_governance",
|
||||
"specify_memory_constitution_principle_1_name",
|
||||
"specify_memory_constitution_principle_2_name",
|
||||
"specify_memory_constitution_principle_3_name",
|
||||
"specify_memory_constitution_principle_4_name",
|
||||
"specify_memory_constitution_principle_5_name",
|
||||
"specify_memory_constitution_project_name_constitution",
|
||||
"specify_memory_constitution_section_2_name",
|
||||
"specify_memory_constitution_section_3_name"
|
||||
],
|
||||
"12": [
|
||||
"specify_templates_constitution_template",
|
||||
"specify_templates_constitution_template_core_principles",
|
||||
"specify_templates_constitution_template_governance",
|
||||
"specify_templates_constitution_template_principle_1_name",
|
||||
"specify_templates_constitution_template_principle_2_name",
|
||||
"specify_templates_constitution_template_principle_3_name",
|
||||
"specify_templates_constitution_template_principle_4_name",
|
||||
"specify_templates_constitution_template_principle_5_name",
|
||||
"specify_templates_constitution_template_project_name_constitution",
|
||||
"specify_templates_constitution_template_section_2_name",
|
||||
"specify_templates_constitution_template_section_3_name"
|
||||
],
|
||||
"13": [
|
||||
"agents_skills_graphify_references_exports",
|
||||
"agents_skills_graphify_references_exports_graphify_reference_extra_exports_and_benchmark",
|
||||
"agents_skills_graphify_references_exports_step_6b_wiki_only_if_wiki_flag",
|
||||
"agents_skills_graphify_references_exports_step_7_neo4j_export_only_if_neo4j_or_neo4j_push_flag",
|
||||
"agents_skills_graphify_references_exports_step_7a_falkordb_export_only_if_falkordb_or_falkordb_push_flag",
|
||||
"agents_skills_graphify_references_exports_step_7b_svg_export_only_if_svg_flag",
|
||||
"agents_skills_graphify_references_exports_step_7c_graphml_export_only_if_graphml_flag",
|
||||
"agents_skills_graphify_references_exports_step_7d_mcp_server_only_if_mcp_flag",
|
||||
"agents_skills_graphify_references_exports_step_8_token_reduction_benchmark_only_if_total_words_5000"
|
||||
],
|
||||
"14": [
|
||||
"agents_skills_ponytail_skill",
|
||||
"agents_skills_ponytail_skill_boundaries",
|
||||
"agents_skills_ponytail_skill_intensity",
|
||||
"agents_skills_ponytail_skill_output",
|
||||
"agents_skills_ponytail_skill_persistence",
|
||||
"agents_skills_ponytail_skill_ponytail",
|
||||
"agents_skills_ponytail_skill_rules",
|
||||
"agents_skills_ponytail_skill_the_ladder",
|
||||
"agents_skills_ponytail_skill_when_not_to_be_lazy"
|
||||
],
|
||||
"15": [
|
||||
"specify_templates_plan_template",
|
||||
"specify_templates_plan_template_complexity_tracking",
|
||||
"specify_templates_plan_template_constitution_check",
|
||||
"specify_templates_plan_template_documentation_this_feature",
|
||||
"specify_templates_plan_template_implementation_plan_feature",
|
||||
"specify_templates_plan_template_project_structure",
|
||||
"specify_templates_plan_template_source_code_repository_root",
|
||||
"specify_templates_plan_template_summary",
|
||||
"specify_templates_plan_template_technical_context"
|
||||
],
|
||||
"16": [
|
||||
"agents_skills_ponytail_help_skill",
|
||||
"agents_skills_ponytail_help_skill_configure_default_mode",
|
||||
"agents_skills_ponytail_help_skill_deactivate",
|
||||
"agents_skills_ponytail_help_skill_levels",
|
||||
"agents_skills_ponytail_help_skill_more",
|
||||
"agents_skills_ponytail_help_skill_ponytail_help",
|
||||
"agents_skills_ponytail_help_skill_skills",
|
||||
"agents_skills_ponytail_help_skill_update"
|
||||
],
|
||||
"17": [
|
||||
"agents_skills_speckit_checklist_skill",
|
||||
"agents_skills_speckit_checklist_skill_anti_examples_what_not_to_do",
|
||||
"agents_skills_speckit_checklist_skill_checklist_purpose_unit_tests_for_english",
|
||||
"agents_skills_speckit_checklist_skill_example_checklist_types_sample_items",
|
||||
"agents_skills_speckit_checklist_skill_execution_steps",
|
||||
"agents_skills_speckit_checklist_skill_post_execution_checks",
|
||||
"agents_skills_speckit_checklist_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_checklist_skill_user_input"
|
||||
],
|
||||
"18": [
|
||||
"agents_skills_speckit_clarify_skill",
|
||||
"agents_skills_speckit_clarify_skill_completion_report",
|
||||
"agents_skills_speckit_clarify_skill_done_when",
|
||||
"agents_skills_speckit_clarify_skill_mandatory_post_execution_hooks",
|
||||
"agents_skills_speckit_clarify_skill_outline",
|
||||
"agents_skills_speckit_clarify_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_clarify_skill_user_input"
|
||||
],
|
||||
"19": [
|
||||
"agents_skills_speckit_implement_skill",
|
||||
"agents_skills_speckit_implement_skill_completion_report",
|
||||
"agents_skills_speckit_implement_skill_done_when",
|
||||
"agents_skills_speckit_implement_skill_mandatory_post_execution_hooks",
|
||||
"agents_skills_speckit_implement_skill_outline",
|
||||
"agents_skills_speckit_implement_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_implement_skill_user_input"
|
||||
],
|
||||
"20": [
|
||||
"agents_skills_graphify_references_query",
|
||||
"agents_skills_graphify_references_query_for_graphify_explain",
|
||||
"agents_skills_graphify_references_query_for_graphify_path",
|
||||
"agents_skills_graphify_references_query_graphify_reference_query_path_explain",
|
||||
"agents_skills_graphify_references_query_step_0_constrained_query_expansion_required_before_traversal",
|
||||
"agents_skills_graphify_references_query_step_1_traversal"
|
||||
],
|
||||
"21": [
|
||||
"agents_skills_speckit_constitution_skill",
|
||||
"agents_skills_speckit_constitution_skill_outline",
|
||||
"agents_skills_speckit_constitution_skill_post_execution_checks",
|
||||
"agents_skills_speckit_constitution_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_constitution_skill_scope_guard",
|
||||
"agents_skills_speckit_constitution_skill_user_input"
|
||||
],
|
||||
"22": [
|
||||
"specify_scripts_powershell_create_new_feature",
|
||||
"specify_scripts_powershell_create_new_feature_convertto_cleanbranchname",
|
||||
"specify_scripts_powershell_create_new_feature_get_branchname",
|
||||
"specify_scripts_powershell_create_new_feature_get_fittedbranchname",
|
||||
"specify_scripts_powershell_create_new_feature_get_highestnumberfromspecs",
|
||||
"specify_scripts_powershell_create_new_feature_test_specprefixinuse"
|
||||
],
|
||||
"23": [
|
||||
"agents_skills_ponytail_audit_skill",
|
||||
"agents_skills_ponytail_audit_skill_boundaries",
|
||||
"agents_skills_ponytail_audit_skill_hunt",
|
||||
"agents_skills_ponytail_audit_skill_output",
|
||||
"agents_skills_ponytail_audit_skill_tags"
|
||||
],
|
||||
"24": [
|
||||
"agents_skills_ponytail_gain_skill",
|
||||
"agents_skills_ponytail_gain_skill_boundaries",
|
||||
"agents_skills_ponytail_gain_skill_honesty_boundary",
|
||||
"agents_skills_ponytail_gain_skill_ponytail_gain",
|
||||
"agents_skills_ponytail_gain_skill_scoreboard"
|
||||
],
|
||||
"25": [
|
||||
"agents_skills_ponytail_review_skill",
|
||||
"agents_skills_ponytail_review_skill_boundaries",
|
||||
"agents_skills_ponytail_review_skill_examples",
|
||||
"agents_skills_ponytail_review_skill_format",
|
||||
"agents_skills_ponytail_review_skill_scoring"
|
||||
],
|
||||
"26": [
|
||||
"agents_skills_speckit_taskstoissues_skill",
|
||||
"agents_skills_speckit_taskstoissues_skill_outline",
|
||||
"agents_skills_speckit_taskstoissues_skill_post_execution_checks",
|
||||
"agents_skills_speckit_taskstoissues_skill_pre_execution_checks",
|
||||
"agents_skills_speckit_taskstoissues_skill_user_input"
|
||||
],
|
||||
"27": [
|
||||
"specify_templates_checklist_template",
|
||||
"specify_templates_checklist_template_category_1",
|
||||
"specify_templates_checklist_template_category_2",
|
||||
"specify_templates_checklist_template_checklist_type_checklist_feature_name",
|
||||
"specify_templates_checklist_template_notes"
|
||||
],
|
||||
"28": [
|
||||
"agents_skills_graphify_references_add_watch",
|
||||
"agents_skills_graphify_references_add_watch_for_graphify_add",
|
||||
"agents_skills_graphify_references_add_watch_for_watch",
|
||||
"agents_skills_graphify_references_add_watch_graphify_reference_add_a_url_and_watch_a_folder"
|
||||
],
|
||||
"29": [
|
||||
"agents_skills_graphify_references_hooks",
|
||||
"agents_skills_graphify_references_hooks_for_git_commit_hook",
|
||||
"agents_skills_graphify_references_hooks_for_native_claude_md_integration",
|
||||
"agents_skills_graphify_references_hooks_graphify_reference_commit_hook_and_native_claude_md_integration"
|
||||
],
|
||||
"30": [
|
||||
"agents_skills_graphify_references_update",
|
||||
"agents_skills_graphify_references_update_for_cluster_only",
|
||||
"agents_skills_graphify_references_update_for_update_incremental_re_extraction",
|
||||
"agents_skills_graphify_references_update_graphify_reference_incremental_update_and_cluster_only"
|
||||
],
|
||||
"31": [
|
||||
"agents_skills_ponytail_debt_skill",
|
||||
"agents_skills_ponytail_debt_skill_boundaries",
|
||||
"agents_skills_ponytail_debt_skill_output",
|
||||
"agents_skills_ponytail_debt_skill_scan"
|
||||
],
|
||||
"32": [
|
||||
"agents_skills_graphify_references_github_and_merge",
|
||||
"agents_skills_graphify_references_github_and_merge_graphify_reference_github_clone_and_cross_repo_merge",
|
||||
"agents_skills_graphify_references_github_and_merge_step_0_clone_github_repo_s_only_if_a_github_url_was_given"
|
||||
],
|
||||
"33": [
|
||||
"agents_skills_graphify_references_transcribe",
|
||||
"agents_skills_graphify_references_transcribe_graphify_reference_transcribe_video_and_audio",
|
||||
"agents_skills_graphify_references_transcribe_step_2_5_transcribe_video_audio_files_only_if_video_files_detected"
|
||||
],
|
||||
"34": [
|
||||
"agents_skills_graphify_references_extraction_spec",
|
||||
"agents_skills_graphify_references_extraction_spec_graphify_reference_extraction_subagent_prompt"
|
||||
],
|
||||
"35": [
|
||||
"specify_scripts_powershell_check_prerequisites"
|
||||
],
|
||||
"36": [
|
||||
"specify_scripts_powershell_resolve_template"
|
||||
],
|
||||
"37": [
|
||||
"specify_scripts_powershell_setup_plan"
|
||||
],
|
||||
"38": [
|
||||
"specify_scripts_powershell_setup_tasks"
|
||||
],
|
||||
"39": [
|
||||
"agents_workflows_graphify",
|
||||
"agents_workflows_graphify_workflow_graphify"
|
||||
]
|
||||
},
|
||||
"cohesion": {
|
||||
"0": 0.07407407407407407,
|
||||
"1": 0.125,
|
||||
"2": 0.225,
|
||||
"3": 0.08,
|
||||
"4": 0.15384615384615385,
|
||||
"5": 0.15384615384615385,
|
||||
"6": 0.15384615384615385,
|
||||
"7": 1.0,
|
||||
"8": 0.18181818181818182,
|
||||
"9": 0.18181818181818182,
|
||||
"10": 0.18181818181818182,
|
||||
"11": 0.18181818181818182,
|
||||
"12": 0.18181818181818182,
|
||||
"13": 0.2222222222222222,
|
||||
"14": 0.2222222222222222,
|
||||
"15": 0.2222222222222222,
|
||||
"16": 0.25,
|
||||
"17": 0.25,
|
||||
"18": 0.2857142857142857,
|
||||
"19": 0.2857142857142857,
|
||||
"20": 0.3333333333333333,
|
||||
"21": 0.3333333333333333,
|
||||
"22": 0.4,
|
||||
"23": 0.4,
|
||||
"24": 0.4,
|
||||
"25": 0.4,
|
||||
"26": 0.4,
|
||||
"27": 0.4,
|
||||
"28": 0.5,
|
||||
"29": 0.5,
|
||||
"30": 0.5,
|
||||
"31": 0.5,
|
||||
"32": 0.6666666666666666,
|
||||
"33": 0.6666666666666666,
|
||||
"34": 1.0,
|
||||
"35": 1.0,
|
||||
"36": 1.0,
|
||||
"37": 1.0,
|
||||
"38": 1.0,
|
||||
"39": 1.0
|
||||
},
|
||||
"gods": [
|
||||
{
|
||||
"id": "specify_templates_tasks_template_tasks_feature_name",
|
||||
"label": "Tasks: [FEATURE NAME]",
|
||||
"degree": 13
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_graphify_skill_what_you_must_do_when_invoked",
|
||||
"label": "What You Must Do When Invoked",
|
||||
"degree": 12
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_graphify_skill_graphify",
|
||||
"label": "/graphify",
|
||||
"degree": 10
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_graphify_references_exports_graphify_reference_extra_exports_and_benchmark",
|
||||
"label": "graphify reference: extra exports and benchmark",
|
||||
"degree": 8
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_ponytail_skill_ponytail",
|
||||
"label": "Ponytail",
|
||||
"degree": 8
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_speckit_converge_skill_execution_steps",
|
||||
"label": "Execution Steps",
|
||||
"degree": 7
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_ponytail_help_skill_ponytail_help",
|
||||
"label": "Ponytail Help",
|
||||
"degree": 7
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_speckit_analyze_skill_4_detection_passes_token_efficient_analysis",
|
||||
"label": "4. Detection Passes (Token-Efficient Analysis)",
|
||||
"degree": 7
|
||||
},
|
||||
{
|
||||
"id": "agents_skills_speckit_analyze_skill_execution_steps",
|
||||
"label": "Execution Steps",
|
||||
"degree": 7
|
||||
},
|
||||
{
|
||||
"id": "specify_memory_constitution_core_principles",
|
||||
"label": "Core Principles",
|
||||
"degree": 6
|
||||
}
|
||||
],
|
||||
"surprises": [],
|
||||
"questions": [
|
||||
{
|
||||
"type": "bridge_node",
|
||||
"question": "Why does `Execution Steps` connect `Analysis Detection` to `Specification Analysis`?",
|
||||
"why": "High betweenness centrality (0.004) - this node is a cross-community bridge."
|
||||
},
|
||||
{
|
||||
"type": "isolated_nodes",
|
||||
"question": "What connects `Format: `[ID] [P?] [Story] Description``, `Implementation for User Story 1`, `Implementation for User Story 2` to the rest of the system?",
|
||||
"why": "212 weakly-connected nodes found - possible documentation gaps or missing edges."
|
||||
},
|
||||
{
|
||||
"type": "low_cohesion",
|
||||
"question": "Should `Task Planning` be split into smaller, more focused modules?",
|
||||
"why": "Cohesion score 0.07407407407407407 - nodes in this community are weakly interconnected."
|
||||
},
|
||||
{
|
||||
"type": "low_cohesion",
|
||||
"question": "Should `Convergence Workflow` be split into smaller, more focused modules?",
|
||||
"why": "Cohesion score 0.125 - nodes in this community are weakly interconnected."
|
||||
},
|
||||
{
|
||||
"type": "low_cohesion",
|
||||
"question": "Should `Graphify Commands` be split into smaller, more focused modules?",
|
||||
"why": "Cohesion score 0.08 - nodes in this community are weakly interconnected."
|
||||
}
|
||||
]
|
||||
}
|
||||
+551
-522
File diff suppressed because it is too large
Load Diff
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
||||
C:\Users\aferr\AppData\Local\Programs\Python\Python314\python.exe
|
||||
@@ -0,0 +1,176 @@
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\rules\graphify.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\assets\template.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\brand-marks\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\examples\dataflow-product-analytics.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\examples\lifecycle-agent-run.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\examples\sequence-cache-miss-request.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\examples\web-app-rendered.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\examples\workflow-agent-tool-call-rendered.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\references\authoring-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\references\brand-marks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\references\delivery-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\references\viewer-runtime.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\renderers\dataflow\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\renderers\lifecycle\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\renderers\sequence\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\renderers\workflow\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\archify\schemas\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\add-watch.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\exports.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\extraction-spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\github-and-merge.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\hooks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\query.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\transcribe.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\graphify\references\update.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail-audit\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail-debt\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail-gain\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail-help\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail-review\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\ponytail\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-analyze\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-checklist\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-clarify\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-constitution\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-converge\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-implement\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-plan\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-specify\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-tasks\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\skills\speckit-taskstoissues\SKILL.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.agents\workflows\graphify.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\memory\constitution.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\templates\checklist-template.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\templates\constitution-template.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\templates\plan-template.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\templates\spec-template.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\templates\tasks-template.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\.specify\workflows\speckit\workflow.yml
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.visual-check.html
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\googlenews_extractor_guia_completo.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\operations.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_convert_json_markdown.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_deterministic_content_selection.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_extrator_artigo_media\adr_001.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_extrator_artigo_media\adr_002.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_extrator_artigo_media\prd.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\prd_extrator_artigos_nlp.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\01_PRD_Runtime_Consolidacao_Artigos.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\02_Arquitetura_Runtime_Consolidacao_Artigos.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\03_ADRs_Runtime_Consolidacao_Artigos.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\04_Plano_Testes_Evals_Runtime.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\05_Metricas_KPIs_Runtime.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\06_Runbook_Producao_Runtime.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\structured_extraction\07_Especificacao_Prompt_Contexto_Harness_Runtime.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\evals\promptfoo.config.yaml
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\examples\content_northvolt_de.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\examples\content_presal_pt.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\examples\content_tangential_es.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\prompts\article_content_hygiene.v1.txt
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\prompts\article_sentiment_tags.v1.txt
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\requirements.txt
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\checklists\poc-readiness.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\contracts\cli-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\001-multilingual-entity-classifier\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\checklists\readiness.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\contracts\cli_contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\002-google-news-extractor\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\checklists\extraction.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\contracts\cli-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\contracts\json-schema.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\003-article-content-extractor\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\checklists\deterministic-selection.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\contracts\cli-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\contracts\json-schema.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\004-deterministic-content-selection\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\checklists\markdown-conversion.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\contracts\cli-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\contracts\markdown-schema.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\005-convert-json-markdown\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\checklists\release-readiness.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\contracts\cli-interface.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\contracts\prompts-contract.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\006-article-consolidation-runtime\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\checklists\media-routing.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\checklists\requirements.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\contracts\cli-interface.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\data-model.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\plan.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\quickstart.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\research.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\spec.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\specs\007-media-article-routing\tasks.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\README.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\de\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\de\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\de\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\de\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\en\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\en\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\en\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\en\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\es\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\es\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\es\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\es\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\fr\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\fr\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\fr\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\fr\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\it\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\it\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\it\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\it\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\pt\contextual.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\pt\direct.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\pt\not_related.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\benchmark_24\pt\tangential.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\markdown_conversion\valid_newspaper4k.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\markdown_conversion\valid_readability.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\fixtures\markdown_conversion\valid_trafilatura.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\tests\runtime\quality\RELEASE_QUALITY_REPORT.md
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.visual-check.1440x900.dark.png
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.visual-check.1440x900.light.png
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.visual-check.2048x1320.dark.png
|
||||
C:\Users\aferr\Projects\AFTech\DunaMedia\TextNLPClassifierApp\docs\archify\runtime-architecture.visual-check.2048x1320.light.png
|
||||
@@ -47,11 +47,11 @@
|
||||
"45": "ClassificationResult",
|
||||
"46": "ECPSnapshot",
|
||||
"47": "LLMFallbackAdapter",
|
||||
"48": "classifier.py",
|
||||
"48": "tools/models.py",
|
||||
"49": "content_northvolt_de.md",
|
||||
"50": "content_presal_pt.md",
|
||||
"51": "content_tangential_es.md",
|
||||
"52": "config.py",
|
||||
"52": "create_schema_registry",
|
||||
"53": "src/__init__.py",
|
||||
"54": "de/contextual.md",
|
||||
"55": "de/direct.md",
|
||||
@@ -98,7 +98,7 @@
|
||||
"96": "readiness.md",
|
||||
"97": "5. 📝 Conversor de Artigo JSON para Markdown",
|
||||
"98": "Extraction Pipeline Checklist: Article Content Multi-Engine Extractor",
|
||||
"99": "ModelGatewayClient",
|
||||
"99": "config.py",
|
||||
"100": "SQLiteStore",
|
||||
"101": "main",
|
||||
"102": "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)",
|
||||
@@ -120,7 +120,7 @@
|
||||
"118": "Tasks: Deterministic Article Content Selection",
|
||||
"119": "select_article_extractor",
|
||||
"120": "process_batch",
|
||||
"121": "detect_language",
|
||||
"121": "classifier.py",
|
||||
"122": "test_select_article_extractor.py",
|
||||
"123": "Feature Specification: Deterministic Content Selection",
|
||||
"124": "2. Entity Descriptions & Fields",
|
||||
@@ -147,14 +147,14 @@
|
||||
"145": "parse_arguments",
|
||||
"146": "Specification Quality Checklist: Convert Article JSON to Markdown",
|
||||
"147": "CLI Contract: `convert_article_to_markdown.py`",
|
||||
"148": "9. Interface CLI",
|
||||
"148": "architecture-delta.mjs",
|
||||
"149": "test_convert_article_to_markdown.py",
|
||||
"150": "13. Estratégia de testes",
|
||||
"150": "render-architecture.mjs",
|
||||
"151": "6. Contrato de entrada",
|
||||
"152": "test_models.py",
|
||||
"153": "convert_html_to_markdown",
|
||||
"154": "properties",
|
||||
"155": "5. Escopo",
|
||||
"155": "geometry.mjs",
|
||||
"156": "Los puntajes de River vs. Independiente Santa Fe, por la Copa Sudamericana - TyC Sports",
|
||||
"157": "valid_newspaper4k.md",
|
||||
"158": "valid_readability.md",
|
||||
@@ -162,7 +162,7 @@
|
||||
"160": "run_preflight_checks",
|
||||
"161": "adapter.py",
|
||||
"162": "validate_and_extract_enrichment",
|
||||
"163": "candidate/parser.py",
|
||||
"163": "render-lifecycle.mjs",
|
||||
"164": "hygiene/harness.py",
|
||||
"165": "SanitizedJsonLogger",
|
||||
"166": "InputSizeExceededError",
|
||||
@@ -176,12 +176,12 @@
|
||||
"174": "properties",
|
||||
"175": "repair-operations.schema.json",
|
||||
"176": "Functional Requirements",
|
||||
"177": "test_equivalence_mapping.py",
|
||||
"177": "properties",
|
||||
"178": "enrichment-response.schema.json",
|
||||
"179": "create_manifest_dict",
|
||||
"179": "properties",
|
||||
"180": "Catálogo de métricas e KPIs — Runtime de consolidação de artigos",
|
||||
"181": "article_content_hygiene",
|
||||
"182": "._get_connection",
|
||||
"182": "archify.mjs",
|
||||
"183": "Documento de Arquitetura — Runtime de consolidação de artigos",
|
||||
"184": "null",
|
||||
"185": "Runbook de produção — Runtime de consolidação de artigos",
|
||||
@@ -193,12 +193,12 @@
|
||||
"191": "properties",
|
||||
"192": "type",
|
||||
"193": "properties",
|
||||
"194": "build_minimal_hygiene_projection",
|
||||
"194": "brand-marks.mjs",
|
||||
"195": "check_zero_regex.py",
|
||||
"196": "properties",
|
||||
"197": "runtime-config.schema.json",
|
||||
"197": "webm-artifact.smoke.mjs",
|
||||
"198": "properties",
|
||||
"199": "object",
|
||||
"199": "provider_versions",
|
||||
"200": "candidates-payload.schema.json",
|
||||
"201": "limits",
|
||||
"202": "langfuse",
|
||||
@@ -211,13 +211,13 @@
|
||||
"209": "object",
|
||||
"210": "null",
|
||||
"211": "properties",
|
||||
"212": "calculate_execution_fingerprint",
|
||||
"213": "model_versions",
|
||||
"212": "check-render-output.mjs",
|
||||
"213": "render-sequence.mjs",
|
||||
"214": "3. Operational Execution Scenarios",
|
||||
"215": "type",
|
||||
"216": "test_cli_subprocess_pipeline.py",
|
||||
"217": "clean_body_images",
|
||||
"218": "ecp",
|
||||
"218": "properties",
|
||||
"219": "paths",
|
||||
"220": "properties",
|
||||
"221": "properties",
|
||||
@@ -275,7 +275,7 @@
|
||||
"273": "7. Checklist de release",
|
||||
"274": "eval_runner.py",
|
||||
"275": "ecp-profile.schema.json",
|
||||
"276": "prompt_versions",
|
||||
"276": "cli.mjs",
|
||||
"277": "12. Preparação determinística",
|
||||
"278": "15. Pequenos reparos textuais",
|
||||
"279": "9. Contrato do ECP",
|
||||
@@ -293,7 +293,10 @@
|
||||
"291": "build_release_metadata.py",
|
||||
"292": "removal_reasons",
|
||||
"293": "kept_image_ids",
|
||||
"294": "generated-validators.mjs",
|
||||
"295": "properties",
|
||||
"296": "test_prompts_contract.py",
|
||||
"297": "properties",
|
||||
"298": "Release Quality Summary Report",
|
||||
"299": "Comandos e Utilitários CLI",
|
||||
"300": "13. Candidatos e proveniência",
|
||||
@@ -326,6 +329,198 @@
|
||||
"327": "test_normalize_scalar_whitespace_collapsing",
|
||||
"328": "test_normalize_scalar_non_string_types",
|
||||
"329": "2. 📰 Extrator de Manchetes do Google News",
|
||||
"330": "properties",
|
||||
"331": "⚙️ Instalação e Setup",
|
||||
"332": "3. 📄 Extrator e Parser Multimotor de Artigos"
|
||||
"332": "3. 📄 Extrator e Parser Multimotor de Artigos",
|
||||
"333": "render-workflow.mjs",
|
||||
"334": "properties",
|
||||
"335": "scripts",
|
||||
"336": "render-dataflow.mjs",
|
||||
"337": "startPreview",
|
||||
"338": "required",
|
||||
"339": "layout-rules.test.mjs",
|
||||
"340": "visual-check.mjs",
|
||||
"341": "scenarios.mjs",
|
||||
"342": "Authoring contract",
|
||||
"343": "required",
|
||||
"344": "properties",
|
||||
"345": "properties",
|
||||
"346": "golden.mjs",
|
||||
"347": "output-path.mjs",
|
||||
"348": "properties",
|
||||
"349": "enum",
|
||||
"350": "adaptive-reader-layout.test.mjs",
|
||||
"351": "required",
|
||||
"352": "required",
|
||||
"353": "legend-contract.test.mjs",
|
||||
"354": "visual-check.test.mjs",
|
||||
"355": "engineering-profiles.mjs",
|
||||
"356": "required",
|
||||
"357": "properties",
|
||||
"358": "generate-brand-marks.mjs",
|
||||
"359": "repository",
|
||||
"360": "$defs",
|
||||
"361": "enum",
|
||||
"362": "dataflow.schema.json",
|
||||
"363": "properties",
|
||||
"364": "lifecycle.schema.json",
|
||||
"365": "workflow.schema.json",
|
||||
"366": "entries",
|
||||
"367": "sequence.schema.json",
|
||||
"368": "items",
|
||||
"369": "semantic-radar.test.mjs",
|
||||
"370": "architecture.schema.json",
|
||||
"371": "legendEntry",
|
||||
"372": "properties",
|
||||
"373": "degraded.test.mjs",
|
||||
"374": "viewer-chrome-layout.test.mjs",
|
||||
"375": "Viewer Runtime reference",
|
||||
"376": "repository-evidence.mjs",
|
||||
"377": "properties",
|
||||
"378": "properties",
|
||||
"379": "enum",
|
||||
"380": "entries",
|
||||
"381": "properties",
|
||||
"382": "properties",
|
||||
"383": "Archify",
|
||||
"384": "ordinary-model-floor.test.mjs",
|
||||
"385": "output-path.test.mjs",
|
||||
"386": "preview-contract.test.mjs",
|
||||
"387": "real-repository-proof.test.mjs",
|
||||
"388": "v1-compatibility.test.mjs",
|
||||
"389": "enum",
|
||||
"390": "Archify JSON IR Schemas",
|
||||
"391": "enum",
|
||||
"392": "enum",
|
||||
"393": "generate-validators.mjs",
|
||||
"394": "authored-reachability.test.mjs",
|
||||
"395": "chapter-delta-preview.test.mjs",
|
||||
"396": "cli.test.mjs",
|
||||
"397": "readme-showcase.test.mjs",
|
||||
"398": "repository-evidence.test.mjs",
|
||||
"399": "run_smoke_test",
|
||||
"400": "open-artifact.mjs",
|
||||
"401": "Delivery contract",
|
||||
"402": "enum",
|
||||
"403": "wraps",
|
||||
"404": "enum",
|
||||
"405": "enum",
|
||||
"406": "enum",
|
||||
"407": "required",
|
||||
"408": "mainPath",
|
||||
"409": "animation.test.mjs",
|
||||
"410": "guided-views.test.mjs",
|
||||
"411": "landing.test.mjs",
|
||||
"412": "preset-tryon.test.mjs",
|
||||
"413": "reach-share-card.test.mjs",
|
||||
"414": "relationship-direct-explorer.test.mjs",
|
||||
"415": "relationship-lens.test.mjs",
|
||||
"416": "relationship-permalink.test.mjs",
|
||||
"417": "release-identity.test.mjs",
|
||||
"418": "repair-receipt.test.mjs",
|
||||
"419": "route-journey.test.mjs",
|
||||
"420": "route-share-card.test.mjs",
|
||||
"421": "semantic-legend-gateway.test.mjs",
|
||||
"422": "sequence-column-fit.test.mjs",
|
||||
"423": "share-card-export.test.mjs",
|
||||
"424": "story-carrier.test.mjs",
|
||||
"425": "story-director-strip.test.mjs",
|
||||
"426": "story-follow-camera.test.mjs",
|
||||
"427": "story-moment-link.test.mjs",
|
||||
"428": "story-shelf.test.mjs",
|
||||
"429": "properties",
|
||||
"430": "PipeCdp",
|
||||
"431": "Sequence Renderer",
|
||||
"432": "enum",
|
||||
"433": "enum",
|
||||
"434": "required",
|
||||
"435": "enum",
|
||||
"436": "chapter-handoff.test.mjs",
|
||||
"437": "chapter-rail.test.mjs",
|
||||
"438": "diagram-guide.test.mjs",
|
||||
"439": "finder.test.mjs",
|
||||
"440": "intent-trace.test.mjs",
|
||||
"441": "motion-governor.test.mjs",
|
||||
"442": "presentation.test.mjs",
|
||||
"443": "preview.test.mjs",
|
||||
"444": "relationship-pulse.test.mjs",
|
||||
"445": "route-probe.test.mjs",
|
||||
"446": "semantic-camera.test.mjs",
|
||||
"447": "semantic-flow.test.mjs",
|
||||
"448": "semantic-lens.test.mjs",
|
||||
"449": "semantic-passport.test.mjs",
|
||||
"450": "semantic-zoom.test.mjs",
|
||||
"451": "story-beat-navigator.test.mjs",
|
||||
"452": "story-horizon.test.mjs",
|
||||
"453": "tags",
|
||||
"454": "test_benchmark_24.py",
|
||||
"455": "Data Flow Renderer",
|
||||
"456": "Lifecycle Renderer",
|
||||
"457": "Workflow Renderer",
|
||||
"458": "enum",
|
||||
"459": "enum",
|
||||
"460": "size",
|
||||
"461": "viewBox",
|
||||
"462": "point",
|
||||
"463": "enum",
|
||||
"464": "viewBox",
|
||||
"465": "automatic-port-spread.test.mjs",
|
||||
"466": "render-output-checks.test.mjs",
|
||||
"467": "skill-metadata.test.mjs",
|
||||
"468": "story-trail.test.mjs",
|
||||
"469": "prompts",
|
||||
"470": "meta",
|
||||
"471": "layout",
|
||||
"472": "enum",
|
||||
"473": "focus",
|
||||
"474": "enum",
|
||||
"475": "meta",
|
||||
"476": "render-examples.mjs",
|
||||
"477": "authoring-safety-contract.test.mjs",
|
||||
"478": "base-input-compatibility.test.mjs",
|
||||
"479": "release-package-gates.test.mjs",
|
||||
"480": "validate_certified_cheap_model",
|
||||
"481": "enum",
|
||||
"482": "enum",
|
||||
"483": "sources",
|
||||
"484": "via",
|
||||
"485": "col",
|
||||
"486": "via",
|
||||
"487": "enum",
|
||||
"488": "enum",
|
||||
"489": "bias",
|
||||
"490": "col",
|
||||
"491": "toCol",
|
||||
"492": "community-proof-intake.test.mjs",
|
||||
"493": "cursor-onboarding.test.mjs",
|
||||
"494": "delivery-contract.test.mjs",
|
||||
"495": "guide-page.test.mjs",
|
||||
"496": "proof-aperture.test.mjs",
|
||||
"497": "start-page.test.mjs",
|
||||
"498": "FaultyMockAdapter",
|
||||
"499": "MockProviderAdapter",
|
||||
"500": "Brand marks",
|
||||
"501": "col",
|
||||
"502": "label",
|
||||
"503": "row",
|
||||
"504": "width",
|
||||
"505": "height",
|
||||
"506": "label",
|
||||
"507": "row",
|
||||
"508": "stage",
|
||||
"509": "width",
|
||||
"510": "cornerRadius",
|
||||
"511": "height",
|
||||
"512": "label",
|
||||
"513": "width",
|
||||
"514": "to",
|
||||
"515": "y",
|
||||
"516": "height",
|
||||
"517": "label",
|
||||
"518": "labelSegment",
|
||||
"519": "width",
|
||||
"520": "generate-validators.test.mjs",
|
||||
"521": "toolbar-polish.test.mjs",
|
||||
"522": ".record_trace",
|
||||
"523": "brand-marks/README.md"
|
||||
}
|
||||
|
||||
+1068
-141
File diff suppressed because it is too large
Load Diff
+72307
-2588
File diff suppressed because it is too large
Load Diff
@@ -1558,5 +1558,857 @@
|
||||
"seen": 1787540618.9104896,
|
||||
"ast_hash": "c72c4699fec119d13d8631b33a62e834",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/bin/archify.mjs": {
|
||||
"mtime": 1787606942.2969265,
|
||||
"seen": 1787607908.2433958,
|
||||
"ast_hash": "adb84b0d00a1ef55e29e3e5d0254f3d3",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/bin/open-artifact.mjs": {
|
||||
"mtime": 1787606942.2969265,
|
||||
"seen": 1787607908.2434044,
|
||||
"ast_hash": "f917a4688be1e532a29f8dd3b07a97ca",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/bin/preview.mjs": {
|
||||
"mtime": 1787606942.297934,
|
||||
"seen": 1787607908.2434072,
|
||||
"ast_hash": "a91c7dbf1d48d3939aa1539c3b9897d0",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/bin/visual-check.mjs": {
|
||||
"mtime": 1787606942.298938,
|
||||
"seen": 1787607908.2434087,
|
||||
"ast_hash": "9beaff4026e194b2ddd24340d2e12430",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/delta/architecture-delta.mjs": {
|
||||
"mtime": 1787606942.3029325,
|
||||
"seen": 1787607908.2434106,
|
||||
"ast_hash": "0e7029313801f54f97192712b4ea68f6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/package.json": {
|
||||
"mtime": 1787606942.325334,
|
||||
"seen": 1787607908.2434123,
|
||||
"ast_hash": "b5cbe125269fe66ddaa6c9ea08871ea9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/recipes/scenarios.mjs": {
|
||||
"mtime": 1787606942.3272111,
|
||||
"seen": 1787607908.2434137,
|
||||
"ast_hash": "59d8762fcfded0193b4fba3bb8c1abb8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/architecture/grid.mjs": {
|
||||
"mtime": 1787606942.331944,
|
||||
"seen": 1787607908.2434154,
|
||||
"ast_hash": "8608b694f270c64bb826be7641b55dba",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/architecture/render-architecture.mjs": {
|
||||
"mtime": 1787606942.332677,
|
||||
"seen": 1787607908.2434165,
|
||||
"ast_hash": "e636e79a7d4a8208d9a9b808322af1a8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/dataflow/render-dataflow.mjs": {
|
||||
"mtime": 1787606942.3338504,
|
||||
"seen": 1787607908.2434182,
|
||||
"ast_hash": "18796b3fb856981c411a671e04215d62",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/lifecycle/render-lifecycle.mjs": {
|
||||
"mtime": 1787606942.3343878,
|
||||
"seen": 1787607908.2434194,
|
||||
"ast_hash": "613dc64b63ea62db9c756361eecc0e3e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/sequence/render-sequence.mjs": {
|
||||
"mtime": 1787606942.33672,
|
||||
"seen": 1787607908.2434208,
|
||||
"ast_hash": "2deb5fde836a9e77d7c82972f2fc91c5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/brand-marks.mjs": {
|
||||
"mtime": 1787606942.3377182,
|
||||
"seen": 1787607908.2434223,
|
||||
"ast_hash": "7bcd29083527e45b2525bae1089a1529",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/cli.mjs": {
|
||||
"mtime": 1787606942.3387196,
|
||||
"seen": 1787607908.2434235,
|
||||
"ast_hash": "848d94f0f0b03e63d34bacc7d2036057",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/desktop-readability.mjs": {
|
||||
"mtime": 1787606942.3397193,
|
||||
"seen": 1787607908.2434247,
|
||||
"ast_hash": "464744a89a7faadb9906782f9b127b84",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/diagnostics.mjs": {
|
||||
"mtime": 1787606942.3407164,
|
||||
"seen": 1787607908.2434266,
|
||||
"ast_hash": "c7fb67f17562ecd79ee5ce13a6aa84f3",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/engineering-profiles.mjs": {
|
||||
"mtime": 1787606942.3407164,
|
||||
"seen": 1787607908.2434278,
|
||||
"ast_hash": "8b0897e45f42590990772743dd52f3c2",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/generated-brand-marks.mjs": {
|
||||
"mtime": 1787606942.343029,
|
||||
"seen": 1787607908.2434292,
|
||||
"ast_hash": "d85c4353afd7fe56c56b519c5fe324bf",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/generated-validators.mjs": {
|
||||
"mtime": 1787606942.344828,
|
||||
"seen": 1787607908.2434306,
|
||||
"ast_hash": "fdcc3b3acfbd089c213eafd23eb6be61",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/geometry.mjs": {
|
||||
"mtime": 1787606942.345882,
|
||||
"seen": 1787607908.243432,
|
||||
"ast_hash": "565b767d2d7fcef0d688ad8a395d3d90",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/layout-report.mjs": {
|
||||
"mtime": 1787606942.3464508,
|
||||
"seen": 1787607908.2434332,
|
||||
"ast_hash": "53a14a9cb190a335f181def5c6c83a3e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/legend.mjs": {
|
||||
"mtime": 1787606942.3475463,
|
||||
"seen": 1787607908.2434347,
|
||||
"ast_hash": "b93529a288dfb476bff8aad2c2111ddd",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/output-path.mjs": {
|
||||
"mtime": 1787606942.3480775,
|
||||
"seen": 1787607908.2434363,
|
||||
"ast_hash": "d232558a3e323ff9f2a836454be8368f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/repository-evidence.mjs": {
|
||||
"mtime": 1787606942.3491492,
|
||||
"seen": 1787607908.2434375,
|
||||
"ast_hash": "d3c203059f0c5bf419158325e4a8a32a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/text-fit.mjs": {
|
||||
"mtime": 1787606942.349677,
|
||||
"seen": 1787607908.243439,
|
||||
"ast_hash": "4f790124cd10ed70d68b197fb6297422",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/utils.mjs": {
|
||||
"mtime": 1787606942.350733,
|
||||
"seen": 1787607908.2434402,
|
||||
"ast_hash": "979bad06c7333c8bfb748e689af347ca",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/shared/validator.mjs": {
|
||||
"mtime": 1787606942.3512383,
|
||||
"seen": 1787607908.2434413,
|
||||
"ast_hash": "704b4b5100701bfefea152237500d4b4",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/workflow/render-workflow.mjs": {
|
||||
"mtime": 1787606942.3531444,
|
||||
"seen": 1787607908.2434428,
|
||||
"ast_hash": "390ee422befab00ad1ba3f546a9eb8ba",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/architecture.schema.json": {
|
||||
"mtime": 1787606942.354152,
|
||||
"seen": 1787607908.2434442,
|
||||
"ast_hash": "f2afb7adcb6a01e88af0ef3ef370022e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/common.schema.json": {
|
||||
"mtime": 1787606942.354152,
|
||||
"seen": 1787607908.2434452,
|
||||
"ast_hash": "a3f28bbfdaae97623612dffb8216f518",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/dataflow.schema.json": {
|
||||
"mtime": 1787606942.3563848,
|
||||
"seen": 1787607908.2434466,
|
||||
"ast_hash": "d95c8e877118083bcb8f24d9815a9eca",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/lifecycle.schema.json": {
|
||||
"mtime": 1787606942.3563848,
|
||||
"seen": 1787607908.2434478,
|
||||
"ast_hash": "92890539e8aa5dd56d21de3b174bd83b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/sequence.schema.json": {
|
||||
"mtime": 1787606942.3563848,
|
||||
"seen": 1787607908.243449,
|
||||
"ast_hash": "67a9f6300075b7925d05e08645b83eb9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/workflow.schema.json": {
|
||||
"mtime": 1787606942.3578615,
|
||||
"seen": 1787607908.2434504,
|
||||
"ast_hash": "010a46752cf73298c2c2a24571da2478",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/scripts/check-render-output.mjs": {
|
||||
"mtime": 1787606942.3593535,
|
||||
"seen": 1787607908.2434516,
|
||||
"ast_hash": "9ef288fb90ed76e67dfd4ae6fcee143b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/scripts/generate-brand-marks.mjs": {
|
||||
"mtime": 1787606942.3604105,
|
||||
"seen": 1787607908.243453,
|
||||
"ast_hash": "080c5ba82d7a26761ca172e140fa1f9d",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/scripts/generate-validators.mjs": {
|
||||
"mtime": 1787606942.3614693,
|
||||
"seen": 1787607908.2434542,
|
||||
"ast_hash": "f42647ac8335d723e429cb7f3022e5e7",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/scripts/render-examples.mjs": {
|
||||
"mtime": 1787606942.3620217,
|
||||
"seen": 1787607908.2434554,
|
||||
"ast_hash": "be6da19091613bd2ddc3de727e6f2924",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/adaptive-reader-layout.test.mjs": {
|
||||
"mtime": 1787606942.363121,
|
||||
"seen": 1787607908.243457,
|
||||
"ast_hash": "ebb8b99ed2607aff54fff46279b96fd9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/animation.test.mjs": {
|
||||
"mtime": 1787606942.3641975,
|
||||
"seen": 1787607908.243458,
|
||||
"ast_hash": "747735a15730f686838ccb853baef926",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/architecture-delta.test.mjs": {
|
||||
"mtime": 1787606942.3658047,
|
||||
"seen": 1787607908.2434592,
|
||||
"ast_hash": "9fe5085753698b6c793e80913f4120f6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/authored-reachability.test.mjs": {
|
||||
"mtime": 1787606942.3664894,
|
||||
"seen": 1787607908.2434607,
|
||||
"ast_hash": "49934afde3d5538acf953dea86e94f20",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/authoring-safety-contract.test.mjs": {
|
||||
"mtime": 1787606942.367573,
|
||||
"seen": 1787607908.2434618,
|
||||
"ast_hash": "feeaff76de2f7e5389a3da990ce85fbb",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/automatic-port-spread.test.mjs": {
|
||||
"mtime": 1787606942.3686311,
|
||||
"seen": 1787607908.2434635,
|
||||
"ast_hash": "5ed5beb946f927407786ee9d486686c4",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/base-input-compatibility.test.mjs": {
|
||||
"mtime": 1787606942.3696902,
|
||||
"seen": 1787607908.2434647,
|
||||
"ast_hash": "3a8c087f5c1f22433d4a11313cdd09ae",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/brand-marks.test.mjs": {
|
||||
"mtime": 1787606942.3707423,
|
||||
"seen": 1787607908.243466,
|
||||
"ast_hash": "f7c8774b13a6e141a05812d011f29fd8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/chapter-delta-preview.test.mjs": {
|
||||
"mtime": 1787606942.3712752,
|
||||
"seen": 1787607908.2434676,
|
||||
"ast_hash": "e99e694a314fd46ce3fab40a3d42bb13",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/chapter-handoff.test.mjs": {
|
||||
"mtime": 1787606942.371803,
|
||||
"seen": 1787607908.2434688,
|
||||
"ast_hash": "9c5ca60e8e087743dca9b9da0a377e8e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/chapter-rail.test.mjs": {
|
||||
"mtime": 1787606942.3728843,
|
||||
"seen": 1787607908.2434702,
|
||||
"ast_hash": "667a6201e4a296f9a66c750a2fade2af",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/cli.test.mjs": {
|
||||
"mtime": 1787606942.3734157,
|
||||
"seen": 1787607908.2434714,
|
||||
"ast_hash": "515f0afe16db3950a238df07be2168bc",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/community-proof-intake.test.mjs": {
|
||||
"mtime": 1787606942.373988,
|
||||
"seen": 1787607908.2434723,
|
||||
"ast_hash": "b34d800299c2bc6b03b8a6f3fd98ccf8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/cursor-onboarding.test.mjs": {
|
||||
"mtime": 1787606942.374722,
|
||||
"seen": 1787607908.2434738,
|
||||
"ast_hash": "3c97170d99a3b69bd8f2a82b1e25e61c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/degraded.test.mjs": {
|
||||
"mtime": 1787606942.3752286,
|
||||
"seen": 1787607908.2434747,
|
||||
"ast_hash": "76c01eb14e60d96a5a30d16811ab987e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/delivery-contract.test.mjs": {
|
||||
"mtime": 1787606942.3757646,
|
||||
"seen": 1787607908.243476,
|
||||
"ast_hash": "65ab7435438710d9268d77b3c672c472",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/desktop-reader-browser.test.mjs": {
|
||||
"mtime": 1787606942.377384,
|
||||
"seen": 1787607908.2434773,
|
||||
"ast_hash": "3cdd9e64c896628d3337b976793f2686",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/diagram-guide.test.mjs": {
|
||||
"mtime": 1787606942.3779619,
|
||||
"seen": 1787607908.2434788,
|
||||
"ast_hash": "fad89a1efbe1ed09e683a5099859cbda",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/engineering-profile.test.mjs": {
|
||||
"mtime": 1787606942.3784935,
|
||||
"seen": 1787607908.24348,
|
||||
"ast_hash": "e50b20ec11d68b58c9b3f732a198841f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/finder.test.mjs": {
|
||||
"mtime": 1787606942.379552,
|
||||
"seen": 1787607908.2434814,
|
||||
"ast_hash": "e587766aac0b8f85eefa5b478cce6c7b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/gallery.test.mjs": {
|
||||
"mtime": 1787606942.3860364,
|
||||
"seen": 1787607908.2434826,
|
||||
"ast_hash": "710f854a8ed201b174996795f33ec647",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/generate-validators.test.mjs": {
|
||||
"mtime": 1787606942.387102,
|
||||
"seen": 1787607908.243484,
|
||||
"ast_hash": "45b632e8f8af8cfe958d38304715b06e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/geometry.test.mjs": {
|
||||
"mtime": 1787606942.3876314,
|
||||
"seen": 1787607908.2434852,
|
||||
"ast_hash": "80ac4e8b3960443d1a70f99f6ad4f5c6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/golden.mjs": {
|
||||
"mtime": 1787606942.3886912,
|
||||
"seen": 1787607908.2434864,
|
||||
"ast_hash": "1e9a9d6df182b0173682289b2f809249",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/grid.test.mjs": {
|
||||
"mtime": 1787606942.3892365,
|
||||
"seen": 1787607908.2434876,
|
||||
"ast_hash": "04346292f4965f824a798ca4a2af7e61",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/guide-page.test.mjs": {
|
||||
"mtime": 1787606942.390305,
|
||||
"seen": 1787607908.2434888,
|
||||
"ast_hash": "5f7030d7c4056f8eafd8901e737f1343",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/guide.test.mjs": {
|
||||
"mtime": 1787606942.3908389,
|
||||
"seen": 1787607908.24349,
|
||||
"ast_hash": "047595c5a23e8eb11fb6217d9198d4e8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/guided-views.test.mjs": {
|
||||
"mtime": 1787606942.3913875,
|
||||
"seen": 1787607908.2434914,
|
||||
"ast_hash": "f28cec8e9da25942da206aaa6c7789bc",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/intent-trace.test.mjs": {
|
||||
"mtime": 1787606942.3919382,
|
||||
"seen": 1787607908.2434926,
|
||||
"ast_hash": "465a91c118045cfd5bf05fabbb37aa68",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/landing.test.mjs": {
|
||||
"mtime": 1787606942.3924935,
|
||||
"seen": 1787607908.2434938,
|
||||
"ast_hash": "e5e389e6325df730ba9f6128482b6b2b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/layout-rules.test.mjs": {
|
||||
"mtime": 1787606942.3933132,
|
||||
"seen": 1787607908.243495,
|
||||
"ast_hash": "45ea36c39ce1c3c90ccbb019a75beb49",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/legend-contract.test.mjs": {
|
||||
"mtime": 1787606942.3938203,
|
||||
"seen": 1787607908.2434962,
|
||||
"ast_hash": "9931c1eeae88ac81d03fc110a3c48628",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/motion-governor.test.mjs": {
|
||||
"mtime": 1787606942.3943624,
|
||||
"seen": 1787607908.2434974,
|
||||
"ast_hash": "c973636650a348b096c264ade2a77aab",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/open-artifact.test.mjs": {
|
||||
"mtime": 1787606942.3954344,
|
||||
"seen": 1787607908.243499,
|
||||
"ast_hash": "1588189d9827309aa392af67a7c7c9d9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/ordinary-model-floor.test.mjs": {
|
||||
"mtime": 1787606942.3959727,
|
||||
"seen": 1787607908.2435002,
|
||||
"ast_hash": "4e333a7d23d11f633f58696319271940",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/output-path.test.mjs": {
|
||||
"mtime": 1787606942.396576,
|
||||
"seen": 1787607908.2435014,
|
||||
"ast_hash": "0615302748576f43fffc39fce8306156",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/presentation.test.mjs": {
|
||||
"mtime": 1787606942.3971064,
|
||||
"seen": 1787607908.2435026,
|
||||
"ast_hash": "8b3981d43872cd5cd2015f1d0f33c8d9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/preset-tryon.test.mjs": {
|
||||
"mtime": 1787606942.3981752,
|
||||
"seen": 1787607908.2435038,
|
||||
"ast_hash": "0b874a06f74863fe3fd9f77374b2ddda",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/preview-contract.test.mjs": {
|
||||
"mtime": 1787606942.3981752,
|
||||
"seen": 1787607908.243505,
|
||||
"ast_hash": "653ca90156e33f1e79d0dc6445f8a372",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/preview.test.mjs": {
|
||||
"mtime": 1787606942.3992372,
|
||||
"seen": 1787607908.2435062,
|
||||
"ast_hash": "5d05016c4fe14b2154a0bc5e7ee00bb1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/proof-aperture.test.mjs": {
|
||||
"mtime": 1787606942.3997834,
|
||||
"seen": 1787607908.2435074,
|
||||
"ast_hash": "518fa80e8ad7d5a666c93a59a1611524",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/reach-share-card.test.mjs": {
|
||||
"mtime": 1787606942.4007993,
|
||||
"seen": 1787607908.243509,
|
||||
"ast_hash": "007616bdf6b8ffa3e158e756f374a3d1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/readme-showcase.test.mjs": {
|
||||
"mtime": 1787606942.4013522,
|
||||
"seen": 1787607908.2435102,
|
||||
"ast_hash": "89b8ebcfb33a025ef336b48a3e768149",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/real-repository-proof.test.mjs": {
|
||||
"mtime": 1787606942.4018612,
|
||||
"seen": 1787607908.2435112,
|
||||
"ast_hash": "67707e3e75b60a3aafe7332bc1f74a06",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/relationship-direct-explorer.test.mjs": {
|
||||
"mtime": 1787606942.4018612,
|
||||
"seen": 1787607908.2435174,
|
||||
"ast_hash": "818d83d811e596f432a6907fe96c29f3",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/relationship-lens.test.mjs": {
|
||||
"mtime": 1787606942.4028716,
|
||||
"seen": 1787607908.2435188,
|
||||
"ast_hash": "2da60eb342310c66805d6f8735212ad7",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/relationship-permalink.test.mjs": {
|
||||
"mtime": 1787606942.40387,
|
||||
"seen": 1787607908.2435198,
|
||||
"ast_hash": "0658033a7facd164fc3b3eb29b349f9a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/relationship-pulse.test.mjs": {
|
||||
"mtime": 1787606942.4044955,
|
||||
"seen": 1787607908.243521,
|
||||
"ast_hash": "168c71ed85f488c36f5d3db00940d068",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/release-identity.test.mjs": {
|
||||
"mtime": 1787606942.4055023,
|
||||
"seen": 1787607908.2435224,
|
||||
"ast_hash": "f2bdc17b706d3cdc19b170ff31940899",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/release-package-gates.test.mjs": {
|
||||
"mtime": 1787606942.4055023,
|
||||
"seen": 1787607908.2435234,
|
||||
"ast_hash": "42e748cf64570eb5c6a41bd97579d108",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/render-output-checks.test.mjs": {
|
||||
"mtime": 1787606942.4075072,
|
||||
"seen": 1787607908.2435246,
|
||||
"ast_hash": "60f0878d277a884adcd96b08133b044b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/repair-receipt.test.mjs": {
|
||||
"mtime": 1787606942.4085078,
|
||||
"seen": 1787607908.243526,
|
||||
"ast_hash": "491f18ccd34d0ce08a9595b518fc8ae3",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/repository-evidence.test.mjs": {
|
||||
"mtime": 1787606942.4085078,
|
||||
"seen": 1787607908.243527,
|
||||
"ast_hash": "a1feed896e93b1b58f58f29c2902693f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/route-journey.test.mjs": {
|
||||
"mtime": 1787606942.4095092,
|
||||
"seen": 1787607908.2435281,
|
||||
"ast_hash": "6876a505dae10e24c6e1ead24cbf9c6e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/route-probe.test.mjs": {
|
||||
"mtime": 1787606942.4099574,
|
||||
"seen": 1787607908.2435293,
|
||||
"ast_hash": "0f62523ce03a3e815faf454faa9c6422",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/route-share-card.test.mjs": {
|
||||
"mtime": 1787606942.4104638,
|
||||
"seen": 1787607908.2435305,
|
||||
"ast_hash": "c999696caa2aab78650bbd839e004521",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-camera.test.mjs": {
|
||||
"mtime": 1787606942.411,
|
||||
"seen": 1787607908.2435317,
|
||||
"ast_hash": "678e6f11f4552adfcc2cf67946bd7fa1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-flow.test.mjs": {
|
||||
"mtime": 1787606942.4115305,
|
||||
"seen": 1787607908.2435331,
|
||||
"ast_hash": "9c98473f0699f3225222e3defd956ad1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-legend-gateway.test.mjs": {
|
||||
"mtime": 1787606942.4120717,
|
||||
"seen": 1787607908.243534,
|
||||
"ast_hash": "6d675a83d6bcd59fe48594305b98240a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-lens.test.mjs": {
|
||||
"mtime": 1787606942.4125986,
|
||||
"seen": 1787607908.2435353,
|
||||
"ast_hash": "cac20cd1e15235981b56505bd3822df6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-passport.test.mjs": {
|
||||
"mtime": 1787606942.4125986,
|
||||
"seen": 1787607908.2435365,
|
||||
"ast_hash": "29d154bf104aea8e7a4719e512dc7282",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-radar.test.mjs": {
|
||||
"mtime": 1787606942.4131448,
|
||||
"seen": 1787607908.2435377,
|
||||
"ast_hash": "4d9f5de537495bbdb73584c687283cc0",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/semantic-zoom.test.mjs": {
|
||||
"mtime": 1787606942.4136765,
|
||||
"seen": 1787607908.2435389,
|
||||
"ast_hash": "763fc2eca549de6c487c488b74f701e5",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/sequence-column-fit.test.mjs": {
|
||||
"mtime": 1787606942.4147384,
|
||||
"seen": 1787607908.2435403,
|
||||
"ast_hash": "01a182d12aa90f3523c6a0e6ce545969",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/settled-flow.test.mjs": {
|
||||
"mtime": 1787606942.4152703,
|
||||
"seen": 1787607908.2435415,
|
||||
"ast_hash": "e79ae3c7d78601b327dffe4f01da7029",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/share-card-export.test.mjs": {
|
||||
"mtime": 1787606942.4163425,
|
||||
"seen": 1787607908.2435424,
|
||||
"ast_hash": "1433ae30bf8bfbb538d2440bdf01ff66",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/skill-metadata.test.mjs": {
|
||||
"mtime": 1787606942.4174168,
|
||||
"seen": 1787607908.2435439,
|
||||
"ast_hash": "2ed4ba5a0a84f8e3e23a5f0b85144356",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/start-page.test.mjs": {
|
||||
"mtime": 1787606942.4179623,
|
||||
"seen": 1787607908.243545,
|
||||
"ast_hash": "a85b05e509de9af7bae60ab0859f7062",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-beat-navigator.test.mjs": {
|
||||
"mtime": 1787606942.4190865,
|
||||
"seen": 1787607908.2435465,
|
||||
"ast_hash": "4b593542dcaa6c5235e3d777c0e33c42",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-carrier.test.mjs": {
|
||||
"mtime": 1787606942.4202332,
|
||||
"seen": 1787607908.2435474,
|
||||
"ast_hash": "5cc1914c4aaf503aacedf3ed12ee6b65",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-director-strip.test.mjs": {
|
||||
"mtime": 1787606942.4213042,
|
||||
"seen": 1787607908.2435486,
|
||||
"ast_hash": "9401178acff38473aec77bf15e206ec8",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-follow-camera.test.mjs": {
|
||||
"mtime": 1787606942.4218316,
|
||||
"seen": 1787607908.24355,
|
||||
"ast_hash": "3255aa1935fd2c295461b2ea2677d215",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-horizon.test.mjs": {
|
||||
"mtime": 1787606942.4228404,
|
||||
"seen": 1787607908.243551,
|
||||
"ast_hash": "024d49efc233c605d5ce0ed43012f762",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-moment-link.test.mjs": {
|
||||
"mtime": 1787606942.423841,
|
||||
"seen": 1787607908.2435522,
|
||||
"ast_hash": "a447e66af10059e4da933fad6429fe04",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-shelf.test.mjs": {
|
||||
"mtime": 1787606942.424844,
|
||||
"seen": 1787607908.2435536,
|
||||
"ast_hash": "0d6899eb50b7f4dd4aa7df3a02131934",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/story-trail.test.mjs": {
|
||||
"mtime": 1787606942.424844,
|
||||
"seen": 1787607908.2435548,
|
||||
"ast_hash": "f2d6fc08e1c10b84fefe45532776dba1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/toolbar-polish.test.mjs": {
|
||||
"mtime": 1787606942.425842,
|
||||
"seen": 1787607908.2435558,
|
||||
"ast_hash": "e1a08f6966b4b4ff7e55374c80d7df6f",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/v1-compatibility.test.mjs": {
|
||||
"mtime": 1787606942.425842,
|
||||
"seen": 1787607908.2435572,
|
||||
"ast_hash": "029acb631f05c01acf4fb2b0091271f1",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/viewer-chrome-layout.test.mjs": {
|
||||
"mtime": 1787606942.4268415,
|
||||
"seen": 1787607908.2435584,
|
||||
"ast_hash": "cf1c079d918f342d385a007710ca4db6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/visual-check.test.mjs": {
|
||||
"mtime": 1787606942.4268415,
|
||||
"seen": 1787607908.2435596,
|
||||
"ast_hash": "2651bdac2a04e8fd70868113b1433d93",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/test/webm-artifact.smoke.mjs": {
|
||||
"mtime": 1787606942.4278405,
|
||||
"seen": 1787607908.2435608,
|
||||
"ast_hash": "7fbaeaafc736403bd8bfd859be41ca38",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/SKILL.md": {
|
||||
"mtime": 1787606942.2916005,
|
||||
"seen": 1787607908.2573073,
|
||||
"ast_hash": "1018c7b1cb10dcaabe1e5fdce31d218e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/assets/template.html": {
|
||||
"mtime": 1787606942.2958257,
|
||||
"seen": 1787607908.2573097,
|
||||
"ast_hash": "a52e0179abd86de2409521e6efdaa18a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/brand-marks/README.md": {
|
||||
"mtime": 1787606942.2999356,
|
||||
"seen": 1787607908.2573109,
|
||||
"ast_hash": "36458d477094c8c0c10ea70b333c004b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/examples/dataflow-product-analytics.html": {
|
||||
"mtime": 1787606942.3084197,
|
||||
"seen": 1787607908.2573125,
|
||||
"ast_hash": "03a1731f1ba4cb45441dd89ab8656cb6",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/examples/lifecycle-agent-run.html": {
|
||||
"mtime": 1787606942.313279,
|
||||
"seen": 1787607908.2573137,
|
||||
"ast_hash": "2768be1baffeff2baed25ea71513ad4a",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/examples/sequence-cache-miss-request.html": {
|
||||
"mtime": 1787606942.3183756,
|
||||
"seen": 1787607908.257315,
|
||||
"ast_hash": "079a511ff5c927c5fddf587d4ded374c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/examples/web-app-rendered.html": {
|
||||
"mtime": 1787606942.3210454,
|
||||
"seen": 1787607908.2573159,
|
||||
"ast_hash": "c1c572489a5e404f4ec34ee4311e7360",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/examples/workflow-agent-tool-call-rendered.html": {
|
||||
"mtime": 1787606942.324779,
|
||||
"seen": 1787607908.2573175,
|
||||
"ast_hash": "d636ee39c05aec820db2f020962b6bdb",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/references/authoring-contract.md": {
|
||||
"mtime": 1787606942.3283193,
|
||||
"seen": 1787607908.2573195,
|
||||
"ast_hash": "06f6d464fd1a769404b454fbd3359647",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/references/brand-marks.md": {
|
||||
"mtime": 1787606942.3288805,
|
||||
"seen": 1787607908.2573202,
|
||||
"ast_hash": "b19005698777cfd581a85406780f6f66",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/references/delivery-contract.md": {
|
||||
"mtime": 1787606942.3306758,
|
||||
"seen": 1787607908.2573214,
|
||||
"ast_hash": "7760ddc41e0e46d9bbc15c6b2cc353c9",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/references/viewer-runtime.md": {
|
||||
"mtime": 1787606942.3312194,
|
||||
"seen": 1787607908.257323,
|
||||
"ast_hash": "99abe43365f07355f654f3ed9aa25fed",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/dataflow/README.md": {
|
||||
"mtime": 1787606942.3332138,
|
||||
"seen": 1787607908.257324,
|
||||
"ast_hash": "16d7de5e3d9696d4dc2ac30bc7dd5d36",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/lifecycle/README.md": {
|
||||
"mtime": 1787606942.3343878,
|
||||
"seen": 1787607908.2573254,
|
||||
"ast_hash": "1a50984ee7fb7cb0930dae9a100d1455",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/sequence/README.md": {
|
||||
"mtime": 1787606942.335704,
|
||||
"seen": 1787607908.2573264,
|
||||
"ast_hash": "bf61961e5e909b7be786a849567d2a45",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/renderers/workflow/README.md": {
|
||||
"mtime": 1787606942.3522985,
|
||||
"seen": 1787607908.2573276,
|
||||
"ast_hash": "5795cfe8e7b65b0b6bafde738314f442",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
".agents/skills/archify/schemas/README.md": {
|
||||
"mtime": 1787606942.354152,
|
||||
"seen": 1787607908.2573287,
|
||||
"ast_hash": "b7310d4e7328a5318dc828f1f9ecf08d",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.html": {
|
||||
"mtime": 1787607879.6105008,
|
||||
"seen": 1787607908.2639914,
|
||||
"ast_hash": "77265db6c6c0ed4afbedb20646eb3fef",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.visual-check.html": {
|
||||
"mtime": 1787607886.0328603,
|
||||
"seen": 1787607908.2639933,
|
||||
"ast_hash": "dba72bf0620fee9de7a08880c6a9536b",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.visual-check.1440x900.dark.png": {
|
||||
"mtime": 1787607885.5348077,
|
||||
"seen": 1787607908.280495,
|
||||
"ast_hash": "e174761d0adb29d2421dc8f9cb53320e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.visual-check.1440x900.light.png": {
|
||||
"mtime": 1787607883.7483766,
|
||||
"seen": 1787607908.280497,
|
||||
"ast_hash": "693d0db4d6a5a70c96c298d820b3c86c",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.visual-check.2048x1320.dark.png": {
|
||||
"mtime": 1787607886.0308325,
|
||||
"seen": 1787607908.2804987,
|
||||
"ast_hash": "d221aa64e5f00f935095a1b1ea66cb7e",
|
||||
"semantic_hash": ""
|
||||
},
|
||||
"docs/runtime-architecture.visual-check.2048x1320.light.png": {
|
||||
"mtime": 1787607885.0050726,
|
||||
"seen": 1787607908.2805004,
|
||||
"ast_hash": "58ea31b656d9efb0285bd41e5b636929",
|
||||
"semantic_hash": ""
|
||||
}
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"runs": [
|
||||
{
|
||||
"date": "2026-08-25T04:09:39.635900+00:00",
|
||||
"input_tokens": 0,
|
||||
"output_tokens": 0,
|
||||
"files": 521
|
||||
}
|
||||
],
|
||||
"total_input_tokens": 0,
|
||||
"total_output_tokens": 0
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
+2243
-2107
File diff suppressed because it is too large
Load Diff
Vendored
+1
-1
File diff suppressed because one or more lines are too long
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"runs": [
|
||||
{
|
||||
"date": "2026-08-25T04:09:39.635900+00:00",
|
||||
"input_tokens": 0,
|
||||
"output_tokens": 0,
|
||||
"files": 521
|
||||
}
|
||||
],
|
||||
"total_input_tokens": 0,
|
||||
"total_output_tokens": 0
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
+72589
-61786
File diff suppressed because it is too large
Load Diff
+1218
-504
File diff suppressed because it is too large
Load Diff
@@ -12,8 +12,11 @@ from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
@@ -30,6 +33,377 @@ from readability import Document
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MediaCandidateInfo:
|
||||
"""Informações estruturais da DOM sobre mídias candidatas identificadas."""
|
||||
|
||||
has_candidate_media: bool
|
||||
has_video: bool = False
|
||||
image_count: int = 0
|
||||
has_embed: bool = False
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MediaClassification:
|
||||
"""Classificação estruturada emitida pelo classificador semântico."""
|
||||
|
||||
content_type: Literal["text", "media"]
|
||||
media_type: Literal["video", "image", "images", "embed", "mixed"] | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"content_type": self.content_type,
|
||||
"media_type": self.media_type,
|
||||
}
|
||||
|
||||
|
||||
MEDIA_CLASSIFIER_SCHEMA: dict[str, Any] = {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"content_type": {
|
||||
"type": "string",
|
||||
"enum": ["text", "media"],
|
||||
"description": "Classification: 'text' for substantive journalistic text, 'media' for predominantly media.",
|
||||
},
|
||||
"media_type": {
|
||||
"type": ["string", "null"],
|
||||
"enum": ["video", "image", "images", "embed", "mixed", None],
|
||||
"description": "Specific media category when content_type is 'media', or null when content_type is 'text'.",
|
||||
},
|
||||
},
|
||||
"required": ["content_type", "media_type"],
|
||||
"additionalProperties": False,
|
||||
}
|
||||
|
||||
|
||||
def validate_classifier_response(data: Any) -> MediaClassification | None:
|
||||
"""Valida estritamente o contrato de 2 campos da resposta do classificador."""
|
||||
if not isinstance(data, dict):
|
||||
return None
|
||||
if set(data.keys()) != {"content_type", "media_type"}:
|
||||
return None
|
||||
|
||||
content_type = data.get("content_type")
|
||||
media_type = data.get("media_type")
|
||||
|
||||
if content_type not in ("text", "media"):
|
||||
return None
|
||||
|
||||
if content_type == "text":
|
||||
if media_type is not None:
|
||||
return None
|
||||
return MediaClassification(content_type="text", media_type=None)
|
||||
|
||||
# content_type == "media"
|
||||
if media_type not in ("video", "image", "images", "embed", "mixed"):
|
||||
return None
|
||||
return MediaClassification(content_type="media", media_type=media_type)
|
||||
|
||||
|
||||
def _http_post_json(
|
||||
url: str,
|
||||
payload: dict[str, Any],
|
||||
headers: dict[str, str],
|
||||
timeout: int,
|
||||
) -> tuple[int, str]:
|
||||
"""Helper de baixo nível para envio de requisições POST JSON via urllib.request."""
|
||||
data_bytes = json.dumps(payload).encode("utf-8")
|
||||
req = urllib.request.Request(url, data=data_bytes, headers=headers, method="POST")
|
||||
with urllib.request.urlopen(req, timeout=timeout) as response:
|
||||
status = getattr(response, "status", response.getcode())
|
||||
body = response.read().decode("utf-8")
|
||||
return status, body
|
||||
|
||||
|
||||
MEDIA_CLASSIFIER_PROMPT: str = (
|
||||
"You are an editorial news classifier. Classify if this news publication is predominantly media or substantive journalistic text.\n\n"
|
||||
"Publication Title: {title}\n"
|
||||
"Structural Media Present: Video={has_video}, ImagesCount={image_count}, Embed={has_embed}\n"
|
||||
"Text Content:\n"
|
||||
"{text_content}\n\n"
|
||||
"Definitions:\n"
|
||||
"- \"media\": The primary informative content is in the media (video, single image, multiple images/gallery, social embed, or mixed), and the text functions essentially as a brief introduction, caption, contextualization, or description.\n"
|
||||
"- \"text\": The publication contains substantive journalistic text on its own, even if accompanied by illustrative media.\n\n"
|
||||
"Respond ONLY with a JSON object matching this exact schema:\n"
|
||||
"{{\"content_type\": \"text\" | \"media\", \"media_type\": \"video\" | \"image\" | \"images\" | \"embed\" | \"mixed\" | null}}\n"
|
||||
"Rules:\n"
|
||||
"- If content_type is \"text\", media_type MUST be null.\n"
|
||||
"- If content_type is \"media\", media_type MUST be one of: \"video\", \"image\", \"images\", \"embed\", \"mixed\"."
|
||||
)
|
||||
|
||||
|
||||
def _find_editorial_region(soup: BeautifulSoup) -> Any:
|
||||
"""Localiza a região editorial da DOM respeitando a ordem de precedência."""
|
||||
article = soup.find("article")
|
||||
if article:
|
||||
return article
|
||||
main = soup.find("main")
|
||||
if main:
|
||||
return main
|
||||
role_main = soup.find("div", attrs={"role": "main"})
|
||||
if role_main:
|
||||
return role_main
|
||||
if soup.body:
|
||||
return soup.body
|
||||
return soup
|
||||
|
||||
|
||||
def detect_candidate_media(soup: BeautifulSoup) -> MediaCandidateInfo:
|
||||
"""
|
||||
Analisa estruturalmente a DOM carregada para identificar elementos candidatos a mídia.
|
||||
Executa exclusivamente via navegação DOM (Zero-Regex).
|
||||
"""
|
||||
region = _find_editorial_region(soup)
|
||||
if not region:
|
||||
return MediaCandidateInfo(has_candidate_media=False)
|
||||
|
||||
# Identifica vídeos: tags <video>
|
||||
videos = region.find_all("video")
|
||||
valid_videos = 0
|
||||
for v in videos:
|
||||
if v.find_parent(["header", "nav", "footer", "aside"]):
|
||||
continue
|
||||
valid_videos += 1
|
||||
has_video = valid_videos > 0
|
||||
|
||||
# Contagem de imagens reais: tags <img>
|
||||
images = region.find_all("img")
|
||||
valid_images = 0
|
||||
for img in images:
|
||||
if img.find_parent(["header", "nav", "footer", "aside"]):
|
||||
continue
|
||||
valid_images += 1
|
||||
|
||||
# Identifica embeds: <iframe>, <embed>, <object>
|
||||
embed_tags = region.find_all(["iframe", "embed", "object"])
|
||||
valid_embeds = 0
|
||||
for emb in embed_tags:
|
||||
if emb.find_parent(["header", "nav", "footer", "aside"]):
|
||||
continue
|
||||
valid_embeds += 1
|
||||
has_embed = valid_embeds > 0
|
||||
|
||||
has_candidate_media = has_video or (valid_images > 0) or has_embed
|
||||
|
||||
return MediaCandidateInfo(
|
||||
has_candidate_media=has_candidate_media,
|
||||
has_video=has_video,
|
||||
image_count=valid_images,
|
||||
has_embed=has_embed,
|
||||
)
|
||||
|
||||
|
||||
def build_compact_payload(soup: BeautifulSoup, candidate_info: MediaCandidateInfo) -> str:
|
||||
"""Monta o payload compacto sem marcações HTML para envio ao classificador semântico."""
|
||||
region = _find_editorial_region(soup)
|
||||
|
||||
title = ""
|
||||
if soup.title and soup.title.string:
|
||||
title = soup.title.string.strip()
|
||||
elif region:
|
||||
h1 = region.find("h1")
|
||||
if h1:
|
||||
title = h1.get_text(strip=True)
|
||||
if not title and soup.find("h1"):
|
||||
h1 = soup.find("h1")
|
||||
if h1:
|
||||
title = h1.get_text(strip=True)
|
||||
|
||||
paragraphs: list[str] = []
|
||||
if region:
|
||||
for p in region.find_all("p"):
|
||||
if p.find_parent(["header", "nav", "footer", "aside"]):
|
||||
continue
|
||||
text = " ".join(p.get_text().split())
|
||||
if text:
|
||||
paragraphs.append(text)
|
||||
|
||||
text_content = "\n\n".join(paragraphs)
|
||||
|
||||
return MEDIA_CLASSIFIER_PROMPT.format(
|
||||
title=title,
|
||||
has_video=candidate_info.has_video,
|
||||
image_count=candidate_info.image_count,
|
||||
has_embed=candidate_info.has_embed,
|
||||
text_content=text_content,
|
||||
)
|
||||
|
||||
|
||||
def classify_media_content(
|
||||
payload: str,
|
||||
metrics: dict[str, int],
|
||||
silent: bool = False,
|
||||
) -> tuple[MediaClassification | None, str | None]:
|
||||
"""
|
||||
Classifica a publicação usando a cadeia sequencial de provedores LLM:
|
||||
Ollama (qwen3.5:2b) -> Groq (openai/gpt-oss-20b) -> OmniRoute (cgpt-web/gpt-5.5).
|
||||
Retorna (MediaClassification, None) na primeira resposta válida ou (None, error_message) em caso de falha cumulativa.
|
||||
"""
|
||||
# --------------------------------------------------------------------------
|
||||
# 1. Provedor Primário: Ollama
|
||||
# --------------------------------------------------------------------------
|
||||
ollama_endpoint = os.environ.get("OLLAMA_ENDPOINT", "http://localhost:11434").rstrip("/")
|
||||
ollama_model = os.environ.get("OLLAMA_MODEL", "qwen3.5:2b")
|
||||
ollama_timeout = int(os.environ.get("OLLAMA_TIMEOUT", "10"))
|
||||
|
||||
ollama_url = f"{ollama_endpoint}/api/chat"
|
||||
ollama_payload = {
|
||||
"model": ollama_model,
|
||||
"messages": [{"role": "user", "content": payload}],
|
||||
"stream": False,
|
||||
"format": MEDIA_CLASSIFIER_SCHEMA,
|
||||
"options": {
|
||||
"temperature": 0.0,
|
||||
},
|
||||
"think": False,
|
||||
}
|
||||
|
||||
try:
|
||||
status, body = _http_post_json(
|
||||
ollama_url,
|
||||
ollama_payload,
|
||||
{"Content-Type": "application/json"},
|
||||
ollama_timeout,
|
||||
)
|
||||
if status == 200:
|
||||
parsed = json.loads(body)
|
||||
content_str = parsed.get("message", {}).get("content", "")
|
||||
data = json.loads(content_str) if isinstance(content_str, str) else content_str
|
||||
classification = validate_classifier_response(data)
|
||||
if classification is not None:
|
||||
if not silent:
|
||||
m_label = f" ({classification.media_type})" if classification.media_type else ""
|
||||
sys.stderr.write(
|
||||
f"[MEDIA] Provedor: Ollama ({ollama_model}) | Classificação: {classification.content_type}{m_label}\n"
|
||||
)
|
||||
sys.stderr.flush()
|
||||
return classification, None
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# 2. Primeiro Fallback: Groq
|
||||
# --------------------------------------------------------------------------
|
||||
metrics["fallback_groq"] += 1
|
||||
if not silent:
|
||||
sys.stderr.write("[MEDIA] Acionando fallback 1: Groq\n")
|
||||
sys.stderr.flush()
|
||||
|
||||
groq_endpoint = os.environ.get("GROQ_ENDPOINT", "https://api.groq.com/openai/v1/chat/completions")
|
||||
groq_api_key = os.environ.get("GROQ_API_KEY", "")
|
||||
groq_model = os.environ.get("GROQ_MODEL", "openai/gpt-oss-20b")
|
||||
groq_timeout = int(os.environ.get("GROQ_TIMEOUT", "15"))
|
||||
|
||||
if groq_endpoint and groq_api_key:
|
||||
groq_payload = {
|
||||
"model": groq_model,
|
||||
"messages": [{"role": "user", "content": payload}],
|
||||
"temperature": 0.0,
|
||||
"reasoning_effort": "low",
|
||||
"response_format": {
|
||||
"type": "json_schema",
|
||||
"json_schema": {
|
||||
"name": "media_classifier",
|
||||
"strict": True,
|
||||
"schema": MEDIA_CLASSIFIER_SCHEMA,
|
||||
},
|
||||
},
|
||||
}
|
||||
groq_headers = {
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {groq_api_key}",
|
||||
}
|
||||
try:
|
||||
status, body = _http_post_json(groq_endpoint, groq_payload, groq_headers, groq_timeout)
|
||||
if status == 200:
|
||||
parsed = json.loads(body)
|
||||
choices = parsed.get("choices", [])
|
||||
if choices:
|
||||
content_str = choices[0].get("message", {}).get("content", "")
|
||||
data = json.loads(content_str) if isinstance(content_str, str) else content_str
|
||||
classification = validate_classifier_response(data)
|
||||
if classification is not None:
|
||||
if not silent:
|
||||
m_label = f" ({classification.media_type})" if classification.media_type else ""
|
||||
sys.stderr.write(
|
||||
f"[MEDIA] Provedor: Groq ({groq_model}) | Classificação: {classification.content_type}{m_label}\n"
|
||||
)
|
||||
sys.stderr.flush()
|
||||
return classification, None
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# 3. Segundo Fallback: OmniRoute
|
||||
# --------------------------------------------------------------------------
|
||||
metrics["fallback_omniroute"] += 1
|
||||
if not silent:
|
||||
sys.stderr.write("[MEDIA] Acionando fallback 2: OmniRoute\n")
|
||||
sys.stderr.flush()
|
||||
|
||||
omniroute_endpoint = os.environ.get("OMNIROUTE_ENDPOINT", "")
|
||||
omniroute_api_key = os.environ.get("OMNIROUTE_API_KEY", "")
|
||||
omniroute_model = os.environ.get("OMNIROUTE_MODEL", "cgpt-web/gpt-5.5")
|
||||
omniroute_timeout = int(os.environ.get("OMNIROUTE_TIMEOUT", "20"))
|
||||
|
||||
if omniroute_endpoint:
|
||||
omniroute_payload = {
|
||||
"model": omniroute_model,
|
||||
"messages": [{"role": "user", "content": payload}],
|
||||
"temperature": 0.0,
|
||||
"response_format": {
|
||||
"type": "json_schema",
|
||||
"json_schema": {
|
||||
"name": "media_classifier",
|
||||
"strict": True,
|
||||
"schema": MEDIA_CLASSIFIER_SCHEMA,
|
||||
},
|
||||
},
|
||||
}
|
||||
omniroute_headers = {
|
||||
"Content-Type": "application/json",
|
||||
}
|
||||
if omniroute_api_key:
|
||||
omniroute_headers["Authorization"] = f"Bearer {omniroute_api_key}"
|
||||
|
||||
try:
|
||||
status, body = _http_post_json(
|
||||
omniroute_endpoint, omniroute_payload, omniroute_headers, omniroute_timeout
|
||||
)
|
||||
if status == 200:
|
||||
parsed = json.loads(body)
|
||||
choices = parsed.get("choices", [])
|
||||
if choices:
|
||||
content_str = choices[0].get("message", {}).get("content", "")
|
||||
data = json.loads(content_str) if isinstance(content_str, str) else content_str
|
||||
classification = validate_classifier_response(data)
|
||||
if classification is not None:
|
||||
if not silent:
|
||||
m_label = f" ({classification.media_type})" if classification.media_type else ""
|
||||
sys.stderr.write(
|
||||
f"[MEDIA] Provedor: OmniRoute ({omniroute_model}) | Classificação: {classification.content_type}{m_label}\n"
|
||||
)
|
||||
sys.stderr.flush()
|
||||
return classification, None
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Falha Total
|
||||
# --------------------------------------------------------------------------
|
||||
error_msg = "All classification providers failed (Ollama, Groq, OmniRoute)."
|
||||
return None, error_msg
|
||||
|
||||
|
||||
def save_media_json(articles: list[dict[str, Any]], output_path: Path) -> None:
|
||||
"""Salva o arquivo de mídia dedicado com envelope mínimo."""
|
||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
payload = {"articles": articles}
|
||||
with output_path.open("w", encoding="utf-8") as f:
|
||||
json.dump(payload, f, ensure_ascii=False, indent=2)
|
||||
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class InputArticle:
|
||||
"""Metadados originais da notícia contida no JSON de entrada."""
|
||||
@@ -39,15 +413,19 @@ class InputArticle:
|
||||
subtitulo: str | None = None
|
||||
quando_publicado: str | None = None
|
||||
pagina: int = 1
|
||||
raw_data: dict[str, Any] = field(default_factory=dict)
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
"titulo": self.titulo,
|
||||
"subtitulo": self.subtitulo,
|
||||
"quando_publicado": self.quando_publicado,
|
||||
"url": self.url,
|
||||
"pagina": self.pagina,
|
||||
}
|
||||
result = dict(self.raw_data)
|
||||
result["titulo"] = self.titulo
|
||||
result["url"] = self.url
|
||||
if self.subtitulo is not None and "subtitulo" not in result:
|
||||
result["subtitulo"] = self.subtitulo
|
||||
if self.quando_publicado is not None and "quando_publicado" not in result:
|
||||
result["quando_publicado"] = self.quando_publicado
|
||||
if "pagina" not in result:
|
||||
result["pagina"] = self.pagina
|
||||
return result
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -175,27 +553,33 @@ class ExtractedArticle:
|
||||
"""Resultado consolidado da extração de um artigo."""
|
||||
|
||||
input_meta: InputArticle
|
||||
extraction_status: Literal["success", "failed"]
|
||||
error_message: str | None
|
||||
crawled_url: str
|
||||
page_title: str | None
|
||||
http_status: int | None
|
||||
extraction_status: Literal["success", "failed"] | None = None
|
||||
classification_status: Literal["failed"] | None = None
|
||||
error_message: str | None = None
|
||||
crawled_url: str = ""
|
||||
page_title: str | None = None
|
||||
http_status: int | None = None
|
||||
trafilatura: TrafilaturaData | None = None
|
||||
newspaper4k: NewspaperData | None = None
|
||||
readability: ReadabilityData | None = None
|
||||
|
||||
def to_dict(self) -> dict[str, Any]:
|
||||
return {
|
||||
res: dict[str, Any] = {
|
||||
"input_meta": self.input_meta.to_dict(),
|
||||
"extraction_status": self.extraction_status,
|
||||
"error_message": self.error_message,
|
||||
"crawled_url": self.crawled_url,
|
||||
"page_title": self.page_title,
|
||||
"http_status": self.http_status,
|
||||
"trafilatura": self.trafilatura.to_dict() if self.trafilatura else None,
|
||||
"newspaper4k": self.newspaper4k.to_dict() if self.newspaper4k else None,
|
||||
"readability": self.readability.to_dict() if self.readability else None,
|
||||
}
|
||||
if self.classification_status is not None:
|
||||
res["classification_status"] = self.classification_status
|
||||
if self.extraction_status is not None:
|
||||
res["extraction_status"] = self.extraction_status
|
||||
res["error_message"] = self.error_message
|
||||
res["crawled_url"] = self.crawled_url
|
||||
res["page_title"] = self.page_title
|
||||
res["http_status"] = self.http_status
|
||||
if self.classification_status is None:
|
||||
res["trafilatura"] = self.trafilatura.to_dict() if self.trafilatura else None
|
||||
res["newspaper4k"] = self.newspaper4k.to_dict() if self.newspaper4k else None
|
||||
res["readability"] = self.readability.to_dict() if self.readability else None
|
||||
return res
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -538,7 +922,7 @@ def log_info(message: str, silent: bool = False) -> None:
|
||||
|
||||
|
||||
def load_search_json(file_path: Path) -> tuple[str | None, str, list[InputArticle]]:
|
||||
"""Carrega o arquivo JSON gerado pelo extrator de notícias."""
|
||||
"""Carrega o arquivo JSON gerado pelo extrator de notícias preservando 100% dos metadados."""
|
||||
if not file_path.exists():
|
||||
raise FileNotFoundError(f"Arquivo de entrada não encontrado: {file_path}")
|
||||
|
||||
@@ -556,9 +940,7 @@ def load_search_json(file_path: Path) -> tuple[str | None, str, list[InputArticl
|
||||
InputArticle(
|
||||
titulo=item["titulo"],
|
||||
url=item["url"],
|
||||
subtitulo=item.get("subtitulo"),
|
||||
quando_publicado=item.get("quando_publicado"),
|
||||
pagina=item.get("pagina", 1),
|
||||
raw_data=dict(item),
|
||||
)
|
||||
)
|
||||
|
||||
@@ -595,11 +977,31 @@ def process_batch(
|
||||
silent=silent,
|
||||
)
|
||||
|
||||
# Determinar caminho de saída padrão se não especificado
|
||||
# Determinar caminho de saída textual e caminho do arquivo de mídia
|
||||
if output_path is None:
|
||||
output_path = input_path.parent / f"{input_path.stem}_extracted.json"
|
||||
text_output_path = input_path.parent / f"{input_path.stem}_extracted.json"
|
||||
media_output_path = input_path.parent / f"{input_path.stem}_media.json"
|
||||
else:
|
||||
text_output_path = output_path
|
||||
media_output_path = output_path.with_name(f"{output_path.stem}_media{output_path.suffix}")
|
||||
|
||||
# Inicializar contadores operacionais das 11 métricas
|
||||
metrics: dict[str, int] = {
|
||||
"total_evaluated": 0,
|
||||
"text": 0,
|
||||
"media": 0,
|
||||
"media/video": 0,
|
||||
"media/image": 0,
|
||||
"media/images": 0,
|
||||
"media/embed": 0,
|
||||
"media/mixed": 0,
|
||||
"fallback_groq": 0,
|
||||
"fallback_omniroute": 0,
|
||||
"classification_failed": 0,
|
||||
}
|
||||
|
||||
extracted_list: list[ExtractedArticle] = []
|
||||
media_articles: list[dict[str, Any]] = []
|
||||
successful_count = 0
|
||||
failed_count = 0
|
||||
start_time = time.time()
|
||||
@@ -611,6 +1013,68 @@ def process_batch(
|
||||
|
||||
try:
|
||||
html, page_title, http_status = crawler.crawl(url)
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
|
||||
metrics["total_evaluated"] += 1
|
||||
candidate_info = detect_candidate_media(soup)
|
||||
|
||||
# Gate estrutural prévio: se houver mídia candidata, envia ao classificador
|
||||
if candidate_info.has_candidate_media:
|
||||
payload = build_compact_payload(soup, candidate_info)
|
||||
classification, error_msg = classify_media_content(
|
||||
payload, metrics, silent=silent
|
||||
)
|
||||
|
||||
if classification is not None and classification.content_type == "media":
|
||||
metrics["media"] += 1
|
||||
if classification.media_type:
|
||||
m_key = f"media/{classification.media_type}"
|
||||
if m_key in metrics:
|
||||
metrics[m_key] += 1
|
||||
|
||||
media_article = {
|
||||
"input_meta": article.to_dict(),
|
||||
"crawled_url": url,
|
||||
"page_title": page_title,
|
||||
"http_status": http_status,
|
||||
"content_type": "media",
|
||||
"media_type": classification.media_type,
|
||||
}
|
||||
media_articles.append(media_article)
|
||||
log_info(
|
||||
f'📹 [{idx}/{total}] Publicação predominantemente de mídia ({classification.media_type}) desviada para *_media.json',
|
||||
silent=silent,
|
||||
)
|
||||
continue
|
||||
elif classification is not None and classification.content_type == "text":
|
||||
metrics["text"] += 1
|
||||
elif classification is None:
|
||||
# Falha total na cadeia de provedores
|
||||
metrics["classification_failed"] += 1
|
||||
failed_count += 1
|
||||
log_info(
|
||||
f"⚠️ [{idx}/{total}] Falha de classificação para URL '{url}': {error_msg}",
|
||||
silent=silent,
|
||||
)
|
||||
if not silent:
|
||||
sys.stderr.write(f"[MEDIA] Falha total da cadeia de classificação: {error_msg}\n")
|
||||
sys.stderr.flush()
|
||||
failed_article = ExtractedArticle(
|
||||
input_meta=article,
|
||||
classification_status="failed",
|
||||
error_message=error_msg,
|
||||
crawled_url=url,
|
||||
page_title=page_title,
|
||||
http_status=http_status,
|
||||
trafilatura=None,
|
||||
newspaper4k=None,
|
||||
readability=None,
|
||||
)
|
||||
extracted_list.append(failed_article)
|
||||
continue
|
||||
else:
|
||||
# Bypass direto do gate estrutural (sem mídia candidata)
|
||||
metrics["text"] += 1
|
||||
|
||||
log_info(
|
||||
f"⚙️ [{idx}/{total}] Processando extratores (Trafilatura, Newspaper4k, Readability)...",
|
||||
@@ -658,25 +1122,47 @@ def process_batch(
|
||||
|
||||
extracted_list.append(extracted_article)
|
||||
|
||||
# Emissão incondicional do arquivo de mídia *_media.json
|
||||
save_media_json(media_articles, media_output_path)
|
||||
|
||||
elapsed = time.time() - start_time
|
||||
now_iso = datetime.now(timezone.utc).isoformat()
|
||||
|
||||
# Contadores do relatório textual refletem estritamente os itens presentes em extracted_list
|
||||
report = ExtractionBatchReport(
|
||||
source_file=str(input_path),
|
||||
processed_at=now_iso,
|
||||
total_articles=total,
|
||||
total_articles=len(extracted_list),
|
||||
successful_articles=successful_count,
|
||||
failed_articles=failed_count,
|
||||
articles=extracted_list,
|
||||
)
|
||||
|
||||
save_extracted_json(report, output_path)
|
||||
log_info(f"💾 Relatório final gravado com sucesso em: '{output_path}'", silent=silent)
|
||||
save_extracted_json(report, text_output_path)
|
||||
log_info(f"💾 Relatório final gravado com sucesso em: '{text_output_path}'", silent=silent)
|
||||
log_info(f"💾 Arquivo de mídia gravado com sucesso em: '{media_output_path}' ({len(media_articles)} artigo(s))", silent=silent)
|
||||
log_info(
|
||||
f"📊 Resumo: {total} total | {successful_count} sucessos | {failed_count} falhas | Tempo: {elapsed:.2f}s",
|
||||
f"📊 Resumo: {len(extracted_list)} no JSON textual | {successful_count} sucessos | {failed_count} falhas | Tempo: {elapsed:.2f}s",
|
||||
silent=silent,
|
||||
)
|
||||
|
||||
if not silent:
|
||||
media_subtypes = ", ".join(
|
||||
f"{k.split('/')[1]}={v}"
|
||||
for k, v in metrics.items()
|
||||
if k.startswith("media/") and v > 0
|
||||
)
|
||||
subtypes_str = f" ({media_subtypes})" if media_subtypes else ""
|
||||
sys.stderr.write(
|
||||
f"[MEDIA] Métricas de Roteamento:\n"
|
||||
f" - Total avaliados: {metrics['total_evaluated']}\n"
|
||||
f" - Texto: {metrics['text']}\n"
|
||||
f" - Mídia: {metrics['media']}{subtypes_str}\n"
|
||||
f" - Fallbacks: Groq={metrics['fallback_groq']}, OmniRoute={metrics['fallback_omniroute']}\n"
|
||||
f" - Falhas de classificação: {metrics['classification_failed']}\n"
|
||||
)
|
||||
sys.stderr.flush()
|
||||
|
||||
return report
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
# Media Routing Requirements Quality & Implementation Planning Checklist
|
||||
|
||||
**Purpose**: Validate requirements quality, architectural fidelity, and implementation plan completeness for the media routing feature
|
||||
**Created**: 2026-08-24
|
||||
**Feature**: [spec.md](../spec.md) | [plan.md](../plan.md) | [research.md](../research.md) | [data-model.md](../data-model.md)
|
||||
|
||||
**Note**: This checklist is a reviewer-owned review artifact. Mark an item `[x]` only when the reviewer determines the quality criterion is satisfied.
|
||||
**Marker Semantics**: `[x]` means the criterion has been reviewed and satisfied. It does not mean implementation work is complete.
|
||||
|
||||
---
|
||||
|
||||
## 1. Requirement & Plan Fidelity
|
||||
|
||||
- [x] CHK001 Does the specification define behavioral requirements for structural media relevance to the publication while the plan and research define the concrete, minimal DOM parsing algorithm without contradiction? [Fidelity, Spec §FR-001, §FR-004; Plan §4.1; Research §Decisão 1]
|
||||
- [x] CHK002 Are all 5 media subtypes (`video`, `image`, `images`, `embed`, `mixed`) exhaustively specified with their definitions? [Completeness, Spec §FR-009]
|
||||
- [x] CHK003 Are the payload fields for compact LLM input generation fully defined across the plan and data model without arbitrary truncation of editorial text? [Completeness, Plan §4.2; Data-Model §2]
|
||||
- [x] CHK004 Are the required fields, mandatory file generation in all successful batch runs, and minimal envelope structure (`articles`, containing `[]` when zero media articles) for `*_media.json` explicitly specified without extra report fields? [Completeness, Spec §FR-014, §FR-015; Plan §4.4; Contract: media-output.schema.json]
|
||||
- [x] CHK005 Are all 11 mandatory operational metrics enumerated with their exact tracking points documented in the plan? [Completeness, Spec §FR-028; Plan §4.5]
|
||||
- [x] CHK006 Are the specific conditions that trigger operational fallback defined for each provider? [Completeness, Spec §FR-019, §FR-020; Plan §4.3]
|
||||
- [x] CHK007 Are error handling and reporting requirements specified when all 3 LLM providers fail, recording the failure inline in the main JSON with `classification_status: "failed"` without creating separate failure files or DTOs? [Completeness, Spec §FR-022; Plan §4.3; Data-Model §2.6]
|
||||
- [x] CHK008 Are all ungrounded technical promises, invented latency SLOs (<1.5s, <10ms) and artificial benchmarks completely absent from the plan? [Fidelity, Plan §2, §4]
|
||||
|
||||
---
|
||||
|
||||
## 2. Minimal Implementation & Code Reusability
|
||||
|
||||
- [x] CHK009 Were real repository files inspected, reusing existing dependencies (`beautifulsoup4`, `urllib.request`) and avoiding new package installations? [Minimalism, Plan §2, §5]
|
||||
- [x] CHK010 Is the implementation organized as lightweight functions in `scripts/extract_article_contents.py` rather than unnecessary class hierarchies or generic provider frameworks? [Minimalism, Plan §5]
|
||||
- [x] CHK011 Are speculative abstractions, generic provider frameworks, and unnecessary domain DTOs (e.g. `MediaBatchReport`, `MediaMetricsCollector`, `FailedArticle`, `ClassificationFailedArticle`) completely absent? [Minimalism, Plan §5, §7]
|
||||
- [x] CHK012 Was `src/tools/adapters/llm.py` evaluated and its non-reuse properly justified due to its coupling to the ECP domain rather than creating redundant provider frameworks? [Minimalism, Plan §5; Research §Decisão 3]
|
||||
|
||||
---
|
||||
|
||||
## 3. Structural DOM Gate
|
||||
|
||||
- [x] CHK013 Is the DOM structural gate defined without treating `<source>` alone as video, without double-counting `<figure>/<picture>` wrappers, and distinguishing editorial content from page structure (`<header>`, `<nav>`, `<footer>`, `<aside>`)? [Gate, Plan §4.1; Research §Decisão 1]
|
||||
- [x] CHK014 Is the structural gate completely free of regular expressions (`re`), language-specific keywords, site-specific selectors, and ad detectors? [Constraint, Spec §FR-002; Plan §4.1]
|
||||
- [x] CHK015 Does the gate bypass the LLM completely when no relevant candidate media is present, sending the article directly to `extract_all_engines()`? [Gate, Spec §FR-003; Plan §4.1]
|
||||
|
||||
---
|
||||
|
||||
## 4. Classifier Payload & Interface Contracts
|
||||
|
||||
- [x] CHK016 Is the compact payload defined with title, normalized editorial text without HTML and without arbitrary character truncation, and structural media summary? [Payload, Spec §FR-005, §FR-006; Plan §4.2]
|
||||
- [x] CHK017 Does the classifier output contract contain exactly 2 fields (`content_type` and `media_type`), enforcing `text -> media_type = null` and `media -> media_type != null` in application validation? [Contract, Spec §FR-007, §FR-008, §FR-009; Contract: classifier-io.schema.json]
|
||||
- [x] CHK018 Is the prompt unique, concise, and language-independent, explicitly prohibiting reasoning, summaries, rationale, and translation requests? [Prompt, Spec §FR-010; Plan §4.2]
|
||||
- [x] CHK019 Does Ollama use JSON Schema in `format`, Groq use native JSON Schema Structured Output with `openai/gpt-oss-20b`, OmniRoute use JSON Schema Structured Output, and the application strictly validate the two-field schema? [Contract, Spec §FR-011; Plan §4.3]
|
||||
- [x] CHK020 Is deliberative reasoning/thinking explicitly disabled/configured per provider (Ollama: `think=false`, Groq: `reasoning_effort="low"`, OmniRoute: sem parâmetro inventado), ensuring no reasoning content appears in the output? [Reasoning, Plan §4.3; Research §Decisão 3]
|
||||
|
||||
---
|
||||
|
||||
## 5. Sequential Provider Chain & Fallback Rules
|
||||
|
||||
- [x] CHK021 Is the provider order strictly Ollama (`qwen3.5:2b`) → Groq (`openai/gpt-oss-20b`) → OmniRoute (`cgpt-web/gpt-5.5`), executing sequentially and stopping immediately on the first valid response? [Providers, Spec §FR-018, §FR-019, §FR-020, §FR-021; Plan §4.3]
|
||||
- [x] CHK022 Are voting, model consensus, second opinions, LLM-as-a-judge, and confidence thresholds strictly excluded? [Boundary, Spec §FR-021; Plan §4.3, §7]
|
||||
- [x] CHK023 Are automatic per-provider retries, exponential backoff, circuit breakers, discovery frameworks, and health services strictly excluded from the provider chain? [Boundary, Plan §4.3, §7]
|
||||
- [x] CHK024 Are endpoints, models, timeouts, and API keys externalized via environment variables, with OmniRoute having no invented default hostname? [Security, Spec §FR-025; Plan §4.3]
|
||||
|
||||
---
|
||||
|
||||
## 6. I/O Persistence, Counter Semantics & CLI Interface
|
||||
|
||||
- [x] CHK025 Are the naming and path derivation rules for `*_media.json` explicit and identical for default (same stem/dir) and custom `-o/--output` paths? [I/O, Spec §FR-014; Plan §4.4; Contract: cli-interface.md]
|
||||
- [x] CHK026 Does every successful batch run generate `*_media.json`, using `{ "articles": [] }` when zero media articles were classified, without extra report counters or status aggregates? [I/O, Spec §FR-014, §FR-015; Plan §4.4; Contract: media-output.schema.json]
|
||||
- [x] CHK027 Do the main JSON counters (`total_articles`, `successful_articles`, `failed_articles`) reflect strictly the items present in that file, excluding media articles and including classification failures? [Counters, Spec §FR-017; Plan §4.4]
|
||||
- [x] CHK028 Does the implementation preserve 100% of the input metadata dictionary in `input_meta` across text, media, and failure records without dropping unknown fields or inventing default values? [Data Integrity, Plan §4.4; Data-Model §2.1]
|
||||
- [x] CHK029 Does the implementation preserve 100% of the existing CLI interface (`-i`, `-o`, `-l`, `--lang`, `-t`, `-s`) without creating new flags like `--media-output`? [CLI, Spec §FR-030; Plan §4.4]
|
||||
|
||||
---
|
||||
|
||||
## 7. Observability, Logging & Security Constraints
|
||||
|
||||
- [x] CHK030 Are all 11 required operational metrics incremented in memory at the exact points defined and emitted via existing stderr logging mechanisms while respecting the `-s/--silent` flag? [Observability, Spec §FR-028; Plan §4.5]
|
||||
- [x] CHK031 Are logs structured to record provider usage, fallback transitions, final classifications, and total failures without logging full text payloads by default? [Observability, Spec §FR-027; Plan §4.5]
|
||||
- [x] CHK032 Are API keys, tokens, and credentials strictly prevented from appearing in logs, error messages, and JSON outputs? [Security, Spec §FR-026; Plan §4.5]
|
||||
|
||||
---
|
||||
|
||||
## 8. Test Coverage & CI Determinism
|
||||
|
||||
- [x] CHK033 Are all 13 normative test scenarios (A through M) plus zero-media batch generation covered in the test plan, spanning text-only, single media, multiple images, mixed media, long text with media, fallbacks, and multilingual cases? [Test Coverage, Spec §User Stories; Plan §6]
|
||||
- [x] CHK034 Are deterministic unit and integration tests defined to simulate the 5 provider states in CI without requiring live network calls to Ollama, Groq, or OmniRoute? [Test CI, Spec §FR-024; Plan §4.3, §6]
|
||||
- [x] CHK035 Is the zero-regex policy verified via static AST inspection in `tests/scripts/check_zero_regex.py` across all feature modules and tests? [Zero-Regex, Spec §SC-007; Plan §5, §6]
|
||||
|
||||
---
|
||||
|
||||
## 9. Global Artifact Consistency Check
|
||||
|
||||
- [x] CHK036 Are `spec.md`, `plan.md`, `research.md`, `data-model.md`, `quickstart.md`, `classifier-io.schema.json`, `media-output.schema.json`, and `cli-interface.md` 100% consistent with each other, confirming that classification failures are recorded inline with `classification_status: "failed"` and without any reintroduction of `MediaBatchReport`, `MediaMetricsCollector`, or DTOs de falha? [Consistency, Plan §5, §7; Data-Model §2]
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- Mark items `[x]` only after review confirms the quality criterion is satisfied
|
||||
- Leave items unchecked when they still require clarification, correction, or reviewer evaluation
|
||||
- `/speckit-implement` reads checklist checkbox state as a gate and must not modify markers
|
||||
- Items are numbered sequentially (CHK001–CHK036) for easy reference
|
||||
@@ -0,0 +1,34 @@
|
||||
# Specification Quality Checklist: Classificação e Roteamento de Notícias Predominantemente de Mídia
|
||||
|
||||
**Purpose**: Validate specification completeness and quality before proceeding to planning
|
||||
**Created**: 2026-08-24
|
||||
**Feature**: [spec.md](../spec.md)
|
||||
|
||||
## Content Quality
|
||||
|
||||
- [x] No implementation details (languages, frameworks, APIs)
|
||||
- [x] Focused on user value and business needs
|
||||
- [x] Written for non-technical stakeholders
|
||||
- [x] All mandatory sections completed
|
||||
|
||||
## Requirement Completeness
|
||||
|
||||
- [x] No [NEEDS CLARIFICATION] markers remain
|
||||
- [x] Requirements are testable and unambiguous
|
||||
- [x] Success criteria are measurable
|
||||
- [x] Success criteria are technology-agnostic (no implementation details)
|
||||
- [x] All acceptance scenarios are defined
|
||||
- [x] Edge cases are identified
|
||||
- [x] Scope is clearly bounded
|
||||
- [x] Dependencies and assumptions identified
|
||||
|
||||
## Feature Readiness
|
||||
|
||||
- [x] All functional requirements have clear acceptance criteria
|
||||
- [x] User scenarios cover primary flows
|
||||
- [x] Feature meets measurable outcomes defined in Success Criteria
|
||||
- [x] No implementation details leak into specification
|
||||
|
||||
## Notes
|
||||
|
||||
- All items passed specification validation. Ready for planning phase (`/speckit-plan`).
|
||||
@@ -0,0 +1,37 @@
|
||||
{
|
||||
"$schema": "http://json-schema.org/draft-07/schema#",
|
||||
"name": "media_classifier",
|
||||
"title": "MediaClassifierOutput",
|
||||
"description": "Contrato estruturado estrito de dois campos emitido pelo classificador semântico",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"content_type",
|
||||
"media_type"
|
||||
],
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"content_type": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"text",
|
||||
"media"
|
||||
],
|
||||
"description": "Classificação da publicação: text para texto jornalístico substancial, media para conteúdo predominantemente de mídia"
|
||||
},
|
||||
"media_type": {
|
||||
"type": [
|
||||
"string",
|
||||
"null"
|
||||
],
|
||||
"enum": [
|
||||
"video",
|
||||
"image",
|
||||
"images",
|
||||
"embed",
|
||||
"mixed",
|
||||
null
|
||||
],
|
||||
"description": "Subtipo de mídia quando content_type=media, ou null quando content_type=text"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,123 @@
|
||||
# CLI Contract & Routing Specification: `scripts/extract_article_contents.py`
|
||||
|
||||
## 1. Interface de Linha de Comando (CLI)
|
||||
|
||||
O script `scripts/extract_article_contents.py` preserva 100% da interface CLI existente sem novos argumentos:
|
||||
|
||||
```bash
|
||||
python scripts/extract_article_contents.py \
|
||||
-i, --input PATH # (Obrigatório) JSON de busca de notícias de entrada
|
||||
[-o, --output PATH] # (Opcional) Caminho do JSON de saída textual
|
||||
[-l, --limit N] # (Opcional) Limite máximo de artigos a processar
|
||||
[--lang, --language CODE] # (Opcional) Código do idioma para NLP (ex: pt, es, en)
|
||||
[-t, --timeout SEC] # (Opcional) Timeout em segundos para navegação (default: 30)
|
||||
[-s, --silent] # (Opcional) Suprime logs no stderr
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Regras de Resolução de Arquivos de Saída
|
||||
|
||||
### Caso A: Sem `-o/--output` (Padrão)
|
||||
|
||||
Para uma entrada: `out/river_plate.json`
|
||||
|
||||
- **Saída Textual**: `out/river_plate_extracted.json`
|
||||
- **Saída de Mídia**: `out/river_plate_media.json` (gerado obrigatoriamente; contém `{"articles": []}` se nenhuma mídia for identificada)
|
||||
|
||||
### Caso B: Com `-o/--output` customizado
|
||||
|
||||
Para uma entrada: `out/noticias.json` com `-o out/processados/brasil_completo.json`
|
||||
|
||||
- **Saída Textual**: `out/processados/brasil_completo.json`
|
||||
- **Saída de Mídia**: `out/processados/brasil_completo_media.json` (gerado obrigatoriamente; contém `{"articles": []}` se nenhuma mídia for identificada)
|
||||
|
||||
Regra:
|
||||
- Diretório: mesmo diretório da saída textual informada.
|
||||
- Nome base (stem): mesmo stem da saída textual informada.
|
||||
- Sufixo: `_media`.
|
||||
- Extensão: `.json`.
|
||||
- Geração: O arquivo `*_media.json` MUST ser gerado em toda execução bem-sucedida do lote, sem omissão condicional.
|
||||
|
||||
---
|
||||
|
||||
## 3. Estrutura do Arquivo de Saída Textual (`*_extracted.json`)
|
||||
|
||||
```json
|
||||
{
|
||||
"source_file": "out/river_plate.json",
|
||||
"processed_at": "2026-08-24T20:30:00.000000+00:00",
|
||||
"total_articles": 8,
|
||||
"successful_articles": 7,
|
||||
"failed_articles": 1,
|
||||
"articles": [
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo": "Notícia textual completa",
|
||||
"url": "https://example.com/noticia-1",
|
||||
"subtitulo": "Subtítulo",
|
||||
"quando_publicado": "há 2 horas",
|
||||
"pagina": 1
|
||||
},
|
||||
"extraction_status": "success",
|
||||
"error_message": null,
|
||||
"crawled_url": "https://example.com/noticia-1",
|
||||
"page_title": "Título no DOM",
|
||||
"http_status": 200,
|
||||
"trafilatura": { "text": "...", "markdown": "...", "title": "..." },
|
||||
"newspaper4k": { "text": "...", "summary": "...", "keywords": [] },
|
||||
"readability": { "cleaned_text": "...", "cleaned_html": "..." }
|
||||
},
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo": "Notícia onde todos provedores falharam",
|
||||
"url": "https://example.com/falha"
|
||||
},
|
||||
"classification_status": "failed",
|
||||
"error_message": "media classification providers unavailable",
|
||||
"crawled_url": "https://example.com/falha",
|
||||
"page_title": null,
|
||||
"http_status": null,
|
||||
"trafilatura": null,
|
||||
"newspaper4k": null,
|
||||
"readability": null
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Estrutura do Arquivo de Saída de Mídia (`*_media.json`)
|
||||
|
||||
```json
|
||||
{
|
||||
"articles": [
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo": "Vídeo dos melhores momentos do jogo",
|
||||
"url": "https://example.com/video-gols",
|
||||
"subtitulo": "Assista ao lance",
|
||||
"quando_publicado": "há 1 hora",
|
||||
"pagina": 1
|
||||
},
|
||||
"crawled_url": "https://example.com/video-gols",
|
||||
"page_title": "Vídeo dos Melhores Momentos",
|
||||
"http_status": 200,
|
||||
"content_type": "media",
|
||||
"media_type": "video"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Códigos de Saída (Exit Codes)
|
||||
|
||||
| Exit Code | Significado |
|
||||
|:---|:---|
|
||||
| `0` | Lote processado com sucesso (arquivos gravados) |
|
||||
| `1` | Arquivo de entrada inexistente ou erro de argumentos CLI |
|
||||
| `2` | Erro fatal não tratado |
|
||||
| `130` | Interrupção pelo usuário (`SIGINT` / Ctrl+C) |
|
||||
@@ -0,0 +1,56 @@
|
||||
{
|
||||
"$schema": "http://json-schema.org/draft-07/schema#",
|
||||
"title": "MediaOutputFileSchema",
|
||||
"description": "Schema de validação do arquivo de saída de mídias (*_media.json)",
|
||||
"type": "object",
|
||||
"required": [
|
||||
"articles"
|
||||
],
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"articles": {
|
||||
"type": "array",
|
||||
"description": "Lista de publicações classificadas como predominantemente de mídia",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": [
|
||||
"input_meta",
|
||||
"crawled_url",
|
||||
"page_title",
|
||||
"http_status",
|
||||
"content_type",
|
||||
"media_type"
|
||||
],
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"input_meta": {
|
||||
"type": "object",
|
||||
"required": [
|
||||
"titulo",
|
||||
"url"
|
||||
],
|
||||
"properties": {
|
||||
"titulo": { "type": "string" },
|
||||
"url": { "type": "string" },
|
||||
"subtitulo": { "type": ["string", "null"] },
|
||||
"quando_publicado": { "type": ["string", "null"] },
|
||||
"pagina": { "type": ["integer", "null"] }
|
||||
},
|
||||
"additionalProperties": true
|
||||
},
|
||||
"crawled_url": { "type": "string" },
|
||||
"page_title": { "type": ["string", "null"] },
|
||||
"http_status": { "type": ["integer", "null"] },
|
||||
"content_type": {
|
||||
"type": "string",
|
||||
"enum": ["media"]
|
||||
},
|
||||
"media_type": {
|
||||
"type": "string",
|
||||
"enum": ["video", "image", "images", "embed", "mixed"]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,138 @@
|
||||
# Data Model & State Transitions: Classificação e Roteamento de Notícias de Mídia
|
||||
|
||||
**Feature**: `007-media-article-routing`
|
||||
**Date**: 2026-08-24
|
||||
**Status**: Completed
|
||||
|
||||
---
|
||||
|
||||
## 1. Diagrama Entidade-Relacionamento e Fluxo de Dados
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Input[InputArticle JSON] --> Crawler[Foxcape Crawler]
|
||||
Crawler --> LoadedPage[DOM Carregada: html, title, status]
|
||||
LoadedPage --> Gate[DOM Media Gate: detecção estrutural]
|
||||
|
||||
Gate -- "Sem mídia candidata relevante" --> TextPipeline[Multimotor Textual: Trafilatura + Newspaper4k + Readability]
|
||||
|
||||
Gate -- "Com mídia candidata relevante" --> CompactBuilder[Montagem de Payload Compacto]
|
||||
CompactBuilder --> Classifier[Cadeia Sequencial LLM: Ollama -> Groq -> OmniRoute]
|
||||
|
||||
Classifier -- "content_type = text" --> TextPipeline
|
||||
TextPipeline --> ExtractedRecord[ExtractedArticle: Sucesso Textual]
|
||||
ExtractedRecord --> MainReport[ExtractionBatchReport -> *_extracted.json]
|
||||
|
||||
Classifier -- "content_type = media" --> MediaRecord[MediaArticle: Registro de Mídia]
|
||||
MediaRecord --> MediaEnvelope[MediaOutputFile -> *_media.json]
|
||||
|
||||
Classifier -- "Falha total dos 3 provedores" --> FailureRecord[Registro de Falha de Classificação]
|
||||
FailureRecord --> MainReport
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Modelos de Dados (Dataclasses e Estruturas de Runtime)
|
||||
|
||||
### 2.1 `InputArticle`
|
||||
Representa a notícia original carregada do JSON de entrada.
|
||||
- `titulo: str` (Obrigatório)
|
||||
- `url: str` (Obrigatório)
|
||||
- `subtitulo: str | None` (Opcional, preservado conforme entrada)
|
||||
- `quando_publicado: str | None` (Opcional, preservado conforme entrada)
|
||||
- `pagina: int | None` (Opcional, preservado conforme entrada)
|
||||
- Campos adicionais da entrada são preservados no dicionário `input_meta`.
|
||||
|
||||
### 2.2 `MediaCandidateInfo` (Estrutura Interna do Gate)
|
||||
Resultado da análise estrutural da DOM:
|
||||
- `has_candidate_media: bool`
|
||||
- `has_video: bool`
|
||||
- `image_count: int`
|
||||
- `has_embed: bool`
|
||||
|
||||
### 2.3 `MediaClassification` (Estrutura Interna de Saída do LLM)
|
||||
Resultado estrito retornado pelo classificador LLM:
|
||||
- `content_type: Literal["text", "media"]`
|
||||
- `media_type: Literal["video", "image", "images", "embed", "mixed"] | None`
|
||||
- Regra semântica: `content_type == "text"` $\iff$ `media_type is None`.
|
||||
|
||||
### 2.4 `MediaArticle` (Persistência em `*_media.json`)
|
||||
Registro consolidado de publicação classificada como mídia:
|
||||
- `input_meta: dict[str, Any]` (Preserva todos os metadados recebidos da entrada)
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `content_type: Literal["media"]`
|
||||
- `media_type: Literal["video", "image", "images", "embed", "mixed"]`
|
||||
|
||||
### 2.5 `ExtractedArticle` (Persistência Textual em `*_extracted.json`)
|
||||
Representa exclusivamente artigos textuais que efetivamente passaram pelo multimotor (`extract_all_engines`):
|
||||
- `input_meta: InputArticle`
|
||||
- `extraction_status: Literal["success", "failed"]`
|
||||
- `error_message: str | None`
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `trafilatura: TrafilaturaData | None`
|
||||
- `newspaper4k: NewspaperData | None`
|
||||
- `readability: ReadabilityData | None`
|
||||
|
||||
### 2.6 Registro de Falha de Classificação (Persistência Inline em `*_extracted.json`)
|
||||
Para artigos que sofram falha operacional dos três provedores LLM, o registro é persistido inline no array `articles` do JSON principal sem ter executado os extratores textuais:
|
||||
- `input_meta: dict[str, Any]`
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `classification_status: Literal["failed"]`
|
||||
- `error_message: str`
|
||||
- `trafilatura: None`
|
||||
- `newspaper4k: None`
|
||||
- `readability: None`
|
||||
|
||||
### 2.7 `ExtractionBatchReport`
|
||||
Envelope consolidado de saída gravado no arquivo principal (`*_extracted.json`):
|
||||
- `source_file: str`
|
||||
- `processed_at: str` (ISO 8601 UTC)
|
||||
- `total_articles: int` (Total de registros presentes no array `articles` do JSON principal)
|
||||
- `successful_articles: int` (Total de artigos textuais processados com sucesso)
|
||||
- `failed_articles: int` (Total de registros de falha presentes, incluindo erros de crawl e de classificação)
|
||||
- `articles: list[dict[str, Any] | ExtractedArticle]`
|
||||
|
||||
### 2.8 `MediaOutputFile`
|
||||
Envelope mínimo gravado no arquivo de mídia (`*_media.json`):
|
||||
- `articles: list[MediaArticle]`
|
||||
|
||||
---
|
||||
|
||||
## 3. Máquina de Estados do Processamento de Artigo
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> Crawling
|
||||
Crawling --> CrawlFailed: Erro HTTP / Timeout Foxcape
|
||||
CrawlFailed --> MainJSONRecord: Registra falha de crawl
|
||||
|
||||
Crawling --> StructuralInspection: DOM carregada com sucesso
|
||||
StructuralInspection --> TextExtraction: Sem mídia candidata relevante
|
||||
StructuralInspection --> LLMClassification: Mídia candidata relevante detectada
|
||||
|
||||
state LLMClassification {
|
||||
[*] --> TryOllama
|
||||
TryOllama --> ValidOutput: Resposta com schema válido
|
||||
TryOllama --> TryGroq: Falha operacional Ollama
|
||||
TryGroq --> ValidOutput: Resposta com schema válido
|
||||
TryGroq --> TryOmniRoute: Falha operacional Groq
|
||||
TryOmniRoute --> ValidOutput: Resposta com schema válido
|
||||
TryOmniRoute --> AllProvidersFailed: Falha operacional OmniRoute
|
||||
}
|
||||
|
||||
ValidOutput --> TextExtraction: content_type == text
|
||||
ValidOutput --> MediaRouting: content_type == media
|
||||
AllProvidersFailed --> MainJSONRecord: Registra classification_status = failed
|
||||
|
||||
TextExtraction --> MainJSONRecord: Executa Trafilatura + Newspaper + Readability
|
||||
MediaRouting --> MediaJSONRecord: Grava em *_media.json
|
||||
|
||||
MainJSONRecord --> [*]
|
||||
MediaJSONRecord --> [*]
|
||||
```
|
||||
@@ -0,0 +1,181 @@
|
||||
# Implementation Plan: Classificação e Roteamento de Notícias Predominantemente de Mídia
|
||||
|
||||
**Branch**: `007-media-article-routing` | **Date**: 2026-08-24 | **Spec**: [spec.md](spec.md)
|
||||
|
||||
**Input**: Feature specification from `/specs/007-media-article-routing/spec.md`
|
||||
|
||||
---
|
||||
|
||||
## 1. Summary
|
||||
|
||||
Esta feature adiciona ao script [`scripts/extract_article_contents.py`](../../scripts/extract_article_contents.py) a capacidade de identificar publicações cujo conteúdo informativo principal seja mídia (vídeo, imagem única, múltiplas imagens/galeria, embed ou mídia mista) e cujo texto atue apenas como introdução ou contextualização.
|
||||
|
||||
Após o crawl com Foxcape, a página sofre uma análise estrutural na DOM (zero-regex via BeautifulSoup). Se nenhuma mídia relevante for detectada, o artigo segue diretamente para o multimotor textual (`extract_all_engines`). Se houver mídia relevante, o texto editorial normalizado e o resumo estrutural são avaliados por uma cadeia sequencial de LLMs com fallback puramente operacional (Ollama/Qwen3.5 2B $\rightarrow$ Groq/openai/gpt-oss-20b $\rightarrow$ OmniRoute/cgpt-web/gpt-5.5). Artigos classificados como `media` são desviados antes dos três motores textuais e gravados no arquivo de mídia dedicado (`*_media.json`), enquanto artigos textuais e eventuais falhas de classificação continuam para o arquivo principal (`*_extracted.json`).
|
||||
|
||||
---
|
||||
|
||||
## 2. Technical Context
|
||||
|
||||
- **Linguagem & Tipagem**: Python `>=3.10` com anotações de tipo completas em conformidade com o princípio de qualidade do repositório (Constitution §Technical Constraints).
|
||||
- **Dependências Reutilizadas**: `beautifulsoup4`, `trafilatura`, `newspaper4k`, `readability-lxml`, `foxcape` (todas já instaladas e ativas no projeto).
|
||||
- **Cliente HTTP para LLMs**: Biblioteca padrão do Python (`urllib.request`, `urllib.error`, `json`), sem novas dependências externas.
|
||||
- **Armazenamento / I/O**: Arquivos JSON no filesystem local (`*_extracted.json` e `*_media.json`).
|
||||
- **Suíte de Testes**: `pytest` com simulação determinística dos 5 estados da cadeia de provedores; script estático existente [`tests/scripts/check_zero_regex.py`](../../tests/scripts/check_zero_regex.py).
|
||||
- **Tipo de Projeto**: Pipeline de linha de comando (CLI) existente.
|
||||
|
||||
---
|
||||
|
||||
## 3. Constitution Check
|
||||
|
||||
| Princípio Constitucional | Avaliação Técnica & Rastreabilidade | Status |
|
||||
|:---|:---|:---:|
|
||||
| **I. Modularity & CLI-First** | Integrado diretamente em `scripts/extract_article_contents.py`, preservando 100% dos parâmetros CLI (`-i`, `-o`, `-l`, `--lang`, `-t`, `-s`) e exit codes (`0`, `1`, `2`, `130`). | **PASS** |
|
||||
| **II. Determinism & Data Integrity** | Contrato estruturado de 2 campos no LLM com configuração determinística; parsing estrito de DOM; preservação intacta dos metadados de entrada (`input_meta`). | **PASS** |
|
||||
| **III. Multi-Engine & Fault-Tolerant Fallback** | Artigos textuais continuam processados pelos 3 motores; cadeia sequencial com 2 níveis de contingência operacional (Ollama $\rightarrow$ Groq $\rightarrow$ OmniRoute); falha isolada por artigo. | **PASS** |
|
||||
| **IV. Test-First & Empirical Validation** | Suíte de testes cobrindo cenários A a M e os 5 estados da cadeia de provedores na CI sem chamadas de rede externas; verificação estática Zero-Regex. | **PASS** |
|
||||
| **V. Observability & Structured Logging** | Contabilização e emissão das 11 métricas operacionais obrigatórias no resumo/log stderr existente; logs de fallbacks e provedor sem expor segredos ou payload integral. | **PASS** |
|
||||
|
||||
---
|
||||
|
||||
## 4. Arquitetura e Decisões Técnicas Fechadas
|
||||
|
||||
### 4.1 Gate Estrutural na DOM (Zero-Regex)
|
||||
1. **Localização da Região Editorial**:
|
||||
- Inspecionar a DOM carregada buscando nós na seguinte ordem de precedência: `<article>`, `<main>`, `<div role="main">`, ou `<body>` caso nenhuma anterior exista.
|
||||
- Descartar nós estruturais fora do conteúdo da matéria (`<header>`, `<nav>`, `<footer>`, `<aside>`).
|
||||
- Sem regex, sem heurísticas de classes CSS, sem seletores específicos de sites e sem detector de anúncios.
|
||||
2. **Critério de Mídia Candidata Relevante**:
|
||||
- **Vídeo**: tags `<video>` $\rightarrow$ `has_video = True`. Tags `<source>` pertencentes a `<video>` não são contadas isoladamente.
|
||||
- **Imagem**: contagem das tags `<img>` reais na região editorial. Wrappers como `<figure>` e `<picture>` não incrementam ou duplicam o contador.
|
||||
- **Embed**: tags `<iframe>`, `<embed>`, `<object>` $\rightarrow$ `has_embed = True`. Sem listas de domínios externos.
|
||||
- **Múltiplas Imagens**: contagem $\ge 2$ de tags `<img>` na região editorial.
|
||||
3. **Decisão do Gate**:
|
||||
- Retorna `MediaCandidateInfo(has_candidate_media, has_video, image_count, has_embed)`.
|
||||
- Se `has_candidate_media == False` $\rightarrow$ desvio direto para `extract_all_engines()` (zero chamadas LLM).
|
||||
- Se `has_candidate_media == True` $\rightarrow$ montagem de payload e execução da cadeia LLM.
|
||||
|
||||
### 4.2 Payload Compacto do Classificador
|
||||
- **Campos Extraídos**:
|
||||
- `title`: Título obtido da tag `<title>` ou `<h1>` da matéria.
|
||||
- `text_content`: Textos dos parágrafos (`<p>`) da região editorial, normalizados sem tags HTML e preservando o conteúdo jornalístico substancial da matéria (sem truncamento arbitrário de caracteres).
|
||||
- `media_summary`: Resumo dos elementos de mídia identificados no gate (`has_video`, `image_count`, `has_embed`).
|
||||
- **Prompt Único Multilíngue**: Instrução concisa solicitando estritamente a classificação em `content_type` (`text` ou `media`) e `media_type` (`video`, `image`, `images`, `embed`, `mixed` ou `null`), sem reasoning deliberativo, sem tradução e sem resumos.
|
||||
|
||||
### 4.3 Cadeia Sequencial de Provedores, Structured Output e Controle de Reasoning
|
||||
1. **Configuração dos Provedores e Reasoning**:
|
||||
- **Ollama (Primário)**:
|
||||
- Endpoint: `os.environ.get("OLLAMA_ENDPOINT", "http://localhost:11434")`
|
||||
- Modelo: `os.environ.get("OLLAMA_MODEL", "qwen3.5:2b")`
|
||||
- Timeout: `int(os.environ.get("OLLAMA_TIMEOUT", "10"))`
|
||||
- Structured Output & Reasoning: passa o JSON Schema diretamente no campo `format` da requisição `/api/chat`, com `options: {"temperature": 0.0}` e `think: false` para desabilitar explicitamente thinking no Qwen3.5 2B.
|
||||
- **Groq (1º Fallback)**:
|
||||
- Endpoint: `os.environ.get("GROQ_ENDPOINT", "https://api.groq.com/openai/v1/chat/completions")`
|
||||
- API Key: `os.environ.get("GROQ_API_KEY")`
|
||||
- Modelo: `os.environ.get("GROQ_MODEL", "openai/gpt-oss-20b")`
|
||||
- Timeout: `int(os.environ.get("GROQ_TIMEOUT", "15"))`
|
||||
- Structured Output & Reasoning: `response_format={"type": "json_schema", "json_schema": {"name": "media_classifier", "strict": True, "schema": <schema>}}`, `temperature: 0.0` e `reasoning_effort: "low"`.
|
||||
- **OmniRoute (2º Fallback)**:
|
||||
- Endpoint: `os.environ.get("OMNIROUTE_ENDPOINT")` (configuração externa obrigatória sem default inventado)
|
||||
- API Key: `os.environ.get("OMNIROUTE_API_KEY")`
|
||||
- Modelo: `os.environ.get("OMNIROUTE_MODEL", "cgpt-web/gpt-5.5")`
|
||||
- Timeout: `int(os.environ.get("OMNIROUTE_TIMEOUT", "20"))`
|
||||
- Structured Output & Reasoning: `response_format` com JSON Schema compatível com a instalação OpenAI-compatible utilizada, `temperature: 0.0` quando suportado, sem parâmetro inventado de reasoning.
|
||||
2. **Regra de Transição e Validação de Schema**:
|
||||
- A aplicação sempre valida os dois campos recebidos:
|
||||
- `content_type` $\in$ `{"text", "media"}`;
|
||||
- se `content_type == "text"`, `media_type` deve ser `None`;
|
||||
- se `content_type == "media"`, `media_type` deve ser um de `{"video", "image", "images", "embed", "mixed"}`.
|
||||
- Resposta válida $\rightarrow$ encerra a cadeia imediatamente (primeira resposta válida conclui).
|
||||
- Falha operacional (timeout, conexão recusada, erro HTTP ou schema incompatível) $\rightarrow$ avança imediatamente para o próximo provedor na ordem estrita Ollama $\rightarrow$ Groq $\rightarrow$ OmniRoute.
|
||||
- Sem retries por provedor, sem exponential backoff, sem circuit breakers e sem discovery dinâmico.
|
||||
- Se todos os 3 provedores falharem $\rightarrow$ artigo recebe `classification_status = "failed"` e `error_message`, sendo gravado inline no JSON principal sem multimotor e sem interromper o lote.
|
||||
|
||||
### 4.4 I/O, Nomenclatura de Arquivos e Contadores
|
||||
- **Regras de Resolução de Caminhos**:
|
||||
- Sem `-o`: entrada `out/river_plate.json` $\rightarrow$ textual `out/river_plate_extracted.json`, mídia `out/river_plate_media.json`.
|
||||
- Com `-o out/dir/saida.json` $\rightarrow$ textual `out/dir/saida.json`, mídia `out/dir/saida_media.json`.
|
||||
- Nenhuma nova flag CLI (sem `--media-output`).
|
||||
- **Arquivo de Mídia (`*_media.json`)**:
|
||||
- O arquivo `*_media.json` MUST ser gerado em toda execução bem-sucedida do lote.
|
||||
- Envelope contendo estritamente `{ "articles": [ MediaArticle... ] }`.
|
||||
- Quando nenhum artigo for classificado como mídia no lote, o arquivo conterá exatamente `{"articles": []}` (sem omissão condicional).
|
||||
- **Arquivo Textual Principal (`*_extracted.json`)**:
|
||||
Envelope `ExtractionBatchReport` onde os contadores refletem estritamente os registros presentes no arquivo:
|
||||
- `total_articles`: contagem de registros presentes no array `articles`;
|
||||
- `successful_articles`: artigos textuais processados com sucesso pelo multimotor;
|
||||
- `failed_articles`: registros de falha presentes (falhas de crawl ou com `classification_status = "failed"`).
|
||||
- **Preservação de Metadados (`input_meta`)**:
|
||||
O dicionário de metadados da entrada é preservado integralmente em `input_meta` para todos os registros (textuais, mídia e falhas), exigindo `titulo` e `url`, sem inventar valores default e sem descartar campos adicionais.
|
||||
|
||||
### 4.5 Observabilidade, Logging e Pontos Exatos de Tracking das 11 Métricas
|
||||
As 11 métricas são contabilizadas em memória (dicionário local dentro de `process_batch`) e registradas nos seguintes pontos exatos:
|
||||
1. `total_evaluated`: incrementado quando um artigo com crawl bem-sucedido entra no gate estrutural da DOM;
|
||||
2. `text`: incrementado quando: (a) não existe mídia candidata no gate (bypass direto), OU (b) LLM retorna `content_type="text"`;
|
||||
3. `media`: incrementado exclusivamente quando o LLM retorna `content_type="media"`;
|
||||
4. `media/video`: incrementado em conjunto com `media` quando `media_type == "video"`;
|
||||
5. `media/image`: incrementado em conjunto com `media` quando `media_type == "image"`;
|
||||
6. `media/images`: incrementado em conjunto com `media` quando `media_type == "images"`;
|
||||
7. `media/embed`: incrementado em conjunto com `media` quando `media_type == "embed"`;
|
||||
8. `media/mixed`: incrementado em conjunto com `media` quando `media_type == "mixed"`;
|
||||
9. `fallback_groq`: incrementado imediatamente antes de despachar a chamada HTTP para o Groq;
|
||||
10. `fallback_omniroute`: incrementado imediatamente antes de despachar a chamada HTTP para o OmniRoute;
|
||||
11. `classification_failed`: incrementado uma única vez por artigo quando os 3 provedores falharem cumulativamente.
|
||||
|
||||
- **Logging no stderr**: Registro de provider utilizado, fallbacks acionados, classificação e falhas. Respeita a flag `-s/--silent`. Segredos e payloads textuais integrais nunca são logados por padrão.
|
||||
|
||||
---
|
||||
|
||||
## 5. Organização de Arquivos (Mínima e Sem Classes Desnecessárias)
|
||||
|
||||
### Funções Implementadas em `scripts/extract_article_contents.py`
|
||||
Para manter o mínimo de código e evitar classes desnecessárias:
|
||||
- `detect_candidate_media(soup: BeautifulSoup) -> MediaCandidateInfo`: Função pura de inspeção estrutural na DOM (sem regex).
|
||||
- `build_compact_payload(soup: BeautifulSoup, candidate_info: MediaCandidateInfo) -> str`: Função pura de extração e normalização do texto editorial e resumo estrutural.
|
||||
- `classify_media_content(payload: str) -> tuple[MediaClassification | None, str | None]`: Função que orquestra a chamada sequencial aos 3 provedores via `urllib.request` e validação estrita.
|
||||
- `save_media_json(articles: list[dict[str, Any]], output_path: Path) -> None`: Gravação do arquivo de mídia com a mesma estratégia de escrita JSON do script existente.
|
||||
- Atualização do loop `process_batch` para incorporar o desvio, gravação de `*_media.json` e contabilidade das 11 métricas.
|
||||
|
||||
*(Nota de Reutilização: O módulo `src/tools/adapters/llm.py` foi inspecionado; ele é altamente especializado na desambiguação de entidades ECP da feature 001 com classes acopladas, de modo que a integração via funções diretas com `urllib.request` em `extract_article_contents.py` é a solução mais desacoplada, limpa e com menor diff).*
|
||||
|
||||
### Arquivo Existente Reutilizado
|
||||
- [`tests/scripts/check_zero_regex.py`](../../tests/scripts/check_zero_regex.py):
|
||||
- Inclusão dos novos arquivos de teste no escopo de validação estática.
|
||||
|
||||
### Novos Arquivos de Teste
|
||||
- `tests/unit/test_media_classifier.py`:
|
||||
- Testes do gate DOM (vídeo sem falso positivo de source, contagem correta de imagens sem duplicar figure/picture, embeds);
|
||||
- Testes de montagem do payload com texto completo;
|
||||
- Testes de validação de schema e simulação dos 5 estados da cadeia de provedores.
|
||||
- `tests/integration/test_media_routing.py`:
|
||||
- Testes de integração em lote cobrindo cenários A a M;
|
||||
- Validação de caminhos `-o`, contadores, geração incondicional de `*_media.json` (com `[]` quando vazio) e persistência de falha inline.
|
||||
|
||||
---
|
||||
|
||||
## 6. Matriz de Rastreabilidade (Requisitos $\rightarrow$ Código $\rightarrow$ Testes)
|
||||
|
||||
| Requisito | Descrição | Implementação em `extract_article_contents.py` | Teste Correspondente |
|
||||
|:---|:---|:---|:---|
|
||||
| **FR-001 - FR-004** | Análise estrutural da DOM e desvio sem LLM | `detect_candidate_media()` | `test_media_classifier.py::test_structural_gate_*` |
|
||||
| **FR-005 - FR-006** | Payload compacto e proibições de mídia binária | `build_compact_payload()` | `test_media_classifier.py::test_compact_payload_*` |
|
||||
| **FR-007 - FR-011** | Contrato estruturado de 2 campos e Structured Output | `classify_media_content()` | `test_media_classifier.py::test_schema_validation_*` |
|
||||
| **FR-012 - FR-017** | Roteamento textual/mídia, `*_media.json` e contadores | `process_batch()`, `save_media_json()` | `test_media_routing.py::test_routing_and_counters_*` |
|
||||
| **FR-018 - FR-024** | Cadeia sequencial Ollama $\rightarrow$ Groq $\rightarrow$ OmniRoute e falha total | `classify_media_content()` | `test_media_classifier.py::test_provider_chain_*` |
|
||||
| **FR-025 - FR-026** | Configuração externa e mascaramento de segredos | `classify_media_content()` | `test_media_routing.py::test_security_secrets_masked` |
|
||||
| **FR-027 - FR-028** | 11 métricas operacionais e logging | `process_batch()` | `test_media_routing.py::test_metrics_logging` |
|
||||
| **FR-029 - FR-030** | Suporte multilíngue e compatibilidade CLI | `parse_arguments()`, `process_batch()` | `test_media_routing.py::test_cli_compatibility` |
|
||||
| **SC-007** | Zero-Regex em toda a nova implementação | AST Checker | `tests/scripts/check_zero_regex.py` |
|
||||
|
||||
---
|
||||
|
||||
## 7. Invariantes e Limites de Escopo
|
||||
|
||||
Fica expressamente estabelecido que a implementação **NÃO DEVE** introduzir:
|
||||
- Bancos de dados, filas de mensagens, DLQ ou novos workers;
|
||||
- Novos serviços, microserviços ou processos autônomos;
|
||||
- Pipelines adicionais de NLP ou frameworks de agentes (LangChain, LangGraph, LLM-as-a-judge);
|
||||
- Votação, consenso entre modelos ou fallbacks por incerteza subjetiva;
|
||||
- Download de mídia binária, OCR, visão computacional ou transcrição de áudio/vídeo;
|
||||
- Interação com carousels ou chamadas a APIs de redes sociais;
|
||||
- Alterações no comportamento de carregamento do Foxcape ou na lógica interna de Trafilatura, Newspaper4k e Readability;
|
||||
- Criação de novas flags CLI como `--media-output`, novos arquivos como `*_failed.json` ou classes desnecessárias como `MediaBatchReport` e `MediaMetricsCollector`.
|
||||
@@ -0,0 +1,68 @@
|
||||
# Quickstart & Validation Guide: Classificação e Roteamento de Mídia
|
||||
|
||||
**Feature**: `007-media-article-routing`
|
||||
**Date**: 2026-08-24
|
||||
**Status**: Completed
|
||||
|
||||
---
|
||||
|
||||
## 1. Configuração de Variáveis de Ambiente (Configuração Externa)
|
||||
|
||||
As credenciais e endpoints dos provedores devem ser configurados no ambiente de execução:
|
||||
|
||||
```bash
|
||||
# Provedor Primário (Ollama Local)
|
||||
OLLAMA_ENDPOINT="http://localhost:11434"
|
||||
OLLAMA_MODEL="qwen3.5:2b"
|
||||
OLLAMA_TIMEOUT="10"
|
||||
|
||||
# 1º Fallback (Groq)
|
||||
GROQ_ENDPOINT="https://api.groq.com/openai/v1/chat/completions"
|
||||
GROQ_API_KEY="gsk_..."
|
||||
GROQ_MODEL="openai/gpt-oss-20b"
|
||||
GROQ_TIMEOUT="15"
|
||||
|
||||
# 2º Fallback (OmniRoute)
|
||||
OMNIROUTE_ENDPOINT="https://<seu-endpoint-omniroute>/v1/chat/completions"
|
||||
OMNIROUTE_API_KEY="omni_..."
|
||||
OMNIROUTE_MODEL="cgpt-web/gpt-5.5"
|
||||
OMNIROUTE_TIMEOUT="20"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Execução dos Testes Automatizados (CI)
|
||||
|
||||
A suíte de testes executa 100% offline em ambiente de CI via simulações controladas:
|
||||
|
||||
```bash
|
||||
# Executar testes unitários e de integração da feature
|
||||
pytest tests/unit/test_media_classifier.py tests/integration/test_media_routing.py -v
|
||||
|
||||
# Validar conformidade com a política Zero-Regex
|
||||
python tests/scripts/check_zero_regex.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Cenários de Validação Manual / E2E
|
||||
|
||||
### Cenário 1: Execução Padrão com Arquivo de Notícias
|
||||
```bash
|
||||
python scripts/extract_article_contents.py -i out/river_plate.json
|
||||
```
|
||||
**Resultados esperados**:
|
||||
- `out/river_plate_extracted.json` gerado contendo artigos textuais e contadores atualizados;
|
||||
- `out/river_plate_media.json` gerado contendo o envelope `{ "articles": [ ... ] }` com as mídias identificadas (ou `{"articles": []}` caso nenhuma mídia seja classificada no lote);
|
||||
- Logs no stderr exibindo as 11 métricas operacionais obrigatórias.
|
||||
|
||||
### Cenário 2: Execução com Caminho de Saída Customizado (`-o/--output`)
|
||||
```bash
|
||||
python scripts/extract_article_contents.py \
|
||||
-i out/river_plate.json \
|
||||
-o out/processados/resultado_custom.json \
|
||||
-l 5
|
||||
```
|
||||
**Resultados esperados**:
|
||||
- Saída textual em `out/processados/resultado_custom.json`;
|
||||
- Saída de mídia em `out/processados/resultado_custom_media.json` (gerado obrigatoriamente, contendo as mídias ou `{"articles": []}`).
|
||||
@@ -0,0 +1,117 @@
|
||||
# Phase 0 Research: Classificação e Roteamento de Notícias Predominantemente de Mídia
|
||||
|
||||
**Feature**: `007-media-article-routing`
|
||||
**Date**: 2026-08-24
|
||||
**Status**: Completed
|
||||
|
||||
---
|
||||
|
||||
## 1. Contexto & Objetivos da Pesquisa
|
||||
|
||||
Esta pesquisa detalha as decisões técnicas para a implementação da classificação semântica e roteamento de artigos predominantemente de mídia em [`scripts/extract_article_contents.py`](../../scripts/extract_article_contents.py), em estrita conformidade com:
|
||||
- `docs/prd_extrator_artigo_media/prd.md`
|
||||
- `docs/prd_extrator_artigo_media/adr_001.md`
|
||||
- `docs/prd_extrator_artigo_media/adr_002.md`
|
||||
- `specs/007-media-article-routing/spec.md`
|
||||
- `.specify/memory/constitution.md`
|
||||
|
||||
---
|
||||
|
||||
## 2. Decisões Técnicas Consolidadas
|
||||
|
||||
### Decisão 1: Algoritmo de Detecção Estrutural de Mídia Candidata na DOM (Zero-Regex)
|
||||
|
||||
- **Decisão**: Utilizar `BeautifulSoup` (parser `html.parser`, biblioteca já instalada e utilizada no repositório) para inspecionar a região da DOM associada ao conteúdo da publicação antes de qualquer chamada LLM ou execução do multimotor textual.
|
||||
- **Algoritmo Estrutural de Localização**:
|
||||
1. **Delimitação da Região da Publicação**:
|
||||
- Localizar nós semânticos de conteúdo editorial na seguinte ordem de precedência: tag `<article>`, tag `<main>`, tag `<div role="main">`, ou tag `<body>` caso nenhuma anterior exista.
|
||||
- Isolar a análise dessa região, descartando elementos estruturais fora do conteúdo da matéria (`<header>`, `<nav>`, `<footer>`, `<aside>`).
|
||||
- Sem regex, sem heurísticas de classes CSS, sem seletores específicos de sites e sem detector de anúncios.
|
||||
2. **Identificação de Mídia Candidata Relevante**:
|
||||
- **Vídeo**: tag `<video>` $\rightarrow$ `has_video = True`. Tags `<source>` pertencentes a `<video>` não são contadas isoladamente.
|
||||
- **Imagem**: contagem de tags `<img>` reais na região editorial. Wrappers como `<figure>` e `<picture>` não incrementam nem duplicam a contagem.
|
||||
- **Embed**: tags `<iframe>`, `<embed>`, `<object>` $\rightarrow$ `has_embed = True`. Sem listas de domínios externos.
|
||||
- **Múltiplas Imagens**: contagem $\ge 2$ de tags `<img>` na região editorial da matéria.
|
||||
3. **Resultado Estrutural**:
|
||||
- Retorna objeto tipado `MediaCandidateInfo(has_candidate_media: bool, has_video: bool, image_count: int, has_embed: bool)`.
|
||||
- Se `has_candidate_media == False`: o artigo segue **diretamente** para `extract_all_engines()` sem qualquer chamada a LLM.
|
||||
- Se `has_candidate_media == True`: constrói o payload compacto e invoca a cadeia sequencial de classificação LLM.
|
||||
- **Rationale**: Filtro prévio determinístico e de custo zero de LLM, sem regex, sem dicionários por idioma e sem seletores amarrados a domínios específicos.
|
||||
|
||||
---
|
||||
|
||||
### Decisão 2: Especificação do Payload Compacto Enviado ao Classificador LLM
|
||||
|
||||
- **Decisão**: Extrair da DOM carregada uma estrutura lógica compacta com os dados essenciais para a decisão semântica:
|
||||
- `title`: Título da matéria extraído da tag `<title>` ou `<h1>` da região editorial.
|
||||
- `text_content`: Textos dos parágrafos (`<p>`) da região editorial normalizados e limpos de marcação HTML, preservando todo o conteúdo jornalístico substancial da matéria sem truncamentos arbitrários de caracteres.
|
||||
- `media_summary`: Indicadores estruturais identificados no gate (`has_video: bool`, `image_count: int`, `has_embed: bool`).
|
||||
- **Formato do Prompt Único Multilíngue**:
|
||||
```text
|
||||
You are an editorial news classifier. Classify if this news publication is predominantly media or substantive journalistic text.
|
||||
|
||||
Publication Title: {title}
|
||||
Structural Media Present: Video={has_video}, ImagesCount={image_count}, Embed={has_embed}
|
||||
Text Content:
|
||||
{text_content}
|
||||
|
||||
Definitions:
|
||||
- "media": The primary informative content is in the media (video, single image, multiple images/gallery, social embed, or mixed), and the text functions essentially as a brief introduction, caption, contextualization, or description.
|
||||
- "text": The publication contains substantive journalistic text on its own, even if accompanied by illustrative media.
|
||||
|
||||
Respond ONLY with a JSON object matching this exact schema:
|
||||
{"content_type": "text" | "media", "media_type": "video" | "image" | "images" | "embed" | "mixed" | null}
|
||||
Rules:
|
||||
- If content_type is "text", media_type MUST be null.
|
||||
- If content_type is "media", media_type MUST be one of: "video", "image", "images", "embed", "mixed".
|
||||
```
|
||||
- **Rationale**: Payload enxuto sem HTML desnecessário, permitindo decisão semântica precisa pelo modelo.
|
||||
|
||||
---
|
||||
|
||||
### Decisão 3: Cliente HTTP e Cadeia Sequencial de Provedores LLM
|
||||
|
||||
- **Decisão de Cliente HTTP**: Utilizar funções diretas com a biblioteca padrão do Python (`urllib.request` / `urllib.error` / `json`), que já é o padrão estabelecido no repositório, sem adicionar novas dependências ao projeto.
|
||||
- **Reutilização de `src/tools/adapters/llm.py`**: O módulo `llm.py` existente é altamente especializado na desambiguação de entidades ECP da feature 001 com classes acopladas (`ECPSnapshot`, `DecisionCategory`); acoplá-lo à classificação de mídia criaria emaranhamento desnecessário de domínios. A implementação com funções diretas via `urllib.request` em `scripts/extract_article_contents.py` é a solução mais desacoplada, limpa e com menor diff.
|
||||
- **Configuração Externa dos Provedores**:
|
||||
1. **Ollama (Primário)**:
|
||||
- Endpoint: `os.environ.get("OLLAMA_ENDPOINT", "http://localhost:11434")`
|
||||
- Modelo: `os.environ.get("OLLAMA_MODEL", "qwen3.5:2b")`
|
||||
- Timeout: `int(os.environ.get("OLLAMA_TIMEOUT", "10"))`
|
||||
- Structured Output: passa o JSON Schema diretamente no campo `format` da requisição `/api/chat` com `options: {"temperature": 0.0}` e `think: false` para desabilitar explicitamente o thinking no Qwen3.5 2B.
|
||||
2. **Groq (1º Fallback)**:
|
||||
- Endpoint: `os.environ.get("GROQ_ENDPOINT", "https://api.groq.com/openai/v1/chat/completions")`
|
||||
- API Key: `os.environ.get("GROQ_API_KEY")`
|
||||
- Modelo: `os.environ.get("GROQ_MODEL", "openai/gpt-oss-20b")`
|
||||
- Timeout: `int(os.environ.get("GROQ_TIMEOUT", "15"))`
|
||||
- Structured Output: `response_format={"type": "json_schema", "json_schema": {"name": "media_classifier", "strict": True, "schema": <schema>}}`, `temperature: 0.0` e `reasoning_effort: "low"`.
|
||||
3. **OmniRoute (2º Fallback)**:
|
||||
- Endpoint: `os.environ.get("OMNIROUTE_ENDPOINT")` (configuração externa obrigatória no ambiente sem default inventado)
|
||||
- API Key: `os.environ.get("OMNIROUTE_API_KEY")`
|
||||
- Modelo: `os.environ.get("OMNIROUTE_MODEL", "cgpt-web/gpt-5.5")`
|
||||
- Timeout: `int(os.environ.get("OMNIROUTE_TIMEOUT", "20"))`
|
||||
- Structured Output: `response_format` com JSON Schema compatível com a instalação OpenAI-compatible utilizada, `temperature: 0.0` quando suportado, sem parâmetro inventado de reasoning.
|
||||
- **Regras de Execução e Fallback**:
|
||||
- Execução estritamente sequencial. Toda resposta que validar contra o schema de 2 campos encerra a cadeia com sucesso.
|
||||
- Falha operacional (timeout, recusa de conexão, erro HTTP ou schema inválido) transita imediatamente para o próximo provedor.
|
||||
- Se todos falharem: atribui `classification_status = "failed"` e mensagem diagnóstica, registrando o artigo inline no JSON principal sem multimotor e sem interromper o lote.
|
||||
|
||||
---
|
||||
|
||||
### Decisão 4: Estrutura de Arquivos e Semântica de Contadores
|
||||
|
||||
- **Saída de Mídia (`*_media.json`)**:
|
||||
Envelope físico contendo estritamente `{ "articles": [ MediaArticle... ] }`.
|
||||
- **Saída Textual (`*_extracted.json`)**:
|
||||
Envelope existente `ExtractionBatchReport` onde `total_articles`, `successful_articles` e `failed_articles` refletem exclusivamente os itens persistidos nesse arquivo (artigos textuais + registros com `classification_status = "failed"`).
|
||||
- **Métricas de Execução**: As 11 métricas obrigatórias são contabilizadas em memória e registradas no log do stderr ao final do lote (respeitando a flag `--silent`).
|
||||
|
||||
---
|
||||
|
||||
### Decisão 5: Estratégia de Testes e Zero-Regex
|
||||
|
||||
- **Testes Determinísticos (CI)**:
|
||||
- Testes unitários (`tests/unit/test_media_classifier.py`) cobrindo gate estrutural DOM, montagem de compact payload e simulação dos 5 estados da cadeia de provedores sem chamadas de rede externas.
|
||||
- Testes de integração (`tests/integration/test_media_routing.py`) cobrindo o fluxo em lote completo, cenários A a M, roteamento com e sem `-o` e contadores.
|
||||
- **Verificação Zero-Regex**:
|
||||
- Reutilização do script existente [`tests/scripts/check_zero_regex.py`](../../tests/scripts/check_zero_regex.py) incluindo os arquivos de teste e módulos da feature no escopo de verificação AST.
|
||||
@@ -0,0 +1,221 @@
|
||||
# Feature Specification: Classificação e Roteamento de Notícias Predominantemente de Mídia
|
||||
|
||||
**Feature Branch**: `007-media-article-routing`
|
||||
|
||||
**Created**: 2026-08-24
|
||||
|
||||
**Status**: Draft
|
||||
|
||||
**Input**: User description: "Classificação e Roteamento de Notícias Predominantemente de Mídia no script scripts/extract_article_contents.py conforme docs/prd_extrator_artigo_media (prd.md, adr_001.md, adr_002.md)"
|
||||
|
||||
---
|
||||
|
||||
## User Scenarios & Testing *(mandatory)*
|
||||
|
||||
### User Story 1 - Roteamento Exclusivo de Notícias Predominantemente de Mídia (Priority: P1)
|
||||
|
||||
Como operador do pipeline de notícias, quero que publicações jornalísticas cujo conteúdo informativo principal seja uma mídia (vídeo, imagem única, coleção/múltiplas imagens, post incorporado/embed ou combinação mista de mídias) e cujo texto atue apenas como introdução, legenda, contextualização ou breve descrição dessa mídia sejam identificadas logo após o crawl e salvas em um arquivo JSON próprio (`*_media.json`), sem passar pelo pipeline multimotor textual (Trafilatura, Newspaper4k e Readability).
|
||||
|
||||
**Why this priority**: É o objetivo central da funcionalidade: evitar processamento desnecessário de páginas não textuais pelo multimotor e separar fisicamente as publicações de mídia dos artigos textuais.
|
||||
|
||||
**Independent Test**: Pode ser testado de forma isolada submetendo URLs cujas páginas possuam mídia predominante e texto meramente descritivo/introdutório. O sistema deve gerar o arquivo `*_media.json` com os registros correspondentes e seus metadados mínimos de rastreabilidade, sem disparar qualquer chamada aos três extratores textuais.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Cenário B (Vídeo)**: **Given** uma página carregada contendo um vídeo e texto curto apenas introdutório/contextual, **When** a detecção estrutural e o classificador processam a publicação, **Then** o classificador retorna `content_type = "media"` e `media_type = "video"`, o multimotor não é executado e o artigo é gravado em `*_media.json`.
|
||||
2. **Cenário C (Imagem Única)**: **Given** uma página carregada contendo uma única imagem e texto curto descritivo, **When** o classificador processa a publicação, **Then** o classificador retorna `content_type = "media"` e `media_type = "image"`, o multimotor não é executado e o artigo é gravado em `*_media.json`.
|
||||
3. **Cenário D (Múltiplas Imagens)**: **Given** uma página carregada contendo múltiplas imagens (em galeria, carousel, slideshow ou sequência vertical) e texto curto, **When** o classificador processa a publicação, **Then** o classificador retorna `content_type = "media"` e `media_type = "images"`, o multimotor não é executado e o artigo é gravado em `*_media.json`.
|
||||
4. **Cenário E (Conteúdo Incorporado / Embed)**: **Given** uma página carregada contendo um post incorporado (ex: Instagram, TikTok, X) e breve contextualização textual, **When** o classificador processa a publicação, **Then** o classificador retorna `content_type = "media"` e `media_type = "embed"`, o multimotor não é executado e o artigo é gravado em `*_media.json`.
|
||||
5. **Cenário F (Mídia Mista)**: **Given** uma página carregada contendo mais de uma categoria relevante de mídia (ex: vídeo e imagens) com texto meramente introdutório, **When** o classificador processa a publicação, **Then** o classificador retorna `content_type = "media"` e `media_type = "mixed"`, o multimotor não é executado e o artigo é gravado em `*_media.json`.
|
||||
|
||||
---
|
||||
|
||||
### User Story 2 - Roteamento Direto e Preservação de Notícias Textuais (Priority: P2)
|
||||
|
||||
Como operador do pipeline textual, quero que notícias cujo conteúdo principal seja texto jornalístico substancial (mesmo que acompanhadas de fotos ilustrativas, vídeos ou infográficos) ou páginas sem qualquer elemento estrutural de mídia continuem sendo processadas normalmente pelos três motores de extração (`extract_all_engines`) e salvas no arquivo de saída textual (`*_extracted.json` ou caminho informado em `-o/--output`), preservando integralmente o formato de saída atual e a semântica de seus contadores.
|
||||
|
||||
**Why this priority**: Garante que o pipeline textual original não sofra quebra de contrato, regressão ou perda de dados em matérias jornalísticas informativas normais.
|
||||
|
||||
**Independent Test**: Pode ser testado submetendo: (a) páginas de texto puro sem mídia, e (b) matérias jornalísticas substanciais de múltiplos parágrafos contendo fotos ou vídeos editoriais. Ambos devem ser encaminhados ao multimotor e consolidados no arquivo de saída textual.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Cenário A (Texto sem Mídia)**: **Given** uma página carregada sem elementos estruturais candidatos a mídia na DOM, **When** a detecção estrutural prévia é executada, **Then** o classificador LLM não é chamado e o artigo é encaminhado diretamente ao multimotor textual.
|
||||
2. **Cenário G (Artigo Textual Longo com Imagem)**: **Given** uma notícia com texto jornalístico substancial contendo uma ou mais imagens ilustrativas, **When** o classificador semântico avalia a publicação, **Then** retorna `content_type = "text"` com `media_type = null`, o artigo é processado pelo multimotor e salvo na saída textual.
|
||||
3. **Cenário H (Artigo Textual Longo com Vídeo)**: **Given** uma notícia com texto jornalístico substancial contendo um vídeo incorporado, **When** o classificador semântico avalia a publicação, **Then** retorna `content_type = "text"` com `media_type = null`, o artigo é processado pelo multimotor e salvo na saída textual.
|
||||
|
||||
---
|
||||
|
||||
### User Story 3 - Resiliência com Fallback Operacional Sequencial e Registro de Falhas (Priority: P3)
|
||||
|
||||
Como operador do sistema, quero que falhas puramente operacionais/técnicas no provedor primário local (Ollama) acionem sequencialmente o primeiro fallback (Groq) e, se este também falhar operacionalmente, o segundo fallback (OmniRoute), e que em caso de indisponibilidade de todos os provedores, o erro seja registrado explicitamente no JSON principal de processamento sem interromper o lote.
|
||||
|
||||
**Why this priority**: Assegura resiliência de produção em execuções de lote sem intervenção manual, mantendo rastreabilidade estrita e isolamento de falhas.
|
||||
|
||||
**Independent Test**: Pode ser testado simulando deterministicamente falhas operacionais e de contrato (timeout, recusa de conexão, erro HTTP, schema inválido) nos provedores intermediários e validando a transição sequencial e o comportamento de falha total, sem necessidade de chamadas a provedores reais durante a suíte normal de CI.
|
||||
|
||||
**Acceptance Scenarios**:
|
||||
|
||||
1. **Cenário I (Falha no Ollama)**: **Given** o provedor primário (Qwen3.5 2B / Ollama) apresentando falha operacional (timeout, erro HTTP ou conexão recusada), **When** um artigo com mídia candidata é classificado, **Then** o sistema aciona o provedor GPT-OSS 20B / Groq e utiliza sua classificação válida para rotear o artigo.
|
||||
2. **Cenário J (Falha no Ollama e Groq)**: **Given** Ollama e Groq apresentando falha operacional, **When** o artigo é classificado, **Then** o sistema aciona o provedor `cgpt-web/gpt-5.5` / OmniRoute e utiliza sua classificação válida para rotear o artigo.
|
||||
3. **Cenário K (Falha Total de Todos os Provedores)**: **Given** Ollama, Groq e OmniRoute apresentando falhas operacionais consecutivas, **When** o artigo é processado, **Then** o sistema atribui `classification_status = "failed"` e mensagem de erro diagnóstica, não assume presunção arbitrária de `text` nem de `media`, não executa o multimotor, não envia o registro para `*_media.json`, registra a falha no arquivo JSON principal de processamento e continua processando os demais artigos do lote.
|
||||
4. **Cenário L (Resposta fora do Schema)**: **Given** um provedor retornando resposta ilegível, truncada ou incompatível com o contrato estruturado obrigatório, **When** a validação de contrato é executada, **Then** a tentativa é tratada como falha operacional do provedor atual e o sistema avança imediatamente para o próximo provedor na cadeia de fallback.
|
||||
5. **Cenário M (Suporte Multilíngue)**: **Given** páginas de notícias redigidas em qualquer um dos até 10 idiomas suportados pelo sistema (ex: espanhol, inglês, português, francês, alemão, italiano, etc.), **When** a detecção e a classificação são executadas, **Then** o sistema classifica corretamente textos curtos com mídia como `media` e textos substanciais com mídia como `text`, sem utilizar regras ou prompts específicos por idioma.
|
||||
|
||||
---
|
||||
|
||||
### Edge Cases
|
||||
|
||||
- **Texto extremamente curto com uma única imagem (notícia curta com imagem)**: Deve ser deliberadamente classificado como `content_type = "media"` e `media_type = "image"`, pois texto com volume insuficiente não é útil para o pipeline textual.
|
||||
- **Avaliação de brevidade textual ("3 a 5 linhas")**: A referência de 3 a 5 linhas de texto é exclusivamente conceitual. O modelo deve avaliar semanticamente se o texto possui conteúdo jornalístico substancial por si próprio ou se funciona apenas como introdução/descrição da mídia, sem depender de contagem literal de linhas renderizadas, viewport, resolução, CSS, número de palavras ou caracteres.
|
||||
- **Coleções de imagens (galerias / carousels / slideshows / sequência vertical)**: O sistema não deve tentar interagir com o carousel, clicar em botões, avançar slides ou extrair URLs das imagens. Deve apenas identificar estruturalmente a presença de múltiplas imagens e classificar como `media_type = "images"`.
|
||||
- **Conteúdo misto sem precedência artificial**: Quando houver mais de um tipo relevante de mídia atuando como elemento informativo principal (ex: vídeo e galeria de fotos), o sistema deve classificar como `media_type = "mixed"`, sem impor regras artificiais de precedência como `video > image`.
|
||||
- **Falha isolada por artigo**: A falha na classificação ou no processamento de um artigo específico não pode interromper nem abortar a execução do lote.
|
||||
|
||||
---
|
||||
|
||||
## Requirements *(mandatory)*
|
||||
|
||||
### Functional Requirements
|
||||
|
||||
#### Detecção Estrutural Prévia
|
||||
- **FR-001**: O sistema MUST analisar estruturalmente a DOM carregada pelo crawler antes de qualquer chamada aos motores Trafilatura, Newspaper4k e Readability.
|
||||
- **FR-002**: A análise estrutural da DOM MUST ser realizada exclusivamente via parser HTML/DOM (navegação por nós, tags, atributos e contagem de elementos), sendo estritamente proibido o uso de regex (`re`), listas de palavras-chave ou heurísticas semânticas por idioma em toda a nova implementação.
|
||||
- **FR-003**: Quando nenhuma mídia candidata estiver presente como elemento estruturalmente relevante ao conteúdo da publicação na DOM carregada (sem imagem, vídeo, iframe/embed, object ou estruturas DOM equivalentes associadas à publicação), o artigo MUST seguir diretamente para o pipeline multimotor textual sem chamada ao classificador LLM.
|
||||
- **FR-004**: Quando houver mídia candidata presente e estruturalmente relevante ao conteúdo da publicação na DOM carregada (como imagem, vídeo, iframe/embed, object ou estruturas equivalentes associadas à publicação), o sistema MUST submeter uma representação compacta da publicação ao classificador de conteúdo. A mera presença de mídia em outras regiões da página não associadas ao conteúdo da publicação não deve, isoladamente, tornar o artigo candidato.
|
||||
|
||||
#### Entrada e Escopo do Classificador
|
||||
- **FR-005**: O classificador MUST receber apenas dados textuais e estruturais compactos já disponíveis na página carregada (título, blocos textuais relevantes, informação estrutural de mídias presentes e quantidade/tipos encontrados), evitando o envio do HTML completo quando a representação compacta contiver a mesma informação.
|
||||
- **FR-006**: O sistema MUST NOT baixar imagens, baixar vídeos, executar OCR, executar visão computacional, assistir vídeos, transcrever áudios, navegar em carousels/slideshows nem chamar APIs externas de plataformas de mídia.
|
||||
|
||||
#### Contrato Estruturado do Classificador
|
||||
- **FR-007**: A saída emitida pelo classificador LLM MUST conter exclusivamente os campos `content_type` e `media_type` em formato JSON estruturado (sem campos como `provider_used`, `status`, `error_message`, `confidence`, `rationale`, `summary`, `keywords`, `evidence` ou `tradução`).
|
||||
- **FR-008**: O campo `content_type` emitido pelo modelo MUST aceitar exclusivamente os valores `"text"` ou `"media"`.
|
||||
- **FR-009**: O campo `media_type` emitido pelo modelo MUST aceitar exclusivamente os valores `"video"`, `"image"`, `"images"`, `"embed"`, `"mixed"` quando `content_type = "media"`, e MUST ser obrigatoriamente `null` quando `content_type = "text"`.
|
||||
- **FR-010**: O prompt enviado ao classificador MUST ser único, conciso, comum a todos os idiomas e solicitar exclusivamente a classificação necessária (`content_type` e `media_type`). O prompt MUST NOT solicitar reasoning, confidence, rationale, summary, keywords, evidence ou tradução. Reasoning deliberativo adicional não deve ser habilitado quando não for necessário para produzir o contrato estruturado da classificação, e novos campos não devem ser adicionados à resposta.
|
||||
- **FR-011**: Sempre que suportado pelo provedor, a requisição ao modelo MUST utilizar structured output nativo (JSON Schema / response_format). Toda resposta MUST ser estritamente validada contra o contrato antes de ser aceita.
|
||||
|
||||
#### Roteamento, Persistência e Regras de Arquivos
|
||||
- **FR-012**: Artigos classificados como `content_type = "text"` MUST seguir normalmente pelo fluxo textual existente, executando os três motores de extração (`extract_all_engines`) e sendo salvos no arquivo JSON de saída textual.
|
||||
- **FR-013**: Artigos classificados como `content_type = "media"` MUST NOT executar os motores Trafilatura, Newspaper4k ou Readability.
|
||||
- **FR-014**: Artigos classificados como `content_type = "media"` MUST ser gravados em arquivo JSON de mídia dedicado com sufixo `_media.json`, obedecendo às seguintes regras de nomenclatura e diretório:
|
||||
- **Sem `-o/--output`**: para uma entrada `<dir>/<stem>.json` (ex: `out/river_plate.json`), a saída textual padrão é `<dir>/<stem>_extracted.json` (`out/river_plate_extracted.json`) e a saída de mídia é `<dir>/<stem>_media.json` (`out/river_plate_media.json`).
|
||||
- **Com `-o/--output` customizado**: para uma saída textual informada como `<custom_dir>/<custom_stem>.json` (ex: `-o out/processados/brasil_completo.json`), a saída de mídia MUST ser gravada em `<custom_dir>/<custom_stem>_media.json` (`out/processados/brasil_completo_media.json`), utilizando o mesmo diretório e stem do output textual com sufixo `_media` e extensão `.json`.
|
||||
- O arquivo `*_media.json` MUST ser gerado em toda execução bem-sucedida do lote. Quando nenhum artigo for classificado como `content_type = "media"`, o arquivo MUST conter exatamente `{"articles": []}`. O sistema MUST NOT adotar comportamento condicional de omissão do arquivo.
|
||||
- A interface CLI existente MUST permanecer inalterada, sem adição de parâmetros como `--media-output`.
|
||||
- **FR-015**: O arquivo `*_media.json` MUST utilizar um envelope físico mínimo contendo exclusivamente a chave `"articles"`, onde cada elemento representa um registro de `MediaArticle` com seus campos mínimos obrigatórios de identificação e rastreabilidade:
|
||||
```json
|
||||
{
|
||||
"articles": [
|
||||
{
|
||||
"input_meta": {
|
||||
"titulo": "...",
|
||||
"url": "..."
|
||||
},
|
||||
"crawled_url": "...",
|
||||
"page_title": "...",
|
||||
"http_status": 200,
|
||||
"content_type": "media",
|
||||
"media_type": "video"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
Quando nenhum artigo de mídia for identificado no lote, o envelope MUST ser emitido exatamente como `{"articles": []}`. `input_meta` preserva os metadados recebidos da entrada (onde `titulo` e `url` são obrigatórios, e `subtitulo`, `quando_publicado`, `pagina` e demais metadados existentes continuam opcionais conforme a entrada). O envelope de mídia MUST NOT conter `total_articles`, `successful_articles`, `failed_articles`, contadores por `media_type`, status agregados, modelo de relatório `MediaBatchReport` ou métricas duplicadas (as métricas operacionais são expostas exclusivamente via logging/resumo de execução conforme FR-028).
|
||||
- **FR-016**: O sistema MUST NOT adicionar em `*_media.json` campos de extração aprofundada de mídia como: `image_urls`, `video_urls`, `embed_urls`, `thumbnails`, `duration`, `captions`, `transcripts`, `OCR`, `alt-text gerado`, `descrição da imagem`, `provider da mídia`, `metadata da mídia`, `summary`, `keywords` ou `sentiment`.
|
||||
- **FR-017**: O fluxo textual existente (`*_extracted.json` ou arquivo informado em `-o/--output`) MUST preservar o envelope de saída existente (`source_file`, `processed_at`, `total_articles`, `successful_articles`, `failed_articles`, `articles`). Como os artigos de mídia são desviados para `*_media.json` e não permanecem no array `articles` do JSON principal, a semântica dos contadores MUST representar estritamente os registros presentes no arquivo JSON principal:
|
||||
- `total_articles`: quantidade total de registros presentes no array `articles` do JSON principal;
|
||||
- `successful_articles`: quantidade de artigos textuais processados com sucesso pelo multimotor e presentes no JSON principal;
|
||||
- `failed_articles`: quantidade de registros de falha presentes no JSON principal (incluindo falhas do crawl ou com `classification_status = "failed"`).
|
||||
O JSON principal MUST NOT conter campos adicionais como `media_articles`, contadores de artigos desviados, referências ao `*_media.json` ou novos campos agregados de relatório.
|
||||
|
||||
#### Cadeia Sequencial de Provedores e Fallback Operacional
|
||||
- **FR-018**: O sistema MUST utilizar Qwen3.5 2B executado localmente via Ollama como classificador primário determinístico.
|
||||
- **FR-019**: O sistema MUST acionar o primeiro fallback operacional (GPT-OSS 20B via Groq) estritamente em caso de falha operacional do Ollama (timeout, erro HTTP, indisponibilidade ou resposta fora do schema).
|
||||
- **FR-020**: O sistema MUST acionar o segundo fallback operacional (`cgpt-web/gpt-5.5` via OmniRoute) estritamente em caso de falha operacional cumulativa do Ollama e do Groq.
|
||||
- **FR-021**: Os provedores de classificação MUST ser executados de forma estritamente sequencial (uma resposta válida encerra imediatamente a cadeia). É expressamente proibida a execução paralela, votação entre modelos, consenso, segunda opinião ou fallback condicionado a valor de classificação ou nível de confiança.
|
||||
- **FR-022**: Em caso de falha operacional dos três provedores (Ollama, Groq e OmniRoute), o artigo MUST receber `classification_status = "failed"` e mensagem diagnóstica de erro sem exposição de credenciais, MUST permanecer registrado no JSON principal de processamento, MUST NOT ser enviado ao arquivo de mídia, MUST NOT executar o multimotor e MUST NOT ter classificação presumida como `text` ou `media`. A falha não deve interromper os demais artigos do lote.
|
||||
- **FR-023**: O sistema MUST NOT criar arquivos físicos adicionais exclusivos para falhas (ex: `*_failed.json`), filas de mensagens, DLQ ou serviços assíncronos de recuperação.
|
||||
- **FR-024**: A cadeia de provedores de classificação MUST ser testável de forma determinística sem depender de chamadas reais aos provedores externos durante a suíte normal automatizada (CI). Devem ser simuláveis/testáveis, no mínimo: (1) Ollama success; (2) Ollama fail → Groq success; (3) Ollama fail → Groq fail → OmniRoute success; (4) todos falham; e (5) provedor retorna resposta fora do schema. Não é exigida presença de Ollama, Groq ou OmniRoute reais na execução regular da suíte de testes.
|
||||
|
||||
#### Segurança, Configuração Externa e Observabilidade
|
||||
- **FR-025**: Endpoints, nomes de modelos, credenciais de API e timeouts dos provedores MUST ser configuráveis externamente ao código-fonte, não sendo permitidos segredos hardcoded.
|
||||
- **FR-026**: Logs, saídas JSON e mensagens de erro persistidas MUST NOT registrar API keys, tokens de autenticação ou credenciais.
|
||||
- **FR-027**: O sistema MUST registrar informações operacionais estruturadas nos logs existentes para identificar: provedor utilizado, acionamento do fallback Ollama → Groq, acionamento do fallback Groq → OmniRoute, classificação final obtida e eventual falha total da cadeia, sem registrar o payload textual integral do artigo por padrão.
|
||||
- **FR-028**: O sistema MUST contabilizar e expor nos mecanismos de logging/resumo de execução existentes as seguintes métricas operacionais obrigatórias:
|
||||
- total de artigos avaliados;
|
||||
- total roteado para `text`;
|
||||
- total roteado para `media`;
|
||||
- total `media/video`;
|
||||
- total `media/image`;
|
||||
- total `media/images`;
|
||||
- total `media/embed`;
|
||||
- total `media/mixed`;
|
||||
- quantidade de acionamentos do fallback Groq;
|
||||
- quantidade de acionamentos do fallback OmniRoute;
|
||||
- total de falhas de classificação.
|
||||
- **FR-029**: O classificador e a detecção estrutural MUST funcionar de forma independente de idioma, cobrindo com precisão os até 10 idiomas processados pelo sistema.
|
||||
- **FR-030**: O script `scripts/extract_article_contents.py` MUST manter compatibilidade com sua interface de linha de comando (`-i/--input`, `-o/--output`, `-l/--limit`, `--lang/--language`, `-t/--timeout`, `-s/--silent`).
|
||||
|
||||
---
|
||||
|
||||
### Key Entities
|
||||
|
||||
- **InputArticle**: Metadados originais da notícia de entrada carregados do arquivo de busca (`titulo` e `url` obrigatórios; `subtitulo`, `quando_publicado`, `pagina` e eventuais campos extras preservados conforme a entrada).
|
||||
- **MediaArticle**: Registro consolidado de publicação predominantemente de mídia gravado no array `articles` de `*_media.json`, contendo os campos mínimos obrigatórios de rastreabilidade (`input_meta`, `crawled_url`, `page_title`, `http_status`, `content_type` e `media_type`).
|
||||
- **ExtractedArticle**: Registro consolidado de artigo textual efetivamente processado pelos motores de extração multimotor (`extract_all_engines`) e gravado no array `articles` do JSON principal de saída textual, contendo `input_meta`, `extraction_status`, `error_message`, `crawled_url`, `page_title`, `http_status`, `trafilatura`, `newspaper4k` e `readability`.
|
||||
- **ExtractionBatchReport**: Relatório consolidado do lote de extração textual gravado no arquivo JSON principal de saída contendo os contadores restritos aos itens presentes nesse arquivo (`total_articles`, `successful_articles`, `failed_articles`) e a lista `articles`.
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria *(mandatory)*
|
||||
|
||||
### Measurable Outcomes
|
||||
|
||||
- **SC-001**: 100% das publicações identificadas como predominantemente de mídia (`video`, `image`, `images`, `embed`, `mixed`) são desviadas antes da execução dos motores Trafilatura, Newspaper4k e Readability.
|
||||
- **SC-002**: 100% das publicações de mídia identificadas são gravadas no arquivo `*_media.json` dentro do envelope `articles` contendo os campos mínimos obrigatórios definidos no contrato de rastreabilidade (`input_meta`, `crawled_url`, `page_title`, `http_status`, `content_type`, `media_type`).
|
||||
- **SC-003**: 0% de alteração de contrato ou degradação de campos nos arquivos JSON de saída textual para artigos normais processados com sucesso, com contadores refletindo estritamente os itens presentes no arquivo.
|
||||
- **SC-004**: 100% das falhas operacionais intermediárias em provedores LLM transitam de forma determinística e sequencial pela cadeia (Ollama → Groq → OmniRoute) sem interrupção manual ou falha espúria.
|
||||
- **SC-005**: 100% das tentativas com falha cumulativa dos três provedores resultam em registro de erro explícito (`classification_status = "failed"`) no arquivo JSON principal de processamento sem parar o processamento dos demais itens do lote.
|
||||
- **SC-006**: 100% dos cenários mínimos e requisitos funcionais obrigatórios possuem cobertura em suíte de testes automatizados (unitários e de integração), contemplando páginas sem mídia, mídias individuais, mídias mistas, fallbacks operacionais simulados deterministicamente e artigos em todos os idiomas suportados pelo sistema.
|
||||
- **SC-007**: 0% de uso de bibliotecas ou chamadas de expressões regulares (`re`) em toda a nova lógica de detecção estrutural e classificação de mídia.
|
||||
- **SC-008**: 0 ocorrências de credenciais, API keys ou tokens expostos em logs, mensagens de erro ou saídas JSON.
|
||||
|
||||
---
|
||||
|
||||
## Assumptions
|
||||
|
||||
- **Configuração de Provedores**: O ambiente de produção deve possuir configuração para o runtime local do Ollama (Qwen3.5 2B) e credenciais para os provedores de fallback Groq (GPT-OSS 20B) e OmniRoute (`cgpt-web/gpt-5.5`). A indisponibilidade operacional de um provedor durante a execução aciona sequencialmente o fallback definido.
|
||||
- **Isolamento de Responsabilidade da Mídia**: O download, visão computacional, OCR, transcrição e indexação da mídia são responsabilidades exclusivas de sistemas ou etapas downstream, não fazendo parte desta funcionalidade.
|
||||
- **Decisão Semântica de Conteúdo**: A distinção entre texto substancial e texto curto introdutório/contextual é uma avaliação semântica realizada pelo modelo de linguagem compacto, não sendo calculada por contagem de linhas visuais renderizadas ou regras lexicais.
|
||||
- **Isolamento de Falha**: A falha operacional na classificação de uma notícia isolada não invalida nem interrompe a execução dos demais artigos presentes no mesmo lote.
|
||||
|
||||
---
|
||||
|
||||
## Out of Scope *(Explicitamente Fora de Escopo)*
|
||||
|
||||
Para evitar overengineering e manter a solução estritamente aderente aos princípios obrigatórios, os seguintes itens estão **explicitamente fora de escopo**:
|
||||
|
||||
- Download de imagens, vídeos ou arquivos binários de mídia;
|
||||
- Extração de URLs internas de imagens, vídeos ou embeds;
|
||||
- OCR (Reconhecimento Óptico de Caracteres) ou visão computacional;
|
||||
- Transcrição de áudio ou vídeo;
|
||||
- Identificação de objetos, fotografias, infográficos ou ilustrações;
|
||||
- Geração de descrições, resumos, palavras-chave ou análise de sentimento;
|
||||
- Interação com carousels, slideshows ou navegação em galerias;
|
||||
- Chamadas a APIs externas dos provedores das mídias (Instagram, TikTok, X, etc.);
|
||||
- Detecção, classificação ou filtragem de publicidade / anúncios;
|
||||
- Armazenamento em banco de dados nesta feature (nenhum banco);
|
||||
- Criação de qualquer worker novo (nenhum worker novo);
|
||||
- Criação de pipeline adicional de NLP (nenhum pipeline NLP adicional);
|
||||
- Criação ou introdução de nova plataforma de observabilidade (nenhuma nova plataforma de observabilidade);
|
||||
- Criação de novos serviços, microserviços ou processos autônomos;
|
||||
- Introdução de filas de mensagens, DLQ ou mecanismos adicionais de retry em fila;
|
||||
- Uso de frameworks de agentes, LangChain, LangGraph ou LLM-as-a-judge;
|
||||
- Votação entre modelos, scoring complexo ou consenso de LLMs;
|
||||
- Classificação baseada em regex ou dicionários de palavras-chave por idioma;
|
||||
- Criação de novos modelos físicos de relatório como `MediaBatchReport`;
|
||||
- Criação de novas entidades de domínio ou DTOs para falhas (ex: `FailedArticle`, `ClassificationFailedArticle`);
|
||||
- Criação de arquivos físicos adicionais como `*_failed.json`;
|
||||
- Criação de novos parâmetros de linha de comando como `--media-output`;
|
||||
- Alterar o comportamento de carregamento/crawl existente do Foxcape/Camoufox, ou a lógica interna de Trafilatura, Newspaper4k e Readability. A nova funcionalidade consome o HTML já retornado pelo crawler e realiza a decisão de roteamento antes da execução de `extract_all_engines()`.
|
||||
@@ -0,0 +1,148 @@
|
||||
# Tasks: Classificação e Roteamento de Notícias Predominantemente de Mídia
|
||||
|
||||
**Feature**: `007-media-article-routing`
|
||||
**Input**: [spec.md](spec.md), [plan.md](plan.md), [data-model.md](data-model.md), [research.md](research.md), [contracts/](contracts/)
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Foundational (Blocking Prerequisites)
|
||||
|
||||
**Purpose**: Estruturas de dados internas, constante do schema e helper de comunicação HTTP via `urllib.request` que bloqueiam a implementação das histórias.
|
||||
|
||||
- [x] T001 Implement `MediaCandidateInfo` dataclass and `MediaClassification` type definition in `scripts/extract_article_contents.py`
|
||||
- [x] T002 Define `MEDIA_CLASSIFIER_SCHEMA` constant (flat 2-field schema matching `contracts/classifier-io.schema.json`) and implement `validate_classifier_response(data: dict) -> MediaClassification | None` in `scripts/extract_article_contents.py`
|
||||
- [x] T003 Implement `urllib.request` JSON HTTP dispatch helper `_http_post_json(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]` in `scripts/extract_article_contents.py`
|
||||
- Serializar request JSON, configurar headers e executar chamada via `urllib.request`;
|
||||
- Retornar tupla `(http_status, response_body)`;
|
||||
- Não converter erros de transporte para status mágicos (ex: 0/-1); exceções operacionais (`URLError`, `HTTPError`, timeout) devem propagar para captura e tratamento centralizado de fallback em `classify_media_content()`.
|
||||
|
||||
**Checkpoint**: Base foundational pronta — implementação das histórias de usuário pode prosseguir.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: User Story 1 - Detecção e Roteamento de Publicações Predominantemente de Mídia (Priority: P1)
|
||||
|
||||
**Goal**: Identificar publicações cujo conteúdo informativo principal seja mídia (vídeo, imagem, imagens, embed ou mídia mista), desviar antes dos extratores textuais e persistir em `*_media.json` com envelope mínimo `{ "articles": [...] }`.
|
||||
|
||||
**Independent Test**: Submeter amostras de páginas de mídia com texto introdutório curto (cenários B, C, D, E, F); verificar que `extract_all_engines` NÃO é chamado e que o arquivo `*_media.json` é gerado contendo o envelope e os registros `MediaArticle`.
|
||||
|
||||
- [x] T004 [US1] Unit tests for schema validation, DOM structural gate (`detect_candidate_media`) and compact payload generation (`build_compact_payload`) in `tests/unit/test_media_classifier.py`
|
||||
- Validar `validate_classifier_response`: `content_type=text` + `media_type=null` (válido), `content_type=text` + `media_type=image` (inválido), `content_type=media` + `media_type=video` (válido), `content_type=media` + `media_type=null` (inválido), valores desconhecidos (inválido), campo obrigatório ausente (inválido), campo extra (inválido);
|
||||
- Validar detecção de `<video>`;
|
||||
- Validar que `<source>` isolado NÃO é vídeo;
|
||||
- Validar que `<figure><img></figure>` e `<picture><img></picture>` contam exatamente 1 imagem;
|
||||
- Validar contagem exata de múltiplos `<img>`;
|
||||
- Validar detecção de `<iframe>`, `<embed>`, `<object>` como embeds;
|
||||
- Validar que mídia em `<header>`, `<nav>`, `<footer>` e `<aside>` não aciona o gate;
|
||||
- Validar ordem estrutural (`<article>` $\rightarrow$ `<main>` $\rightarrow$ `[role="main"]` $\rightarrow$ `<body>`);
|
||||
- Validar que `build_compact_payload` extrai `title`, `text_content` (todos os `<p>` editoriais normalizados sem truncamento arbitrário e preservando Unicode) e `media_summary` (`has_video`, `image_count`, `has_embed`).
|
||||
- [x] T005 [US1] Implement `detect_candidate_media(soup: BeautifulSoup) -> MediaCandidateInfo` identifying editorial region (`<article>`, `<main>`, `[role=main]`, `<body>`) and media elements (`<video>`, real `<img>` count without duplicate wrappers, `<iframe>`/`<embed>`/`<object>`) in `scripts/extract_article_contents.py`
|
||||
- [x] T006 [US1] Implement `build_compact_payload(soup: BeautifulSoup, candidate_info: MediaCandidateInfo) -> str` extracting title and normalized editorial text in `scripts/extract_article_contents.py`
|
||||
- [x] T007 [US1] Implement single prompt constant and initial `classify_media_content(payload: str, metrics: dict, silent: bool = False) -> tuple[MediaClassification | None, str | None]` in `scripts/extract_article_contents.py`
|
||||
- Definir constante de prompt único `MEDIA_CLASSIFIER_PROMPT` em `scripts/extract_article_contents.py` reutilizada por Ollama, Groq e OmniRoute, comum a todos os idiomas, solicitando exclusivamente `content_type` e `media_type`, contendo definições semânticas de text vs media, e sem solicitar reasoning, confidence, rationale, summary, keywords, evidence ou tradução;
|
||||
- Despachar provedor primário Ollama (`OLLAMA_ENDPOINT`, `OLLAMA_MODEL="qwen3.5:2b"`, `OLLAMA_TIMEOUT=10`, `think=false`, `temperature=0.0`, `format=MEDIA_CLASSIFIER_SCHEMA`) e validar schema da resposta;
|
||||
- Registrar no stderr o provedor utilizado e a classificação final (`text` ou `media` + `media_type`) respeitando `silent`, sem logar o payload textual integral.
|
||||
- [x] T008 [US1] Implement media output path resolution, metrics dict initialization, `save_media_json`, and batch routing in `scripts/extract_article_contents.py`
|
||||
- Inicializar em `process_batch` o dicionário local com as 11 métricas operacionais zeradas;
|
||||
- Implementar resolução do caminho do arquivo de mídia: sem `-o` $\rightarrow$ `<input_dir>/<input_stem>_media.json` (ex: `out/river_plate.json` $\rightarrow$ `out/river_plate_media.json`); com `-o` $\rightarrow$ `<output_dir>/<output_stem>_media.json` (ex: `-o out/processados/resultado.json` $\rightarrow$ `out/processados/resultado_media.json`) sem novas flags CLI;
|
||||
- Implementar `save_media_json(articles: list[dict[str, Any]], output_path: Path) -> None` gravando `{ "articles": [...] }`;
|
||||
- Integrar o fluxo em `process_batch`: após crawl bem-sucedido, carregar HTML no BeautifulSoup e executar `detect_candidate_media()`; se `has_candidate_media == True`, executar `build_compact_payload()` e `classify_media_content(payload, metrics, silent)`; se retornar `content_type == "media"`, adicionar `MediaArticle` em `media_articles` e NÃO executar `extract_all_engines()`.
|
||||
- [x] T009 [US1] Integration tests for media routing (Cenários B, C, D, E, F), `_media.json` naming resolution, and `input_meta` preservation with custom unknown fields in `tests/integration/test_media_routing.py`
|
||||
- Validar que campos desconhecidos arbitrários (ex: `custom_field="preserve-me"`, `custom_number=42`) são integralmente preservados em `MediaArticle`.
|
||||
|
||||
**Checkpoint**: User Story 1 completa e testável de forma independente com Ollama mockado.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: User Story 2 - Roteamento Direto e Preservação de Notícias Textuais (Priority: P2)
|
||||
|
||||
**Goal**: Garantir que artigos textuais substantivos (com ou sem mídias ilustrativas) ou páginas sem mídia candidata continuem sendo processados pelos 3 motores textuais em `*_extracted.json`, gerando incondicionalmente `*_media.json` com `{"articles": []}` caso não haja mídias.
|
||||
|
||||
**Independent Test**: Submeter página sem mídia candidata (Cenário A) e matérias jornalísticas longas com fotos/vídeos (Cenários G, H); verificar execução dos 3 motores, gravação no JSON textual e presença de `*_media.json` com `{"articles": []}`.
|
||||
|
||||
- [x] T010 [US2] Unit tests for direct gate bypass (`has_candidate_media == False`) and textual classification handling in `tests/unit/test_media_classifier.py`
|
||||
- [x] T011 [US2] Implement direct gate bypass in `process_batch` when `has_candidate_media == False` routing directly to `extract_all_engines` without LLM calls, ensure `content_type == "text"` passes to `extract_all_engines`, and execute `save_media_json` ensuring `_media.json` is always generated (with `{"articles": []}` when zero media articles) in `scripts/extract_article_contents.py`
|
||||
- [x] T012 [US2] Integration tests for textual articles (Cenários A, G, H), main JSON report counters (`total_articles`, `successful_articles`, `failed_articles`), mandatory generation of `_media.json` with `{"articles": []}`, and `input_meta` preservation in `tests/integration/test_media_routing.py`
|
||||
- Validar que os mesmos campos desconhecidos arbitrários em `input_meta` sobrevivem integralmente nos registros de artigos textuais em `*_extracted.json`.
|
||||
|
||||
**Checkpoint**: User Stories 1 e 2 funcionais e integradas, com preservação estrita do pipeline textual e arquivo de mídia incondicional.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: User Story 3 - Resiliência com Fallback Operacional Sequencial e Registro de Falhas (Priority: P3)
|
||||
|
||||
**Goal**: Implementar a cadeia sequencial de contingência (Ollama $\rightarrow$ Groq $\rightarrow$ OmniRoute), registrando falhas operacionais cumulativas inline no JSON principal com `classification_status: "failed"` sem abortar o lote e mantendo suporte multilíngue.
|
||||
|
||||
**Independent Test**: Simular os 5 estados de provedores (Cenários I, J, K, L, M) via mocks offline; verificar transição Groq/OmniRoute, first-valid-wins, registro de erro inline com `classification_status: "failed"` e continuidade do lote.
|
||||
|
||||
- [x] T013 [US3] Unit tests for sequential provider chain in `tests/unit/test_media_classifier.py`
|
||||
- Validar first-valid-wins (Ollama sucesso $\rightarrow$ Groq e OmniRoute com 0 chamadas; Ollama falha + Groq sucesso $\rightarrow$ OmniRoute com 0 chamadas);
|
||||
- Validar transição imediata para o próximo provider em caso de timeout, connection error, HTTP error ou schema inválido;
|
||||
- Validar que providers recebem parâmetros corretos (Ollama: `think: false`, `temperature: 0.0`; Groq: `openai/gpt-oss-20b`, `reasoning_effort: "low"`, `temperature: 0.0`, strict JSON Schema; OmniRoute: `cgpt-web/gpt-5.5`, `temperature: 0.0`, JSON Schema);
|
||||
- Validar que os 3 providers utilizam o mesmo prompt único, sem variantes por idioma e sem solicitações de campos proibidos (reasoning, confidence, rationale, summary, keywords, evidence, tradução);
|
||||
- Validar ausência de retries, votação, juiz ou confidence thresholds.
|
||||
- [x] T014 [US3] Extend `classify_media_content(payload, metrics, silent)` with Groq (`GROQ_ENDPOINT`, `GROQ_API_KEY`, `GROQ_MODEL="openai/gpt-oss-20b"`, `GROQ_TIMEOUT=15`, `reasoning_effort="low"`, `temperature=0.0`, strict JSON Schema) and OmniRoute (`OMNIROUTE_ENDPOINT`, `OMNIROUTE_API_KEY`, `OMNIROUTE_MODEL="cgpt-web/gpt-5.5"`, `OMNIROUTE_TIMEOUT=20`, `temperature=0.0`, JSON Schema), incrementing `fallback_groq` e `fallback_omniroute` imediatamente antes de cada chamada, registrando logs operacionais de fallback no stderr respeitando `silent`, e tratando configuração ausente como falha operacional do provider in `scripts/extract_article_contents.py`
|
||||
- [x] T015 [US3] Implement cumulative failure handling in `process_batch` recording inline in main JSON with `classification_status = "failed"`, `error_message`, incrementing `classification_failed`, registrando log de falha total no stderr respeitando `silent`, preservando crawl metadata, skipping multimotor, não adicionando o registro de falha ao `media_articles` nem ao array `articles` de `*_media.json`, e incrementing `failed_articles` in `scripts/extract_article_contents.py`
|
||||
- [x] T016 [US3] Integration tests for cumulative provider failure (Cenário K), fallback transitions (Cenários I, J, L), `input_meta` preservation in failure records, and multilingual support (Cenário M) across supported pipeline languages without language-specific prompts in `tests/integration/test_media_routing.py`
|
||||
- Validar no Cenário K que `classification_status == "failed"`, `error_message` existe, `extract_all_engines` NÃO é chamado, o registro NÃO entra em `*_media.json`, `failed_articles` incrementa exatamente uma vez, `input_meta` integral é preservado com campos desconhecidos, `crawled_url`, `page_title` e `http_status` disponíveis são preservados, e o lote continua normalmente para o próximo artigo.
|
||||
|
||||
**Checkpoint**: As 3 User Stories estão completas, com resiliência total, fallback sequencial determinístico e tratamento de falhas.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Polish & Observabilidade
|
||||
|
||||
**Purpose**: Métricas operacionais, segurança de credenciais, compatibilidade CLI e validação estática de Zero-Regex.
|
||||
|
||||
- [x] T017 [Polish] Implement in-memory tracking of remaining operational metrics (`total_evaluated`, `text`, `media`, and `media/<subtype>`) and output summary to stderr respecting `--silent` in `scripts/extract_article_contents.py`
|
||||
- Garantir que as métricas `total_evaluated`, `text`, `media`, `media/video`, `media/image`, `media/images`, `media/embed`, `media/mixed`, `fallback_groq`, `fallback_omniroute` e `classification_failed` sejam incrementadas uma única vez nos pontos definidos, sem duplicidade;
|
||||
- Ao final do lote, emitir o summary formatado no stderr respeitando `--silent`.
|
||||
- [x] T018 [P] Verify and update `tests/scripts/check_zero_regex.py` to include `scripts/extract_article_contents.py` and new test files in the scoped AST check
|
||||
- [x] T019 Integration tests verifying the 11 metrics values, logging in stderr, `--silent` suppression, secret non-exposure, CLI flag compatibility, and exit code preservation in `tests/integration/test_media_routing.py`
|
||||
- Validar contagem exata das 11 métricas nos pontos definidos;
|
||||
- Validar que o provider utilizado, acionamentos de fallback (Ollama $\rightarrow$ Groq, Groq $\rightarrow$ OmniRoute), classificação final e eventuais falhas totais são emitidos no stderr quando não silencioso;
|
||||
- Validar que `--silent` suprime todos os logs operacionais e o resumo de métricas;
|
||||
- Validar que API keys, tokens e secrets NUNCA aparecem em stderr, JSON de saída ou `error_message`;
|
||||
- Validar que o payload textual integral enviado ao classificador NÃO aparece no log/stderr por padrão (sem proibir o texto normal pertencente ao contrato de `*_extracted.json`);
|
||||
- Validar compatibilidade dos argumentos CLI (`-i`, `-o`, `-l`, `--lang`, `-t`, `-s`), resolução dos outputs e preservação dos exit codes existentes (`0`, `1`, `2`), sem criar novos testes complexos de SIGINT/Ctrl+C/exit 130 exclusivamente para esta feature.
|
||||
- [x] T020 Execute validation commands from `specs/007-media-article-routing/quickstart.md` (`pytest` suite and `python tests/scripts/check_zero_regex.py`)
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Execution Order
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Foundational[Phase 1: Foundational T001-T003] --> US1[Phase 2: User Story 1 - P1 T004-T009]
|
||||
US1 --> US2[Phase 3: User Story 2 - P2 T010-T012]
|
||||
US2 --> US3[Phase 4: User Story 3 - P3 T013-T016]
|
||||
US3 --> Polish[Phase 5: Polish & Observability T017-T020]
|
||||
Polish17[T017] --> Polish19[T019]
|
||||
Polish18[T018] --> Polish20[T020]
|
||||
Polish19 --> Polish20
|
||||
```
|
||||
|
||||
### Phase Dependencies
|
||||
- **Foundational (Phase 1)**: Sem dependências de histórias; implementa tipos, schema e cliente HTTP básico.
|
||||
- **User Story 1 (Phase 2 - P1)**: Depende de Foundational. Implementa detecção DOM, payload, Ollama e persistência `*_media.json`.
|
||||
- **User Story 2 (Phase 3 - P2)**: Depende de US1. Implementa bypass direto sem LLM, preservação do pipeline textual e emissão incondicional de `*_media.json` com `[]`.
|
||||
- **User Story 3 (Phase 4 - P3)**: Depende de US2. Implementa fallbacks Groq/OmniRoute na mesma função `classify_media_content`, tratamento de falhas inline e multilíngue.
|
||||
- **Polish (Phase 5)**: Depende de US1, US2 e US3. `T017` implementa o fechamento das métricas, `T019` testa métricas, segurança e CLI, `T018` atualiza o checker AST de forma independente, e `T020` executa a validação global final.
|
||||
|
||||
---
|
||||
|
||||
## Parallel Execution Opportunities
|
||||
|
||||
- **Phase 5 (Polish)**: `T018` (atualização do checker AST de zero-regex em `tests/scripts/check_zero_regex.py`) está marcado com `[P]` pois opera em arquivo independente e pode rodar em paralelo às demais tarefas.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
A entrega da feature será realizada de forma incremental e ordenada:
|
||||
1. **Fundação**: Dataclasses, constante do schema e helper HTTP `_http_post_json`.
|
||||
2. **História 1**: Detecção DOM, compact payload, Ollama, inicialização das 11 métricas e geração de `*_media.json`.
|
||||
3. **História 2**: Gate bypass sem LLM, preservação textual e geração incondicional de `*_media.json` (com `[]` quando vazio).
|
||||
4. **História 3**: Extensão da cadeia sequencial com Groq (`reasoning_effort="low"`) e OmniRoute, com registro inline de falha no JSON principal.
|
||||
5. **Polimento**: Fechamento da contagem e resumo das 11 métricas no stderr, verificação de não-exposição de segredos, compatibilidade CLI e verificação estática Zero-Regex.
|
||||
6. **Conclusão**: 100% das 20 tarefas concluídas e validadas contra a suíte de testes e o checker Zero-Regex.
|
||||
@@ -0,0 +1,225 @@
|
||||
"""
|
||||
Testes End-to-End (E2E) reais com Ollama e modelo qwen3.5:2b ao vivo.
|
||||
|
||||
Executa o funil completo de ponta a ponta:
|
||||
1. Comunicação HTTP real com Ollama local (qwen3.5:2b);
|
||||
2. Classificação semântica real para publicações de mídia vs texto;
|
||||
3. Execução de lote com Foxcape + Ollama + Multimotor (Trafilatura, Newspaper4k, Readability);
|
||||
4. Validação dos arquivos de saída (*_extracted.json e *_media.json).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from bs4 import BeautifulSoup
|
||||
|
||||
# Adiciona a raiz do projeto ao path
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent.parent))
|
||||
|
||||
from scripts.extract_article_contents import (
|
||||
MediaCandidateInfo,
|
||||
MediaClassification,
|
||||
build_compact_payload,
|
||||
classify_media_content,
|
||||
detect_candidate_media,
|
||||
process_batch,
|
||||
)
|
||||
|
||||
|
||||
def is_ollama_available() -> bool:
|
||||
"""Verifica se o servidor Ollama está ativo localmente e com qwen3.5:2b disponível."""
|
||||
try:
|
||||
url = os.environ.get("OLLAMA_ENDPOINT", "http://localhost:11434").rstrip("/") + "/api/tags"
|
||||
req = urllib.request.Request(url)
|
||||
with urllib.request.urlopen(req, timeout=3) as resp:
|
||||
if resp.status == 200:
|
||||
data = json.loads(resp.read().decode("utf-8"))
|
||||
models = [m.get("name", "") for m in data.get("models", [])]
|
||||
return any("qwen3.5:2b" in m for m in models)
|
||||
except Exception:
|
||||
return False
|
||||
return False
|
||||
|
||||
|
||||
@pytest.mark.skipif(not is_ollama_available(), reason="Ollama local com qwen3.5:2b não está acessível")
|
||||
def test_live_ollama_video_media_classification() -> None:
|
||||
"""Valida classificação real via Ollama para publicação de vídeo esportivo com texto descritivo curto."""
|
||||
html = """
|
||||
<html>
|
||||
<head><title>Golaço de falta na Copa Libertadores - Melhores Momentos</title></head>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Golaço de falta na Copa Libertadores</h1>
|
||||
<video controls src="https://example.com/videos/gol.mp4"></video>
|
||||
<p>Confira no vídeo acima o golaço de falta marcado no último minuto da partida.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
candidate_info = detect_candidate_media(soup)
|
||||
assert candidate_info.has_candidate_media is True
|
||||
assert candidate_info.has_video is True
|
||||
|
||||
payload = build_compact_payload(soup, candidate_info)
|
||||
metrics = {
|
||||
"fallback_groq": 0,
|
||||
"fallback_omniroute": 0,
|
||||
"classification_failed": 0,
|
||||
}
|
||||
|
||||
classification, err = classify_media_content(payload, metrics, silent=False)
|
||||
assert err is None
|
||||
assert classification is not None
|
||||
assert classification.content_type == "media"
|
||||
assert classification.media_type in ("video", "mixed")
|
||||
# Nenhum fallback deve ter sido acionado
|
||||
assert metrics["fallback_groq"] == 0
|
||||
assert metrics["fallback_omniroute"] == 0
|
||||
|
||||
|
||||
@pytest.mark.skipif(not is_ollama_available(), reason="Ollama local com qwen3.5:2b não está acessível")
|
||||
def test_live_ollama_substantive_textual_article_classification() -> None:
|
||||
"""Valida classificação real via Ollama para matéria jornalística densa acompanhada de foto ilustrativa."""
|
||||
html = """
|
||||
<html>
|
||||
<head><title>Banco Central mantém taxa de juros e sinaliza cautela econômica</title></head>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Banco Central mantém taxa de juros e sinaliza cautela econômica</h1>
|
||||
<img src="https://example.com/fotos/reuniao_copom.jpg" alt="Reunião do Copom">
|
||||
<p>O Comitê de Política Monetária (Copom) do Banco Central decidiu por unanimidade manter a taxa básica de juros (Selic) inalterada, conforme amplamente esperado pelos analistas de mercado.</p>
|
||||
<p>Em seu comunicado oficial, a autoridade monetária destacou que o cenário global permanece volátil e que a conjuntura doméstica exige perseverança na condução da política monetária para consolidar a desinflação.</p>
|
||||
<p>Economistas ouvidos pelo relatório Focus projetam que os primeiros cortes poderão ocorrer no próximo trimestre, caso as expectativas de inflação continuem ancoradas nas metas estabelecidas pelo Conselho Monetário Nacional.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
candidate_info = detect_candidate_media(soup)
|
||||
assert candidate_info.has_candidate_media is True
|
||||
assert candidate_info.image_count == 1
|
||||
|
||||
payload = build_compact_payload(soup, candidate_info)
|
||||
metrics = {
|
||||
"fallback_groq": 0,
|
||||
"fallback_omniroute": 0,
|
||||
"classification_failed": 0,
|
||||
}
|
||||
|
||||
classification, err = classify_media_content(payload, metrics, silent=False)
|
||||
assert err is None
|
||||
assert classification is not None
|
||||
assert classification.content_type == "text"
|
||||
assert classification.media_type is None
|
||||
assert metrics["fallback_groq"] == 0
|
||||
|
||||
|
||||
@pytest.mark.skipif(not is_ollama_available(), reason="Ollama local com qwen3.5:2b não está acessível")
|
||||
def test_live_e2e_batch_funnel_with_ollama_and_multimotor(tmp_path: Path) -> None:
|
||||
"""
|
||||
Executa o funil completo em lote:
|
||||
- 1 artigo textual puro (sem mídia) -> bypass direto do gate -> extraído pelos 3 motores
|
||||
- 1 artigo predominantemente de mídia (vídeo) -> classificado pelo Ollama ao vivo -> desviado para *_media.json
|
||||
- 1 artigo textual longo com foto ilustrativa -> classificado pelo Ollama ao vivo -> extraído pelos 3 motores
|
||||
"""
|
||||
input_file = tmp_path / "batch_funnel.json"
|
||||
batch_data = {
|
||||
"query": "noticias ao vivo",
|
||||
"language": "pt",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Artigo Textual Puro",
|
||||
"url": "https://example.com/texto-puro",
|
||||
"custom_id": "item-001",
|
||||
},
|
||||
{
|
||||
"titulo": "Vídeo Highlights do Jogo",
|
||||
"url": "https://example.com/video-highlights",
|
||||
"custom_id": "item-002",
|
||||
},
|
||||
{
|
||||
"titulo": "Análise Econômica com Foto",
|
||||
"url": "https://example.com/analise-economica",
|
||||
"custom_id": "item-003",
|
||||
},
|
||||
],
|
||||
}
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(batch_data, f, ensure_ascii=False, indent=2)
|
||||
|
||||
html_puro = """
|
||||
<html><body><article>
|
||||
<h1>Artigo Textual Puro</h1>
|
||||
<p>Primeiro parágrafo longo e substancial sobre política pública e saúde coletiva.</p>
|
||||
<p>Segundo parágrafo detalhando as medidas sanitárias aprovadas pelo congresso.</p>
|
||||
</article></body></html>
|
||||
"""
|
||||
|
||||
html_video = """
|
||||
<html><body><article>
|
||||
<h1>Vídeo Highlights do Jogo</h1>
|
||||
<video src="highlights.mp4"></video>
|
||||
<p>Veja o resumo em vídeo com os melhores momentos da rodada esportiva.</p>
|
||||
</article></body></html>
|
||||
"""
|
||||
|
||||
html_foto = """
|
||||
<html><body><article>
|
||||
<h1>Análise Econômica com Foto</h1>
|
||||
<img src="foto_grafico.jpg">
|
||||
<p>O mercado de capitais registrou forte valorização nesta segunda-feira impulsionado por dados positivos do setor industrial.</p>
|
||||
<p>As ações das principais empresas de commodities lideraram os ganhos no pregão, enquanto o câmbio operou em estabilidade.</p>
|
||||
<p>Investidores estrangeiros aportaram mais de 500 milhões de reais no mercado à vista ao longo da última semana.</p>
|
||||
</article></body></html>
|
||||
"""
|
||||
|
||||
def mock_crawl(url: str) -> tuple[str, str, int]:
|
||||
if "texto-puro" in url:
|
||||
return html_puro, "Artigo Textual Puro", 200
|
||||
elif "video-highlights" in url:
|
||||
return html_video, "Vídeo Highlights do Jogo", 200
|
||||
else:
|
||||
return html_foto, "Análise Econômica com Foto", 200
|
||||
|
||||
from unittest.mock import patch, MagicMock
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler:
|
||||
crawler = MagicMock()
|
||||
crawler.__enter__.return_value = crawler
|
||||
crawler.__exit__.return_value = None
|
||||
crawler.crawl.side_effect = mock_crawl
|
||||
MockCrawler.return_value = crawler
|
||||
|
||||
# Executa process_batch usando o Ollama real em localhost:11434
|
||||
report = process_batch(input_file, silent=False)
|
||||
|
||||
# 1. Verifica arquivo textual *_extracted.json
|
||||
# Devem existir 2 artigos textuais no relatório principal (item-001 e item-003)
|
||||
assert report.total_articles == 2
|
||||
assert report.successful_articles == 2
|
||||
assert report.failed_articles == 0
|
||||
assert len(report.articles) == 2
|
||||
|
||||
# Preservação de input_meta com custom_id nos artigos textuais
|
||||
assert report.articles[0].input_meta.to_dict()["custom_id"] == "item-001"
|
||||
assert report.articles[1].input_meta.to_dict()["custom_id"] == "item-003"
|
||||
|
||||
# 2. Verifica arquivo de mídia *_media.json
|
||||
# Deve conter exatamente 1 artigo desviado (item-002)
|
||||
media_file = tmp_path / "batch_funnel_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
|
||||
assert len(media_data["articles"]) == 1
|
||||
media_item = media_data["articles"][0]
|
||||
assert media_item["input_meta"]["custom_id"] == "item-002"
|
||||
assert media_item["content_type"] == "media"
|
||||
assert media_item["media_type"] in ("video", "mixed")
|
||||
@@ -0,0 +1,567 @@
|
||||
"""
|
||||
Testes de integração para roteamento de artigos de mídia e texto.
|
||||
|
||||
Cobre os Cenários B a F (User Story 1), resolução de caminhos padrão e com -o,
|
||||
e preservação integral de metadados em input_meta com campos customizados.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from unittest.mock import patch, MagicMock
|
||||
|
||||
import pytest
|
||||
|
||||
# Adiciona o diretório raiz ao sys.path
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent.parent))
|
||||
|
||||
from scripts.extract_article_contents import process_batch, ArticleCrawler
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_input_json(tmp_path: Path) -> Path:
|
||||
"""Cria um arquivo de entrada de busca sintético para testes."""
|
||||
input_file = tmp_path / "search_sample.json"
|
||||
data = {
|
||||
"query": "noticias futebol",
|
||||
"language": "pt",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Gol Histórico do River Plate",
|
||||
"url": "https://example.com/video-gol",
|
||||
"subtitulo": "Veja o lance em vídeo",
|
||||
"quando_publicado": "2026-08-24T12:00:00Z",
|
||||
"pagina": 1,
|
||||
"custom_field": "preserve-me",
|
||||
"custom_number": 42,
|
||||
}
|
||||
],
|
||||
}
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(data, f, ensure_ascii=False, indent=2)
|
||||
return input_file
|
||||
|
||||
|
||||
def mock_crawl_html(html_body: str, title: str = "Página Notícia", status_code: int = 200) -> MagicMock:
|
||||
mock = MagicMock()
|
||||
mock.__enter__.return_value = mock
|
||||
mock.__exit__.return_value = None
|
||||
mock.crawl.return_value = (html_body, title, status_code)
|
||||
return mock
|
||||
|
||||
|
||||
def test_media_routing_scenario_b_video(tmp_path: Path) -> None:
|
||||
"""Cenário B: Publicação com vídeo e texto curto descritivo -> roteada para *_media.json."""
|
||||
input_file = tmp_path / "river_plate.json"
|
||||
input_data = {
|
||||
"query": "river plate",
|
||||
"language": "es",
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Video del Golazo",
|
||||
"url": "https://example.com/video",
|
||||
"custom_tracker_id": "trk-999",
|
||||
}
|
||||
],
|
||||
}
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(input_data, f)
|
||||
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Video del Golazo</h1>
|
||||
<video src="gol.mp4"></video>
|
||||
<p>Mira el resumen del partido.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
|
||||
# Mock crawler e mock LLM (Ollama respondendo content_type=media, media_type=video)
|
||||
mock_llm_response = (
|
||||
200,
|
||||
json.dumps(
|
||||
{
|
||||
"message": {
|
||||
"content": json.dumps({"content_type": "media", "media_type": "video"})
|
||||
}
|
||||
}
|
||||
),
|
||||
)
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", return_value=mock_llm_response), \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
crawler_instance = mock_crawl_html(html, title="Video del Golazo", status_code=200)
|
||||
MockCrawler.return_value = crawler_instance
|
||||
|
||||
report = process_batch(input_file, silent=True)
|
||||
|
||||
# Multimotor NÃO deve ter sido executado
|
||||
assert mock_extractors.call_count == 0
|
||||
|
||||
# Relatório textual deve ter 0 artigos no JSON principal
|
||||
assert report.total_articles == 0
|
||||
assert report.successful_articles == 0
|
||||
assert len(report.articles) == 0
|
||||
|
||||
# Arquivo de mídia river_plate_media.json deve ter sido criado
|
||||
media_file = tmp_path / "river_plate_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
|
||||
assert "articles" in media_data
|
||||
assert len(media_data["articles"]) == 1
|
||||
article = media_data["articles"][0]
|
||||
assert article["content_type"] == "media"
|
||||
assert article["media_type"] == "video"
|
||||
assert article["crawled_url"] == "https://example.com/video"
|
||||
assert article["http_status"] == 200
|
||||
assert article["input_meta"]["titulo"] == "Video del Golazo"
|
||||
assert article["input_meta"]["custom_tracker_id"] == "trk-999"
|
||||
|
||||
|
||||
def test_media_routing_scenarios_c_d_e_f(tmp_path: Path) -> None:
|
||||
"""Cenários C, D, E, F: image, images, embed, mixed."""
|
||||
cases = [
|
||||
("image", "<article><img src='foto.jpg'><p>Legenda curta.</p></article>"),
|
||||
("images", "<article><img src='1.jpg'><img src='2.jpg'><p>Galeria.</p></article>"),
|
||||
("embed", "<article><iframe src='insta.com/p/123'></iframe><p>Post.</p></article>"),
|
||||
("mixed", "<article><video src='v.mp4'></video><img src='f.jpg'><p>Misto.</p></article>"),
|
||||
]
|
||||
|
||||
for media_type, html in cases:
|
||||
input_file = tmp_path / f"test_{media_type}.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"titulo": f"Titulo {media_type}",
|
||||
"url": f"https://example.com/{media_type}",
|
||||
}
|
||||
]
|
||||
},
|
||||
f,
|
||||
)
|
||||
|
||||
mock_llm_response = (
|
||||
200,
|
||||
json.dumps(
|
||||
{
|
||||
"message": {
|
||||
"content": json.dumps({"content_type": "media", "media_type": media_type})
|
||||
}
|
||||
}
|
||||
),
|
||||
)
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", return_value=mock_llm_response), \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
process_batch(input_file, silent=True)
|
||||
|
||||
assert mock_extractors.call_count == 0
|
||||
|
||||
media_file = tmp_path / f"test_{media_type}_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
|
||||
assert len(media_data["articles"]) == 1
|
||||
assert media_data["articles"][0]["media_type"] == media_type
|
||||
|
||||
|
||||
def test_media_routing_custom_output_path_resolution(tmp_path: Path) -> None:
|
||||
"""Valida resolução de caminho com -o customizado: out/processados/resultado.json -> resultado_media.json."""
|
||||
input_file = tmp_path / "raw_input.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump({"items": [{"titulo": "T", "url": "https://example.com/item"}]}, f)
|
||||
|
||||
custom_out = tmp_path / "custom_dir" / "saida_final.json"
|
||||
|
||||
html = "<article><video src='v.mp4'></video><p>Texto curto.</p></article>"
|
||||
mock_llm_response = (
|
||||
200,
|
||||
json.dumps(
|
||||
{"message": {"content": json.dumps({"content_type": "media", "media_type": "video"})}}
|
||||
),
|
||||
)
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", return_value=mock_llm_response):
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
process_batch(input_file, output_path=custom_out, silent=True)
|
||||
|
||||
assert custom_out.exists()
|
||||
expected_media_file = tmp_path / "custom_dir" / "saida_final_media.json"
|
||||
assert expected_media_file.exists()
|
||||
|
||||
|
||||
def test_textual_routing_scenario_a_no_media_bypasses_llm(tmp_path: Path) -> None:
|
||||
"""Cenário A: Artigo sem mídia candidata segue direto ao multimotor sem chamar LLM."""
|
||||
input_file = tmp_path / "pure_text.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Notícia Textual Sem Mídia",
|
||||
"url": "https://example.com/texto-puro",
|
||||
"custom_field": "preserve-me",
|
||||
"custom_number": 42,
|
||||
}
|
||||
]
|
||||
},
|
||||
f,
|
||||
)
|
||||
|
||||
html = "<article><h1>Texto Puro</h1><p>Notícia densa com múltiplos parágrafos.</p></article>"
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json") as mock_http, \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
mock_extractors.return_value = (None, None, None)
|
||||
|
||||
report = process_batch(input_file, silent=True)
|
||||
|
||||
# LLM NÃO deve ser chamado (bypass)
|
||||
assert mock_http.call_count == 0
|
||||
|
||||
# Multimotor DEVE ser chamado
|
||||
assert mock_extractors.call_count == 1
|
||||
|
||||
# JSON textual principal contém o artigo
|
||||
assert report.total_articles == 1
|
||||
assert report.successful_articles == 1
|
||||
assert report.failed_articles == 0
|
||||
assert len(report.articles) == 1
|
||||
assert report.articles[0].input_meta.to_dict()["custom_field"] == "preserve-me"
|
||||
assert report.articles[0].input_meta.to_dict()["custom_number"] == 42
|
||||
|
||||
# *_media.json MUST ser gerado com {"articles": []}
|
||||
media_file = tmp_path / "pure_text_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
assert media_data == {"articles": []}
|
||||
|
||||
|
||||
def test_textual_routing_scenarios_g_and_h_with_illustrative_media(tmp_path: Path) -> None:
|
||||
"""Cenários G e H: Notícias longas com imagem/vídeo ilustrativo recebem content_type=text."""
|
||||
cases = [
|
||||
("cenario_g_image", "<article><img src='foto.jpg'><p>Texto jornalístico longo e substancial 1.</p><p>Texto longo 2.</p></article>"),
|
||||
("cenario_h_video", "<article><video src='v.mp4'></video><p>Texto jornalístico longo e substancial 1.</p><p>Texto longo 2.</p></article>"),
|
||||
]
|
||||
|
||||
mock_llm_text_response = (
|
||||
200,
|
||||
json.dumps(
|
||||
{"message": {"content": json.dumps({"content_type": "text", "media_type": None})}}
|
||||
),
|
||||
)
|
||||
|
||||
for case_name, html in cases:
|
||||
input_file = tmp_path / f"{case_name}.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"titulo": f"Matéria {case_name}",
|
||||
"url": f"https://example.com/{case_name}",
|
||||
"custom_author": "Repórter Especial",
|
||||
}
|
||||
]
|
||||
},
|
||||
f,
|
||||
)
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", return_value=mock_llm_text_response) as mock_http, \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
mock_extractors.return_value = (None, None, None)
|
||||
|
||||
report = process_batch(input_file, silent=True)
|
||||
|
||||
# LLM foi chamado pois havia mídia candidata
|
||||
assert mock_http.call_count == 1
|
||||
# Multimotor foi executado pois a classificação retornou text
|
||||
assert mock_extractors.call_count == 1
|
||||
|
||||
assert report.total_articles == 1
|
||||
assert report.successful_articles == 1
|
||||
assert report.articles[0].input_meta.to_dict()["custom_author"] == "Repórter Especial"
|
||||
|
||||
# *_media.json gerado incondicionalmente vazio
|
||||
media_file = tmp_path / f"{case_name}_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
assert media_data == {"articles": []}
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Testes de Integração de Fallback e Falha Cumulativa (User Story 3 / T016)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_scenario_k_cumulative_failure(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Cenário K: Falha em todos os 3 provedores -> registro inline de falha no JSON principal."""
|
||||
monkeypatch.setenv("GROQ_API_KEY", "test-key")
|
||||
monkeypatch.setenv("OMNIROUTE_ENDPOINT", "http://omniroute.test")
|
||||
|
||||
input_file = tmp_path / "fail_input.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump(
|
||||
{
|
||||
"items": [
|
||||
{
|
||||
"titulo": "Artigo com Falha de LLM",
|
||||
"url": "https://example.com/fail-item",
|
||||
"custom_id": "cust-77",
|
||||
}
|
||||
]
|
||||
},
|
||||
f,
|
||||
)
|
||||
|
||||
html = "<article><video src='v.mp4'></video><p>Texto.</p></article>"
|
||||
|
||||
def mock_http_fail(*args: Any, **kwargs: Any) -> tuple[int, str]:
|
||||
raise ConnectionResetError("Connection reset")
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", side_effect=mock_http_fail), \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html, title="Artigo com Falha de LLM", status_code=200)
|
||||
|
||||
report = process_batch(input_file, silent=True)
|
||||
|
||||
# Multimotor NÃO deve ter sido executado
|
||||
assert mock_extractors.call_count == 0
|
||||
|
||||
# Artigo presente no JSON textual como falha
|
||||
assert report.total_articles == 1
|
||||
assert report.successful_articles == 0
|
||||
assert report.failed_articles == 1
|
||||
assert len(report.articles) == 1
|
||||
|
||||
failed_art = report.articles[0]
|
||||
assert failed_art.classification_status == "failed"
|
||||
assert failed_art.error_message is not None
|
||||
assert failed_art.crawled_url == "https://example.com/fail-item"
|
||||
assert failed_art.page_title == "Artigo com Falha de LLM"
|
||||
assert failed_art.http_status == 200
|
||||
assert failed_art.input_meta.to_dict()["custom_id"] == "cust-77"
|
||||
assert failed_art.trafilatura is None
|
||||
assert failed_art.newspaper4k is None
|
||||
assert failed_art.readability is None
|
||||
|
||||
# NÃO deve estar em *_media.json
|
||||
media_file = tmp_path / "fail_input_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
assert media_data == {"articles": []}
|
||||
|
||||
|
||||
def test_scenario_i_j_l_fallbacks_and_multilingual(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Cenários I, J, L e M: Transição para Groq/OmniRoute e suporte multilíngue (pt, es, en)."""
|
||||
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
|
||||
monkeypatch.setenv("OMNIROUTE_ENDPOINT", "http://omniroute.test")
|
||||
|
||||
input_file = tmp_path / "multilingual.json"
|
||||
items = [
|
||||
{"titulo": "Notícia em Português", "url": "https://example.com/pt", "lang": "pt"},
|
||||
{"titulo": "Noticia en Español", "url": "https://example.com/es", "lang": "es"},
|
||||
{"titulo": "English News Article", "url": "https://example.com/en", "lang": "en"},
|
||||
]
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump({"items": items}, f)
|
||||
|
||||
html = "<article><video src='v.mp4'></video><p>Texto.</p></article>"
|
||||
|
||||
# 1º item: Ollama falha, Groq retorna media
|
||||
# 2º item: Ollama falha, Groq falha, OmniRoute retorna media
|
||||
# 3º item: Ollama falha, Groq retorna text -> vai para multimotor
|
||||
step_state = {"call_idx": 0}
|
||||
|
||||
def mock_http_chain(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
if "11434" in url:
|
||||
# Ollama sempre falha
|
||||
raise TimeoutError("Ollama timeout")
|
||||
if "groq.com" in url:
|
||||
step_state["call_idx"] += 1
|
||||
if step_state["call_idx"] == 1:
|
||||
# 1º item via Groq: media
|
||||
return 200, json.dumps({"choices": [{"message": {"content": json.dumps({"content_type": "media", "media_type": "video"})}}]})
|
||||
if step_state["call_idx"] == 2:
|
||||
# 2º item: Groq falha
|
||||
raise RuntimeError("Groq rate limit")
|
||||
if step_state["call_idx"] >= 3:
|
||||
# 3º item via Groq: text
|
||||
return 200, json.dumps({"choices": [{"message": {"content": json.dumps({"content_type": "text", "media_type": None})}}]})
|
||||
if "omniroute.test" in url:
|
||||
# OmniRoute responde media para o 2º item
|
||||
return 200, json.dumps({"choices": [{"message": {"content": json.dumps({"content_type": "media", "media_type": "embed"})}}]})
|
||||
return 500, "error"
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", side_effect=mock_http_chain), \
|
||||
patch("scripts.extract_article_contents.extract_all_engines") as mock_extractors:
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
mock_extractors.return_value = (None, None, None)
|
||||
|
||||
report = process_batch(input_file, silent=True)
|
||||
|
||||
# 3º item foi para multimotor
|
||||
assert mock_extractors.call_count == 1
|
||||
assert report.total_articles == 1
|
||||
assert report.successful_articles == 1
|
||||
assert report.articles[0].input_meta.to_dict()["lang"] == "en"
|
||||
|
||||
# 1º e 2º itens foram para *_media.json
|
||||
media_file = tmp_path / "multilingual_media.json"
|
||||
assert media_file.exists()
|
||||
with open(media_file, "r", encoding="utf-8") as f:
|
||||
media_data = json.load(f)
|
||||
|
||||
assert len(media_data["articles"]) == 2
|
||||
assert media_data["articles"][0]["input_meta"]["lang"] == "pt"
|
||||
assert media_data["articles"][0]["media_type"] == "video"
|
||||
assert media_data["articles"][1]["input_meta"]["lang"] == "es"
|
||||
assert media_data["articles"][1]["media_type"] == "embed"
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# Testes de Observabilidade, Segurança e CLI (Phase 5 / T019)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_metrics_values_and_stderr_summary_logging(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""Valida valores das métricas e emissão formatada no stderr quando silent=False."""
|
||||
input_file = tmp_path / "metrics_test.json"
|
||||
items = [
|
||||
{"titulo": "Texto Puro", "url": "https://example.com/puro"},
|
||||
{"titulo": "Video Direct", "url": "https://example.com/video"},
|
||||
]
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump({"items": items}, f)
|
||||
|
||||
def mock_crawl(url: str) -> tuple[str, str, int]:
|
||||
if "puro" in url:
|
||||
return "<article><p>Densa matéria pura.</p></article>", "Puro", 200
|
||||
return "<article><video src='v.mp4'></video><p>Vídeo.</p></article>", "Vídeo", 200
|
||||
|
||||
def mock_http(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
return 200, json.dumps({"message": {"content": json.dumps({"content_type": "media", "media_type": "video"})}})
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", side_effect=mock_http), \
|
||||
patch("scripts.extract_article_contents.extract_all_engines", return_value=(None, None, None)):
|
||||
|
||||
crawler = MagicMock()
|
||||
crawler.__enter__.return_value = crawler
|
||||
crawler.__exit__.return_value = None
|
||||
crawler.crawl.side_effect = mock_crawl
|
||||
MockCrawler.return_value = crawler
|
||||
|
||||
process_batch(input_file, silent=False)
|
||||
|
||||
captured = capsys.readouterr()
|
||||
assert "[MEDIA] Métricas de Roteamento:" in captured.err
|
||||
assert "Total avaliados: 2" in captured.err
|
||||
assert "Texto: 1" in captured.err
|
||||
assert "Mídia: 1" in captured.err
|
||||
assert "video=1" in captured.err
|
||||
|
||||
|
||||
def test_silent_mode_suppresses_all_media_logs(tmp_path: Path, capsys: pytest.CaptureFixture[str]) -> None:
|
||||
"""Valida que flag silent=True suprime todos os logs em stderr."""
|
||||
input_file = tmp_path / "silent_test.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump({"items": [{"titulo": "T", "url": "https://example.com/item"}]}, f)
|
||||
|
||||
html = "<article><video src='v.mp4'></video><p>Texto.</p></article>"
|
||||
mock_resp = (200, json.dumps({"message": {"content": json.dumps({"content_type": "media", "media_type": "video"})}}))
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", return_value=mock_resp):
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
process_batch(input_file, silent=True)
|
||||
|
||||
captured = capsys.readouterr()
|
||||
assert "[MEDIA]" not in captured.err
|
||||
|
||||
|
||||
def test_api_key_security_not_leaked_in_logs(
|
||||
tmp_path: Path, monkeypatch: pytest.MonkeyPatch, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
"""Valida que chaves secretas de API (Groq/OmniRoute) nunca são expostas em logs/stderr."""
|
||||
secret_groq = "gsk_secret_token_12345"
|
||||
secret_omni = "omni_secret_token_67890"
|
||||
monkeypatch.setenv("GROQ_API_KEY", secret_groq)
|
||||
monkeypatch.setenv("OMNIROUTE_API_KEY", secret_omni)
|
||||
monkeypatch.setenv("OMNIROUTE_ENDPOINT", "http://omniroute.test")
|
||||
|
||||
input_file = tmp_path / "sec_test.json"
|
||||
with open(input_file, "w", encoding="utf-8") as f:
|
||||
json.dump({"items": [{"titulo": "T", "url": "https://example.com/sec"}]}, f)
|
||||
|
||||
html = "<article><video src='v.mp4'></video><p>Texto de teste confidencial.</p></article>"
|
||||
|
||||
# Ollama falha -> Groq falha -> OmniRoute responde
|
||||
def mock_http(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
if "11434" in url:
|
||||
raise ConnectionRefusedError()
|
||||
if "groq.com" in url:
|
||||
raise RuntimeError("Groq rate limit")
|
||||
return 200, json.dumps({"choices": [{"message": {"content": json.dumps({"content_type": "media", "media_type": "mixed"})}}]})
|
||||
|
||||
with patch("scripts.extract_article_contents.ArticleCrawler") as MockCrawler, \
|
||||
patch("scripts.extract_article_contents._http_post_json", side_effect=mock_http):
|
||||
|
||||
MockCrawler.return_value = mock_crawl_html(html)
|
||||
process_batch(input_file, silent=False)
|
||||
|
||||
captured = capsys.readouterr()
|
||||
assert secret_groq not in captured.err
|
||||
assert secret_omni not in captured.err
|
||||
assert secret_groq not in captured.out
|
||||
assert secret_omni not in captured.out
|
||||
# Garante que payload textual não foi vazado no log
|
||||
assert "Texto de teste confidencial." not in captured.err
|
||||
|
||||
|
||||
def test_cli_argument_compatibility() -> None:
|
||||
"""Valida que argumentos de linha de comando existentes continuam totalmente compatíveis."""
|
||||
from scripts.extract_article_contents import parse_arguments
|
||||
|
||||
args = parse_arguments(["-i", "input.json", "-o", "custom_out.json", "--limit", "10", "--lang", "es", "--silent"])
|
||||
assert args.input == "input.json"
|
||||
assert args.output == "custom_out.json"
|
||||
assert args.limit == 10
|
||||
assert args.language == "es"
|
||||
assert args.silent is True
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -27,9 +27,15 @@ import yaml
|
||||
SCOPED_PYTHON_PATHS = [
|
||||
"src/runtime",
|
||||
"tests/runtime",
|
||||
"scripts/extract_article_contents.py",
|
||||
"tests/unit",
|
||||
"tests/integration",
|
||||
]
|
||||
|
||||
CONTRACT_SCHEMA_DIR = "specs/006-article-consolidation-runtime/contracts"
|
||||
CONTRACT_SCHEMA_DIRS = [
|
||||
"specs/006-article-consolidation-runtime/contracts",
|
||||
"specs/007-media-article-routing/contracts",
|
||||
]
|
||||
PROMPTFOO_CONFIG_PATHS = [
|
||||
"evals/promptfoo.config.yaml",
|
||||
]
|
||||
@@ -90,10 +96,11 @@ class RegexASTVisitor(ast.NodeVisitor):
|
||||
def check_python_ast(root: Path) -> List[str]:
|
||||
violations: List[str] = []
|
||||
for scoped_rel in SCOPED_PYTHON_PATHS:
|
||||
target_dir = root / scoped_rel
|
||||
if not target_dir.exists():
|
||||
target_path = root / scoped_rel
|
||||
if not target_path.exists():
|
||||
continue
|
||||
for py_file in target_dir.rglob("*.py"):
|
||||
py_files = [target_path] if target_path.is_file() else list(target_path.rglob("*.py"))
|
||||
for py_file in py_files:
|
||||
try:
|
||||
content = py_file.read_text(encoding="utf-8")
|
||||
tree = ast.parse(content, filename=str(py_file))
|
||||
@@ -107,9 +114,10 @@ def check_python_ast(root: Path) -> List[str]:
|
||||
|
||||
def check_json_schemas(root: Path) -> List[str]:
|
||||
violations: List[str] = []
|
||||
schema_dir = root / CONTRACT_SCHEMA_DIR
|
||||
for schema_rel in CONTRACT_SCHEMA_DIRS:
|
||||
schema_dir = root / schema_rel
|
||||
if not schema_dir.exists():
|
||||
return violations
|
||||
continue
|
||||
|
||||
for schema_file in schema_dir.glob("*.schema.json"):
|
||||
try:
|
||||
|
||||
@@ -379,6 +379,9 @@ def test_funnel_live_api_execution_if_configured():
|
||||
content = "O River Plate empatou em 1 a 1 em Bogotá com gols de Otamendi na Copa Sul-Americana."
|
||||
|
||||
refined = adapter.disambiguate(ecp, content, initial_res)
|
||||
if refined is None and initial_res.warnings:
|
||||
pytest.skip(f"Live API unavailable or blocked: {initial_res.warnings}")
|
||||
|
||||
assert refined is not None
|
||||
assert refined.decision in [
|
||||
DecisionCategory.DIRECT_INHERENT,
|
||||
|
||||
@@ -18,6 +18,7 @@ from scripts.extract_article_contents import (
|
||||
ExtractedArticle,
|
||||
ExtractionBatchReport,
|
||||
InputArticle,
|
||||
MediaClassification,
|
||||
NewspaperData,
|
||||
NewspaperExtractor,
|
||||
ReadabilityData,
|
||||
@@ -321,7 +322,11 @@ def test_e2e_live_article_extraction(tmp_path: Path):
|
||||
|
||||
out_file = tmp_path / "e2e_live_extracted.json"
|
||||
|
||||
# Executa o batch real com limite de 1 notícia
|
||||
# Executa o batch real com limite de 1 notícia simulando classificação textual
|
||||
with patch(
|
||||
"scripts.extract_article_contents.classify_media_content",
|
||||
return_value=(MediaClassification(content_type="text", media_type=None), None),
|
||||
):
|
||||
report = process_batch(
|
||||
input_path=live_input_file,
|
||||
output_path=out_file,
|
||||
@@ -364,9 +369,14 @@ def test_e2e_cli_live_execution(tmp_path: Path):
|
||||
pytest.skip("Arquivo out/river_plate.json não encontrado para teste E2E.")
|
||||
|
||||
out_file = tmp_path / "e2e_cli_live.json"
|
||||
with patch(
|
||||
"scripts.extract_article_contents.classify_media_content",
|
||||
return_value=(MediaClassification(content_type="text", media_type=None), None),
|
||||
):
|
||||
exit_code = main(["-i", str(live_input_file), "-o", str(out_file), "--limit", "1", "-s"])
|
||||
|
||||
assert exit_code == 0
|
||||
assert out_file.exists()
|
||||
data = json.loads(out_file.read_text(encoding="utf-8"))
|
||||
assert data["successful_articles"] == 1
|
||||
|
||||
|
||||
@@ -0,0 +1,365 @@
|
||||
"""
|
||||
Testes unitários do classificador de mídia e gate estrutural DOM.
|
||||
|
||||
Valida schema de 2 campos, navegação estrutural na DOM (Zero-Regex)
|
||||
e montagem do payload compacto.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import sys
|
||||
import urllib.error
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from bs4 import BeautifulSoup
|
||||
|
||||
# Adiciona o diretório raiz ao sys.path para import dos scripts
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent.parent))
|
||||
|
||||
from scripts.extract_article_contents import (
|
||||
MEDIA_CLASSIFIER_PROMPT,
|
||||
MEDIA_CLASSIFIER_SCHEMA,
|
||||
MediaCandidateInfo,
|
||||
MediaClassification,
|
||||
build_compact_payload,
|
||||
classify_media_content,
|
||||
detect_candidate_media,
|
||||
validate_classifier_response,
|
||||
)
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 1. Testes de Validação do Contrato Estruturado (validate_classifier_response)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_validate_classifier_response_text_valid() -> None:
|
||||
data = {"content_type": "text", "media_type": None}
|
||||
result = validate_classifier_response(data)
|
||||
assert result == MediaClassification(content_type="text", media_type=None)
|
||||
|
||||
|
||||
def test_validate_classifier_response_text_with_media_type_invalid() -> None:
|
||||
data = {"content_type": "text", "media_type": "image"}
|
||||
assert validate_classifier_response(data) is None
|
||||
|
||||
|
||||
def test_validate_classifier_response_media_valid() -> None:
|
||||
for m_type in ("video", "image", "images", "embed", "mixed"):
|
||||
data = {"content_type": "media", "media_type": m_type}
|
||||
result = validate_classifier_response(data)
|
||||
assert result == MediaClassification(content_type="media", media_type=m_type)
|
||||
|
||||
|
||||
def test_validate_classifier_response_media_with_null_invalid() -> None:
|
||||
data = {"content_type": "media", "media_type": None}
|
||||
assert validate_classifier_response(data) is None
|
||||
|
||||
|
||||
def test_validate_classifier_response_unknown_values() -> None:
|
||||
assert validate_classifier_response({"content_type": "audio", "media_type": None}) is None
|
||||
assert validate_classifier_response({"content_type": "media", "media_type": "podcast"}) is None
|
||||
|
||||
|
||||
def test_validate_classifier_response_missing_or_extra_fields() -> None:
|
||||
assert validate_classifier_response({"content_type": "text"}) is None
|
||||
assert validate_classifier_response({"content_type": "text", "media_type": None, "extra": 123}) is None
|
||||
assert validate_classifier_response("not_a_dict") is None
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 2. Testes do Gate Estrutural DOM (detect_candidate_media)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_detect_candidate_media_video() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Título com Vídeo</h1>
|
||||
<video controls><source src="movie.mp4" type="video/mp4"></video>
|
||||
<p>Texto curto da matéria.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is True
|
||||
assert info.has_video is True
|
||||
assert info.image_count == 0
|
||||
assert info.has_embed is False
|
||||
|
||||
|
||||
def test_detect_candidate_media_isolated_source_not_video() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Artigo com source isolado</h1>
|
||||
<audio><source src="podcast.mp3"></audio>
|
||||
<p>Apenas texto e áudio.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_video is False
|
||||
assert info.has_candidate_media is False
|
||||
|
||||
|
||||
def test_detect_candidate_media_image_wrappers_single_count() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Notícia com Imagens em Wrappers</h1>
|
||||
<figure>
|
||||
<img src="foto1.jpg" alt="Foto 1">
|
||||
<figcaption>Legenda 1</figcaption>
|
||||
</figure>
|
||||
<picture>
|
||||
<source srcset="foto2.webp">
|
||||
<img src="foto2.jpg" alt="Foto 2">
|
||||
</picture>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is True
|
||||
assert info.image_count == 2
|
||||
assert info.has_video is False
|
||||
|
||||
|
||||
def test_detect_candidate_media_multiple_images() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<main>
|
||||
<h1>Galeria de Fotos</h1>
|
||||
<img src="img1.jpg">
|
||||
<img src="img2.jpg">
|
||||
<img src="img3.jpg">
|
||||
<p>Galeria completa do evento.</p>
|
||||
</main>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is True
|
||||
assert info.image_count == 3
|
||||
|
||||
|
||||
def test_detect_candidate_media_embeds() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<div role="main">
|
||||
<h1>Post Incorporado</h1>
|
||||
<iframe src="https://platform.twitter.com/embed/Tweet.html"></iframe>
|
||||
<p>Veja o post acima.</p>
|
||||
</div>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is True
|
||||
assert info.has_embed is True
|
||||
|
||||
|
||||
def test_detect_candidate_media_in_page_furniture_ignored() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<header>
|
||||
<img src="logo.png" alt="Logo">
|
||||
<nav><img src="icon.png"></nav>
|
||||
</header>
|
||||
<article>
|
||||
<h1>Matéria Textual Pura</h1>
|
||||
<p>Texto jornalístico longo sem mídia interna no corpo.</p>
|
||||
<p>Segundo parágrafo informativo.</p>
|
||||
</article>
|
||||
<aside>
|
||||
<iframe src="banner.html"></iframe>
|
||||
<img src="ad.jpg">
|
||||
</aside>
|
||||
<footer>
|
||||
<img src="partner.png">
|
||||
</footer>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is False
|
||||
assert info.image_count == 0
|
||||
assert info.has_video is False
|
||||
assert info.has_embed is False
|
||||
|
||||
|
||||
def test_detect_candidate_media_structural_region_precedence() -> None:
|
||||
# <article> deve ter precedência sobre <main> e <body>
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<main>
|
||||
<img src="outside_article.jpg">
|
||||
<article>
|
||||
<h1>Dentro do Article</h1>
|
||||
<video src="article_video.mp4"></video>
|
||||
<p>Parágrafo do artigo.</p>
|
||||
</article>
|
||||
</main>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is True
|
||||
assert info.has_video is True
|
||||
assert info.image_count == 0
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 3. Testes de Montagem do Compact Payload (build_compact_payload)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_build_compact_payload_preserves_text_and_unicode() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<head><title>Título da Publicação — Jornal Exemplo</title></head>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Título da Publicação</h1>
|
||||
<p>Primeiro parágrafo com acentuação: <em>ação, notícia & análise</em>.</p>
|
||||
<p>Segundo parágrafo com texto em espanhol: ¿Cómo estás? Fútbol y pasión.</p>
|
||||
<video src="clip.mp4"></video>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
candidate_info = MediaCandidateInfo(has_candidate_media=True, has_video=True, image_count=0, has_embed=False)
|
||||
payload = build_compact_payload(soup, candidate_info)
|
||||
|
||||
assert "Título da Publicação" in payload
|
||||
assert "ação, notícia & análise" in payload
|
||||
assert "¿Cómo estás? Fútbol y pasión." in payload
|
||||
assert "Video=True" in payload
|
||||
assert "ImagesCount=0" in payload
|
||||
assert "Embed=False" in payload
|
||||
assert "<p>" not in payload
|
||||
assert "<article>" not in payload
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 4. Testes de Bypass Direto do Gate Sem LLM (User Story 2 / T010)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_detect_candidate_media_pure_text_bypasses_llm() -> None:
|
||||
html = """
|
||||
<html>
|
||||
<body>
|
||||
<article>
|
||||
<h1>Editorial Textual Pura</h1>
|
||||
<p>Primeiro parágrafo de notícia densa.</p>
|
||||
<p>Segundo parágrafo detalhando o acontecimento político e econômico.</p>
|
||||
<p>Terceiro parágrafo com as conclusões e declarações oficiais.</p>
|
||||
</article>
|
||||
</body>
|
||||
</html>
|
||||
"""
|
||||
soup = BeautifulSoup(html, "html.parser")
|
||||
info = detect_candidate_media(soup)
|
||||
assert info.has_candidate_media is False
|
||||
assert info.has_video is False
|
||||
assert info.image_count == 0
|
||||
assert info.has_embed is False
|
||||
|
||||
|
||||
# ==============================================================================
|
||||
# 5. Testes da Cadeia Sequencial de Provedores e Invariância de Prompt (User Story 3 / T013)
|
||||
# ==============================================================================
|
||||
|
||||
|
||||
def test_classify_media_content_first_valid_wins_ollama(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
# Valida que prompt não solicita reasoning, keywords, summary, etc.
|
||||
forbidden = ["reasoning", "confidence", "rationale", "summary", "keywords", "evidence", "tradução"]
|
||||
for word in forbidden:
|
||||
assert word not in MEDIA_CLASSIFIER_PROMPT.lower()
|
||||
|
||||
metrics = {"fallback_groq": 0, "fallback_omniroute": 0, "classification_failed": 0}
|
||||
call_log: list[str] = []
|
||||
|
||||
def mock_http(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
call_log.append(url)
|
||||
# Ollama responde válido
|
||||
return 200, json.dumps({"message": {"content": json.dumps({"content_type": "media", "media_type": "video"})}})
|
||||
|
||||
monkeypatch.setattr("scripts.extract_article_contents._http_post_json", mock_http)
|
||||
|
||||
res, err = classify_media_content("payload", metrics, silent=True)
|
||||
assert res == MediaClassification(content_type="media", media_type="video")
|
||||
assert err is None
|
||||
assert len(call_log) == 1
|
||||
assert "11434" in call_log[0]
|
||||
assert metrics["fallback_groq"] == 0
|
||||
assert metrics["fallback_omniroute"] == 0
|
||||
|
||||
|
||||
def test_classify_media_content_ollama_fail_groq_success(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
monkeypatch.setenv("GROQ_API_KEY", "test-key")
|
||||
|
||||
metrics = {"fallback_groq": 0, "fallback_omniroute": 0, "classification_failed": 0}
|
||||
call_log: list[str] = []
|
||||
|
||||
def mock_http(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
call_log.append(url)
|
||||
if "11434" in url:
|
||||
# Ollama falha
|
||||
raise urllib.error.URLError("Connection refused")
|
||||
# Groq responde válido
|
||||
assert payload["model"] == "openai/gpt-oss-20b"
|
||||
assert payload["reasoning_effort"] == "low"
|
||||
assert payload["temperature"] == 0.0
|
||||
return 200, json.dumps({"choices": [{"message": {"content": json.dumps({"content_type": "text", "media_type": None})}}]})
|
||||
|
||||
monkeypatch.setattr("scripts.extract_article_contents._http_post_json", mock_http)
|
||||
|
||||
res, err = classify_media_content("payload", metrics, silent=True)
|
||||
assert res == MediaClassification(content_type="text", media_type=None)
|
||||
assert err is None
|
||||
assert len(call_log) == 2
|
||||
assert metrics["fallback_groq"] == 1
|
||||
assert metrics["fallback_omniroute"] == 0
|
||||
|
||||
|
||||
def test_classify_media_content_all_fail(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
monkeypatch.setenv("GROQ_API_KEY", "test-key")
|
||||
monkeypatch.setenv("OMNIROUTE_ENDPOINT", "http://omniroute.local")
|
||||
|
||||
metrics = {"fallback_groq": 0, "fallback_omniroute": 0, "classification_failed": 0}
|
||||
|
||||
def mock_http(url: str, payload: dict, headers: dict, timeout: int) -> tuple[int, str]:
|
||||
raise urllib.error.URLError("Server unreachable")
|
||||
|
||||
monkeypatch.setattr("scripts.extract_article_contents._http_post_json", mock_http)
|
||||
|
||||
res, err = classify_media_content("payload", metrics, silent=True)
|
||||
assert res is None
|
||||
assert err is not None
|
||||
assert metrics["fallback_groq"] == 1
|
||||
assert metrics["fallback_omniroute"] == 1
|
||||
|
||||
|
||||
Reference in New Issue
Block a user