feat(media-routing): implement 007 media article routing, runtime architecture diagram and update graphify knowledge graph
This commit is contained in:
@@ -0,0 +1,138 @@
|
||||
# Data Model & State Transitions: Classificação e Roteamento de Notícias de Mídia
|
||||
|
||||
**Feature**: `007-media-article-routing`
|
||||
**Date**: 2026-08-24
|
||||
**Status**: Completed
|
||||
|
||||
---
|
||||
|
||||
## 1. Diagrama Entidade-Relacionamento e Fluxo de Dados
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Input[InputArticle JSON] --> Crawler[Foxcape Crawler]
|
||||
Crawler --> LoadedPage[DOM Carregada: html, title, status]
|
||||
LoadedPage --> Gate[DOM Media Gate: detecção estrutural]
|
||||
|
||||
Gate -- "Sem mídia candidata relevante" --> TextPipeline[Multimotor Textual: Trafilatura + Newspaper4k + Readability]
|
||||
|
||||
Gate -- "Com mídia candidata relevante" --> CompactBuilder[Montagem de Payload Compacto]
|
||||
CompactBuilder --> Classifier[Cadeia Sequencial LLM: Ollama -> Groq -> OmniRoute]
|
||||
|
||||
Classifier -- "content_type = text" --> TextPipeline
|
||||
TextPipeline --> ExtractedRecord[ExtractedArticle: Sucesso Textual]
|
||||
ExtractedRecord --> MainReport[ExtractionBatchReport -> *_extracted.json]
|
||||
|
||||
Classifier -- "content_type = media" --> MediaRecord[MediaArticle: Registro de Mídia]
|
||||
MediaRecord --> MediaEnvelope[MediaOutputFile -> *_media.json]
|
||||
|
||||
Classifier -- "Falha total dos 3 provedores" --> FailureRecord[Registro de Falha de Classificação]
|
||||
FailureRecord --> MainReport
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Modelos de Dados (Dataclasses e Estruturas de Runtime)
|
||||
|
||||
### 2.1 `InputArticle`
|
||||
Representa a notícia original carregada do JSON de entrada.
|
||||
- `titulo: str` (Obrigatório)
|
||||
- `url: str` (Obrigatório)
|
||||
- `subtitulo: str | None` (Opcional, preservado conforme entrada)
|
||||
- `quando_publicado: str | None` (Opcional, preservado conforme entrada)
|
||||
- `pagina: int | None` (Opcional, preservado conforme entrada)
|
||||
- Campos adicionais da entrada são preservados no dicionário `input_meta`.
|
||||
|
||||
### 2.2 `MediaCandidateInfo` (Estrutura Interna do Gate)
|
||||
Resultado da análise estrutural da DOM:
|
||||
- `has_candidate_media: bool`
|
||||
- `has_video: bool`
|
||||
- `image_count: int`
|
||||
- `has_embed: bool`
|
||||
|
||||
### 2.3 `MediaClassification` (Estrutura Interna de Saída do LLM)
|
||||
Resultado estrito retornado pelo classificador LLM:
|
||||
- `content_type: Literal["text", "media"]`
|
||||
- `media_type: Literal["video", "image", "images", "embed", "mixed"] | None`
|
||||
- Regra semântica: `content_type == "text"` $\iff$ `media_type is None`.
|
||||
|
||||
### 2.4 `MediaArticle` (Persistência em `*_media.json`)
|
||||
Registro consolidado de publicação classificada como mídia:
|
||||
- `input_meta: dict[str, Any]` (Preserva todos os metadados recebidos da entrada)
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `content_type: Literal["media"]`
|
||||
- `media_type: Literal["video", "image", "images", "embed", "mixed"]`
|
||||
|
||||
### 2.5 `ExtractedArticle` (Persistência Textual em `*_extracted.json`)
|
||||
Representa exclusivamente artigos textuais que efetivamente passaram pelo multimotor (`extract_all_engines`):
|
||||
- `input_meta: InputArticle`
|
||||
- `extraction_status: Literal["success", "failed"]`
|
||||
- `error_message: str | None`
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `trafilatura: TrafilaturaData | None`
|
||||
- `newspaper4k: NewspaperData | None`
|
||||
- `readability: ReadabilityData | None`
|
||||
|
||||
### 2.6 Registro de Falha de Classificação (Persistência Inline em `*_extracted.json`)
|
||||
Para artigos que sofram falha operacional dos três provedores LLM, o registro é persistido inline no array `articles` do JSON principal sem ter executado os extratores textuais:
|
||||
- `input_meta: dict[str, Any]`
|
||||
- `crawled_url: str`
|
||||
- `page_title: str | None`
|
||||
- `http_status: int | None`
|
||||
- `classification_status: Literal["failed"]`
|
||||
- `error_message: str`
|
||||
- `trafilatura: None`
|
||||
- `newspaper4k: None`
|
||||
- `readability: None`
|
||||
|
||||
### 2.7 `ExtractionBatchReport`
|
||||
Envelope consolidado de saída gravado no arquivo principal (`*_extracted.json`):
|
||||
- `source_file: str`
|
||||
- `processed_at: str` (ISO 8601 UTC)
|
||||
- `total_articles: int` (Total de registros presentes no array `articles` do JSON principal)
|
||||
- `successful_articles: int` (Total de artigos textuais processados com sucesso)
|
||||
- `failed_articles: int` (Total de registros de falha presentes, incluindo erros de crawl e de classificação)
|
||||
- `articles: list[dict[str, Any] | ExtractedArticle]`
|
||||
|
||||
### 2.8 `MediaOutputFile`
|
||||
Envelope mínimo gravado no arquivo de mídia (`*_media.json`):
|
||||
- `articles: list[MediaArticle]`
|
||||
|
||||
---
|
||||
|
||||
## 3. Máquina de Estados do Processamento de Artigo
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> Crawling
|
||||
Crawling --> CrawlFailed: Erro HTTP / Timeout Foxcape
|
||||
CrawlFailed --> MainJSONRecord: Registra falha de crawl
|
||||
|
||||
Crawling --> StructuralInspection: DOM carregada com sucesso
|
||||
StructuralInspection --> TextExtraction: Sem mídia candidata relevante
|
||||
StructuralInspection --> LLMClassification: Mídia candidata relevante detectada
|
||||
|
||||
state LLMClassification {
|
||||
[*] --> TryOllama
|
||||
TryOllama --> ValidOutput: Resposta com schema válido
|
||||
TryOllama --> TryGroq: Falha operacional Ollama
|
||||
TryGroq --> ValidOutput: Resposta com schema válido
|
||||
TryGroq --> TryOmniRoute: Falha operacional Groq
|
||||
TryOmniRoute --> ValidOutput: Resposta com schema válido
|
||||
TryOmniRoute --> AllProvidersFailed: Falha operacional OmniRoute
|
||||
}
|
||||
|
||||
ValidOutput --> TextExtraction: content_type == text
|
||||
ValidOutput --> MediaRouting: content_type == media
|
||||
AllProvidersFailed --> MainJSONRecord: Registra classification_status = failed
|
||||
|
||||
TextExtraction --> MainJSONRecord: Executa Trafilatura + Newspaper + Readability
|
||||
MediaRouting --> MediaJSONRecord: Grava em *_media.json
|
||||
|
||||
MainJSONRecord --> [*]
|
||||
MediaJSONRecord --> [*]
|
||||
```
|
||||
Reference in New Issue
Block a user