feat(runtime): implement single-article consolidation runtime and modularize codebase

This commit is contained in:
2026-08-24 00:14:07 -03:00
parent e1e0be1353
commit 23de7d8fe7
176 changed files with 266754 additions and 10179 deletions
@@ -0,0 +1,202 @@
# Quickstart Guide: Article Consolidation and Hygiene Runtime
**Feature Branch**: `006-article-consolidation-runtime`
**Date**: 2026-08-23
**Status**: Complete
---
## 1. Prerequisites & Environment Setup
The runtime is built using the repository's standard Python packaging defined in `pyproject.toml` (`requires-python = ">=3.10"`).
### Installation
```bash
# Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate # ou .venv\Scripts\activate
# Install package and dependencies in editable mode
pip install -e .
```
### Environment Secrets
Set provider credentials and Langfuse secrets via environment variables:
```bash
# Model Gateway Provider Keys
export GROQ_API_KEY="gsk_samplekey1234567890abcdef"
export DEEPSEEK_API_KEY="sk_samplekey1234567890abcdef"
# Observability
export LANGFUSE_PUBLIC_KEY="pk-lf-samplepublickey123456"
export LANGFUSE_SECRET_KEY="sk-lf-samplesecretkey123456"
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
```
---
## 2. Local Validation Configuration (Fixture Only)
> [!NOTE]
> The configuration below is a local development fixture (`runtime_config.local.json`), certified against packaged `src/core/release-metadata.json`. Timeouts, retries, pricing, and limits shown here are exclusive to local testing fixtures and do not represent production calibrated defaults.
```json
{
"config_version": "1.0.0",
"paths": {
"output_dir": "./out",
"sqlite_db": "./out/runtime_state.db"
},
"roles": {
"runtime_primary": {
"role_config_version": "1.0.0",
"provider": "groq",
"model": "openai/gpt-oss-20b",
"endpoint_url": "https://api.groq.com/openai/v1/chat/completions",
"timeout_seconds": 15,
"max_retries": 2,
"parameters": { "temperature": 0.0 },
"hygiene_prompt_version": "1.0.0",
"hygiene_schema_version": "1.0.0",
"enrichment_prompt_version": "1.0.0",
"enrichment_schema_version": "1.0.0"
},
"runtime_fallback": {
"role_config_version": "1.0.0",
"provider": "deepseek",
"model": "deepseek-v4-flash",
"endpoint_url": "https://api.deepseek.com/v1/chat/completions",
"timeout_seconds": 15,
"max_retries": 2,
"parameters": { "temperature": 0.0 },
"hygiene_prompt_version": "1.0.0",
"hygiene_schema_version": "1.0.0",
"enrichment_prompt_version": "1.0.0",
"enrichment_schema_version": "1.0.0"
}
},
"prompts": {
"article_content_hygiene": {
"path": "prompts/article_content_hygiene.v1.txt",
"version": "1.0.0",
"hash": "f8a9c2b1d3e4f5a6b7c8d9e0f1a2b3c4d5e6f7a8b9c0d1e2f3a4b5c6d7e8f9a0"
},
"article_sentiment_tags": {
"path": "prompts/article_sentiment_tags.v1.txt",
"version": "1.0.0",
"hash": "d4e1b7a2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0"
}
},
"ecp": {
"canonical_schema_reference": "src/adapters/ecp/schemas/ecp-profile.schema.json",
"classifier_module": "src.classifier.InherenceClassifier"
},
"limits": {
"max_input_bytes": 5242880,
"context_strategy": "fail_before_provider"
},
"pricing": {
"primary_input_1k": 0.0001,
"primary_output_1k": 0.0002,
"fallback_input_1k": 0.0001,
"fallback_output_1k": 0.0002
},
"langfuse": {
"environment": "local_dev",
"trace_content_policy": "metadata_only"
},
"sqlite": {
"busy_timeout_ms": 5000
}
}
```
---
## 3. Operational Execution Scenarios
### Scenario A: Preflight Validation
Validates local configuration file exact byte hash against packaged `src/core/release-metadata.json`, prompt file hashes, schema compatibility, filesystem permissions, clock sync, and validated credentials before accepting live traffic.
```bash
python -m src.cli.preflight --config runtime_config.local.json
```
*Expected Outcome*: Exit code `0`, structured JSON confirmation on stdout that all preflight checks passed.
---
### Scenario B: End-to-End Processing (Approved Article)
Executes consolidation on an article with direct ECP inherence.
```bash
python -m src.cli.consolidate \
--input-article examples/sample_article_valid.json \
--ecp-snapshot examples/sample_ecp_snapshot.json \
--config runtime_config.local.json
```
*Expected Outcome*:
- Exit code `0`.
- Output manifest `./out/<fingerprint>.result.json` created with `final_status: "completed_text"` and `generate_markdown: true`.
- Published Markdown `./out/<fingerprint>.md` created with YAML front matter.
- SQLite state table updated to `completed_text`.
- stdout emits the complete JSON manifest matching `manifest-output.schema.json`.
---
### Scenario C: ECP Rejection (Zero Markdown)
Executes consolidation on a non-inherent article (`TANGENTIAL` or `NOT_RELATED`).
```bash
python -m src.cli.consolidate \
--input-article examples/sample_article_tangential.json \
--ecp-snapshot examples/sample_ecp_snapshot.json \
--config runtime_config.local.json
```
*Expected Outcome*:
- Exit code `0`.
- Manifest `./out/<fingerprint>.result.json` created with `final_status: "rejected_ecp"` and `generate_markdown: false`.
- **Zero `<fingerprint>.md` file created for this execution fingerprint**.
- SQLite state table updated to `ecp_rejected`.
- stdout emits the complete JSON manifest matching `manifest-output.schema.json`.
---
### Scenario D: Idempotent Execution & Replay
Re-executes the command on a previously completed article fingerprint.
```bash
python -m src.cli.consolidate \
--input-article examples/sample_article_valid.json \
--ecp-snapshot examples/sample_ecp_snapshot.json \
--config runtime_config.local.json
```
*Expected Outcome*:
- Immediate return without re-invoking LLM providers.
- Output JSON manifest on stdout matches previously persisted result.
---
## 4. Automated Quality Evaluation & Test Commands
```bash
# 1. Run static multi-parser zero-regex check (Python AST + JSON pattern check + Promptfoo YAML check)
python -m tests.scripts.check_zero_regex
# 2. Run unit and all 9 contract tests (including the 20 real reference units)
pytest tests/unit tests/contract -v
# 3. Run security tests (SEC-001 to SEC-008)
pytest tests/security -v
# 4. Run load test (100 articles/hour benchmark)
pytest tests/load -v
# 5. Run mock integration tests, fault injection (10 scenarios), and operations resilience (FR-081)
pytest tests/integration tests/fault_injection -v
# 6. Run offline Promptfoo evaluations
npx promptfoo eval -c evals/promptfoo.config.yaml
```