feat(classifier): add multilingual ECP inherence classifier POC

This commit is contained in:
2026-08-20 00:51:02 -03:00
parent d371b81aa4
commit 67cc40f91a
175 changed files with 30399 additions and 703 deletions
+132 -37
View File
@@ -1,20 +1,26 @@
# Graph Report - TextNLPClassifierApp (2026-08-19)
# Graph Report - TextNLPClassifierApp (2026-08-20)
## Corpus Check
- cluster-only mode — file stats not available
- 133 files · ~54,759 words
- Verdict: corpus is large enough that graph structure adds value.
## Summary
- 310 nodes · 284 edges · 40 communities (34 shown, 6 thin omitted)
- Extraction: 100% EXTRACTED · 0% INFERRED · 0% AMBIGUOUS
- Token cost: 1,734 input · 302 output
- 595 nodes · 704 edges · 80 communities (43 shown, 37 thin omitted)
- Extraction: 96% EXTRACTED · 4% INFERRED · 0% AMBIGUOUS · INFERRED: 31 edges (avg confidence: 0.95)
- Token cost: 0 input · 0 output
## Graph Freshness
- Built from commit: `d371b81a`
- Run `git rev-parse HEAD` and compare to check if the graph is stale.
- Run `graphify update .` after code changes (no API cost).
## Community Hubs (Navigation)
- Task Planning
- Convergence Workflow
- SpecKit Utilities
- Graphify Commands
- Specification Analysis
- Analysis Detection
- speckit-analyze/SKILL.md
- Tasks: Multilingual NLP Entity Inherence Classifier (POC)
- Feature Specification Template
- Graphify Rules
- Implementation Planning
@@ -45,26 +51,75 @@
- Media Transcription
- Extraction Specification
- Graphify Workflows
- Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)
- 1. Technical Decisions & Tradeoffs
- 1. Input Schemas
- 2. Basic CLI Usage Examples
- 2. Standard Streams & Exit Codes
- ECPSnapshot
- test_models.py
- detect_language
- main
- content_northvolt_de.md
- content_presal_pt.md
- content_tangential_es.md
- adapters/__init__.py
- src/__init__.py
- de/contextual.md
- de/direct.md
- de/not_related.md
- de/tangential.md
- en/contextual.md
- en/direct.md
- en/not_related.md
- en/tangential.md
- es/contextual.md
- es/direct.md
- es/not_related.md
- es/tangential.md
- fr/contextual.md
- fr/direct.md
- fr/not_related.md
- fr/tangential.md
- it/contextual.md
- it/direct.md
- it/not_related.md
- it/tangential.md
- pt/contextual.md
- pt/direct.md
- pt/not_related.md
- pt/tangential.md
- tests/__init__.py
- text-nlp-classifier
## God Nodes (most connected - your core abstractions)
1. `Tasks: [FEATURE NAME]` - 13 edges
2. `What You Must Do When Invoked` - 12 edges
3. `/graphify` - 10 edges
4. `graphify reference: extra exports and benchmark` - 8 edges
5. `Ponytail` - 8 edges
6. `Execution Steps` - 7 edges
7. `Ponytail Help` - 7 edges
8. `4. Detection Passes (Token-Efficient Analysis)` - 7 edges
9. `Execution Steps` - 7 edges
10. `Core Principles` - 6 edges
1. `ECPSnapshot` - 31 edges
2. `InherenceClassifier` - 25 edges
3. `DecisionCategory` - 18 edges
4. `ClassificationResult` - 17 edges
5. `LocalEmbeddingsAdapter` - 14 edges
6. `LLMFallbackAdapter` - 14 edges
7. `detect_language()` - 14 edges
8. `main()` - 13 edges
9. `Tasks: [FEATURE NAME]` - 13 edges
10. `BaseNLPAdapter` - 12 edges
## Surprising Connections (you probably didn't know these)
- None detected - all connections are within the same source files.
- `emit_error()` --uses--> `ErrorCode` [INFERRED]
classify.py → src/models.py
- `main()` --uses--> `ECPSnapshot` [INFERRED]
classify.py → src/models.py
- `main()` --uses--> `ErrorCode` [INFERRED]
classify.py → src/models.py
- `test_classification_error_serialization()` --uses--> `ErrorCode` [INFERRED]
tests/test_models.py → src/models.py
- `test_ecp_snapshot_defaults()` --uses--> `ECPSnapshot` [INFERRED]
tests/test_models.py → src/models.py
## Import Cycles
- None detected.
## Communities (40 total, 6 thin omitted)
## Communities (80 total, 37 thin omitted)
### Community 0 - "Task Planning"
Cohesion: 0.07
@@ -82,13 +137,13 @@ Nodes (13): Find-SpecifyRoot(), Format-SpecKitCommand(), Get-CurrentBranch(), Ge
Cohesion: 0.08
Nodes (24): For /graphify add and --watch, For /graphify query, For the commit hook and native CLAUDE.md integration, For --update and --cluster-only, /graphify, Honesty Rules, Interpreter guard for subcommands, Part A - Structural extraction for code files (+16 more)
### Community 4 - "Specification Analysis"
Cohesion: 0.15
Nodes (12): 7. Provide Next Actions, 8. Offer Remediation, 9. Check for extension hooks, Analysis Guidelines, Context, Context Efficiency, Goal, Operating Constraints (+4 more)
### Community 4 - "speckit-analyze/SKILL.md"
Cohesion: 0.08
Nodes (25): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, 7. Provide Next Actions, 8. Offer Remediation (+17 more)
### Community 5 - "Analysis Detection"
Cohesion: 0.15
Nodes (13): 1. Initialize Analysis Context, 2. Load Artifacts (Progressive Disclosure), 3. Build Semantic Models, 4. Detection Passes (Token-Efficient Analysis), 5. Severity Assignment, 6. Produce Compact Analysis Report, A. Duplication Detection, B. Ambiguity Detection (+5 more)
### Community 5 - "Tasks: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.06
Nodes (33): 1. Requirement Completeness & Scope Boundaries, 2. Requirement Clarity & Decision Semantics, 3. Requirement Consistency & Alignment, 4. Acceptance Criteria & Measurability, 5. Scenario & Edge Case Coverage, Notes, POC Readiness & Requirements Quality Checklist: Multilingual NLP Entity Inherence Classifier, Content Quality (+25 more)
### Community 6 - "Feature Specification Template"
Cohesion: 0.15
@@ -186,21 +241,61 @@ Nodes (3): For --cluster-only, For --update (incremental re-extraction), graphif
Cohesion: 0.50
Nodes (3): Boundaries, Output, Scan
### Community 40 - "Feature Specification: Multilingual NLP Entity Inherence Classifier (POC)"
Cohesion: 0.14
Nodes (14): Assumptions, Assumptions & Scope, Clarifications, Explicit Out of Scope (POC), Feature Specification: Multilingual NLP Entity Inherence Classifier (POC), Functional Requirements, Key Entities *(data models & domain entities)*, Measurable Outcomes (+6 more)
### Community 41 - "1. Technical Decisions & Tradeoffs"
Cohesion: 0.22
Nodes (8): 1. Technical Decisions & Tradeoffs, 2. Standardized Error Handling Strategy, Decision 1: Execution Engine & CLI Architecture, Decision 2: Multilingual Language Detection & Normalization (Tier 1 Core), Decision 3: Materialized ECP Snapshot Contract & Matching Logic, Decision 4: Tier 2 (Embeddings) & Tier 3 (LLM) Optional Adapters, Decision 5: Controlled 24-Case POC Benchmark Suite, Technical Research & Architecture Decisions (POC)
### Community 42 - "1. Input Schemas"
Cohesion: 0.25
Nodes (7): 1.1 ECP Snapshot Schema (`snapshot.json`), 1.2 Content Item Schema (`content.md`), 1. Input Schemas, 2.1 Classification Success Result Schema (`result.json`), 2.2 Error Result Schema, 2. Output Schemas, Data Models & Schemas (POC)
### Community 43 - "2. Basic CLI Usage Examples"
Cohesion: 0.25
Nodes (7): 1. Prerequisites & Installation, 2.1 Direct Inherence (Portuguese), 2.2 Contextual Inherence via Graph Snapshot (German), 2.3 Tangential Mention (Spanish), 2. Basic CLI Usage Examples, 3. Running the Controlled 24-Case Benchmark, Quickstart & Validation Guide (POC)
### Community 44 - "2. Standard Streams & Exit Codes"
Cohesion: 0.29
Nodes (6): 1.1 Arguments & Options, 1. Command Line Interface, 2.1 Exit Codes, 2.2 Standard Output (`stdout`) / Standard Error (`stderr`), 2. Standard Streams & Exit Codes, CLI Contract & Interface Specification (POC)
### Community 45 - "ECPSnapshot"
Cohesion: 0.05
Nodes (63): ABC, Enum, parametrize, BaseNLPAdapter, Base abstract adapter interface for optional Tier 2 / Tier 3 NLP enhancers., Abstract interface for pluggable NLP classification adapters., Return True if the underlying provider or model is installed and configured., Compute semantic similarity score between text and a set of candidate terms. (+55 more)
### Community 46 - "test_models.py"
Cohesion: 0.12
Nodes (17): Any, emit_error(), ClassificationError, extract_evidence_snippets(), extract_sentences(), Markdown content parser and excerpt extraction utilities., Remove markdown syntax markers (headers, bold, italics, links, code blocks) to…, Split text into individual sentences. (+9 more)
### Community 47 - "detect_language"
Cohesion: 0.19
Nodes (16): detect_language(), extract_words(), normalize_text(), Lightweight multilingual language detection and text normalization., Normalize text by converting to lowercase and stripping combining diacritical…, Tokenize text into lowercase alphanumeric words., Detect the ISO-639-1 language code of text among supported languages (pt, en,…, Unit tests for language detection and text normalization. (+8 more)
### Community 48 - "main"
Cohesion: 0.31
Nodes (9): main(), parse_args(), Namespace, CLI execution tests covering flags, arguments, stdout, and error handling., test_cli_empty_content_file(), test_cli_missing_ecp_file(), test_cli_missing_required_ecp_field(), test_cli_output_file() (+1 more)
## Knowledge Gaps
- **212 isolated node(s):** `Format: `[ID] [P?] [Story] Description``, `Implementation for User Story 1`, `Implementation for User Story 2`, `Implementation for User Story 3`, `Incremental Delivery` (+207 more)
- **291 isolated node(s):** `text-nlp-classifier`, `MatchedGraphEntity`, `graphify`, `Usage`, `What graphify is for` (+286 more)
These have ≤1 connection - possible missing edges or undocumented components.
- **6 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
- **37 thin communities (<3 nodes) omitted from report** — run `graphify query` to explore isolated nodes.
## Suggested Questions
_Questions this graph is uniquely positioned to answer:_
- **Why does `Execution Steps` connect `Analysis Detection` to `Specification Analysis`?**
_High betweenness centrality (0.004) - this node is a cross-community bridge._
- **What connects `Format: `[ID] [P?] [Story] Description``, `Implementation for User Story 1`, `Implementation for User Story 2` to the rest of the system?**
_212 weakly-connected nodes found - possible documentation gaps or missing edges._
- **Should `Task Planning` be split into smaller, more focused modules?**
_Cohesion score 0.07407407407407407 - nodes in this community are weakly interconnected._
- **Should `Convergence Workflow` be split into smaller, more focused modules?**
_Cohesion score 0.125 - nodes in this community are weakly interconnected._
- **Should `Graphify Commands` be split into smaller, more focused modules?**
_Cohesion score 0.08 - nodes in this community are weakly interconnected._
- **Why does `ECPSnapshot` connect `ECPSnapshot` to `main`, `test_models.py`?**
_High betweenness centrality (0.015) - this node is a cross-community bridge._
- **Why does `InherenceClassifier` connect `ECPSnapshot` to `main`?**
_High betweenness centrality (0.008) - this node is a cross-community bridge._
- **Why does `detect_language()` connect `detect_language` to `ECPSnapshot`?**
_High betweenness centrality (0.007) - this node is a cross-community bridge._
- **Are the 10 inferred relationships involving `ECPSnapshot` (e.g. with `main()` and `BaseNLPAdapter`) actually correct?**
_`ECPSnapshot` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 6 inferred relationships involving `InherenceClassifier` (e.g. with `LocalEmbeddingsAdapter` and `LLMFallbackAdapter`) actually correct?**
_`InherenceClassifier` has 6 INFERRED edges - model-reasoned connections that need verification._
- **Are the 10 inferred relationships involving `DecisionCategory` (e.g. with `InherenceClassifier` and `test_adversarial_apple_fruit_recipe()`) actually correct?**
_`DecisionCategory` has 10 INFERRED edges - model-reasoned connections that need verification._
- **Are the 4 inferred relationships involving `ClassificationResult` (e.g. with `BaseNLPAdapter` and `LocalEmbeddingsAdapter`) actually correct?**
_`ClassificationResult` has 4 INFERRED edges - model-reasoned connections that need verification._