Files
TextNLPClassifierApp/specs/004-deterministic-content-selection/spec.md
T

12 KiB

Feature Specification: Deterministic Content Selection

Feature Branch: 004-deterministic-content-selection
Created: 2026-08-20
Status: Draft
Input: User description: "usando o PRD: docs/prd_deterministic_content_selection.md"

User Scenarios & Testing (mandatory)

User Story 1 - Deterministic Selection with Text Consensus (Priority: P1)

As a data pipeline consumer or analyst, I want the system to automatically analyze the extracted text from Trafilatura, Newspaper4k, and Readability for each article and pick the single best extractor using consensus and coverage scoring, so that our dataset has high-quality, standardized content without human review.

Why this priority: Core value of the feature. Resolves the primary dilemma of choosing between 3 extractor outputs per article based on mutual agreement (consensus shingles) and concise content.

Independent Test: Can be tested independently by running the selection algorithm on articles where extractors have high agreement or partial variations, verifying that the extractor with highest F1 score (or closest score with fewest excess shingles) is selected.

Acceptance Scenarios:

  1. Given an article with usable extracts from all 3 libraries where 2 or 3 libraries agree closely, When selection is evaluated, Then the library with the highest consensus F1-score (or the more concise candidate within a 0.03 technical tie margin) is set in selected_extractor.
  2. Given an extractor with excess boilerplate/noise and two extractors with clean common content, When selection is evaluated, Then the noisy extractor suffers lower support score and the clean agreeing extractor is selected.
  3. Given an extractor with only a small snippet and two extractors with complete text, When selection is evaluated, Then the short snippet loses due to low consensus coverage.

User Story 2 - Resilient Decision Under Total Disagreement or Degradation (Priority: P2)

As a pipeline maintainer, I want the selection algorithm to make a deterministic and sensible fallback choice even when extractors completely disagree, produce errors, or return empty/degraded content, so that the pipeline never halts or leaves an article without a chosen extractor.

Why this priority: Essential for pipeline stability. The system must guarantee that every article gets an unambiguous winner without throwing runtime exceptions or generating null/ambiguous states.

Independent Test: Can be tested with synthetic articles representing edge cases: all extractors returning non-overlapping text, extractors reporting errors, or all extractors failing.

Acceptance Scenarios:

  1. Given 3 active candidates with 0 consensus shingles, When selection runs, Then the candidate with the median shingle length is selected.
  2. Given 2 active candidates with 0 consensus shingles, When selection runs, Then the candidate with the larger shingle count is selected.
  3. Given an article where all usable candidates are absent but degraded candidates exist, When selection runs, Then the algorithm evaluates only the degraded candidates.
  4. Given an article where all 3 extractors failed or returned empty content, When selection runs, Then newspaper4k is selected via the mandatory final fallback rule.

User Story 3 - Non-Destructive JSON Batch Processing (Priority: P3)

As a system operator, I want to pass a JSON file with an articles array, execute the deterministic selector, and receive a new file <original_name>_selected.json with all original data and order intact plus the selected_extractor field, leaving the original file completely untouched.

Why this priority: Guarantees data preservation, idempotency, and clean pipeline integration.

Independent Test: Can be tested by running the process on a full batch JSON file (such as out/river_plate_extracted.json) and comparing input vs output keys, element counts, article order, and field contents.

Acceptance Scenarios:

  1. Given a valid JSON file with N articles, When the batch selection is executed, Then a new file <original_name>_selected.json is generated containing exactly N articles in identical order, each with all original fields plus selected_extractor.
  2. Given an input JSON file where selected_extractor already exists, When the batch selection is executed, Then selected_extractor is recalculated and updated.
  3. Given an invalid JSON file or a file where articles is not a list, When execution runs, Then the process terminates with an error and does not produce a partial or corrupted output file.

Edge Cases

  • Empty articles list ([]): Produces a valid output JSON containing an empty articles: [] list without errors.
  • Exact score & shingle count tie: Resolved deterministically by the strict fallback hierarchy: newspaper4k > readability > trafilatura.
  • Single active candidate: When only 1 library produces usable output, it is selected immediately without computing consensus.
  • Short texts (< 5 tokens): When candidate text has between 1 and 4 tokens, the entire token sequence forms a single shingle.
  • Malformed fields / Type mismatch: If a content field is not a string or missing, the candidate is classified as unavailable.
  • Missing library block: If an article does not contain a trafilatura, newspaper4k, or readability block, that candidate is treated as unavailable.

Requirements (mandatory)

Functional Requirements

  • FR-001: System MUST accept a valid JSON file path containing a root object with an articles array.
  • FR-002: System MUST validate input structure (root is object, articles is list) and terminate immediately without creating an output file if validation fails.
  • FR-003: System MUST process all articles in articles, preserving their exact sequence and all existing fields and values without modification.
  • FR-004: System MUST evaluate extractor candidates using exclusively:
    • trafilatura.text for Trafilatura
    • newspaper4k.text for Newspaper4k
    • readability.cleaned_text for Readability
  • FR-005: System MUST categorize each candidate into one of three states:
    • Usable: Content is non-empty string after normalization and extractor error is null/empty.
    • Degraded: Content is non-empty string after normalization but extractor error is non-null.
    • Unavailable: Content is missing, not a string, or empty after normalization.
  • FR-006: System MUST form the active candidate set per article: Usable candidates if any exist; otherwise Degraded candidates if any exist; otherwise trigger final fallback.
  • FR-007: System MUST perform deterministic in-memory normalization for candidate comparisons:
    1. Decode HTML entities.
    2. Strip Markdown images (![alt](url)), removing non-textual media embeds.
    3. In Markdown links ([text](url)), preserve anchor text and strip URL targets.
    4. Strip HTML tags, maintaining spacing between adjacent words.
    5. Apply Unicode NFKC normalization.
    6. Convert to lowercase.
    7. Collapse multiple whitespace/newlines/tabs into a single space.
    8. Tokenize retaining Unicode letters and digits.
    9. Ignore punctuation symbols.
  • FR-008: System MUST generate 5-token sliding window shingles from the ordered token sequence of each candidate (or single $N$-token shingle if 1 \le N \le 4).
  • FR-009: System MUST construct the consensus shingle set (shingles appearing in at least 2 active candidates).
  • FR-010: System MUST compute coverage, support, and score (F_1 = 2 \times \text{coverage} \times \text{support} / (\text{coverage} + \text{support})) for each active candidate against the consensus shingles (or 0 if denominator is 0).
  • FR-011: When consensus shingles exist, the system MUST:
    1. Sort candidates descending by score.
    2. Identify all candidates within a 0.03 difference from the top score (technical tie pool).
    3. If technical tie pool has 1 candidate, select it.
    4. If multiple candidates are in technical tie, select the one with the smallest total shingle count (least surplus).
    5. If shingle count is also tied, apply priority hierarchy: newspaper4k > readability > trafilatura.
  • FR-012: When 0 consensus shingles exist, the system MUST:
    • With 3 active candidates: select candidate with median shingle count.
    • With 2 active candidates: select candidate with maximum shingle count.
    • With 1 active candidate: select that single candidate.
    • In shingle count ties: apply priority hierarchy (newspaper4k > readability > trafilatura).
  • FR-013: When 0 active candidates exist (all unavailable), system MUST assign newspaper4k.
  • FR-014: System MUST inject or replace selected_extractor in each article item with exactly one value from {"trafilatura", "newspaper4k", "readability"}.
  • FR-015: System MUST never output null, empty string, ambiguous, or leave an article without a selection.
  • FR-016: System MUST write the result atomically to <original_name_without_extension>_selected.json in the same directory or specified target, leaving the input file unchanged.
  • FR-017: System MUST produce 100% deterministic and identical outputs across repeated runs with identical inputs.

Key Entities (include if feature involves data)

  • Article Input Batch: Root JSON container with metadata and an ordered list of articles.
  • Article Record: Object representing an article, containing source metadata, extraction results from the 3 extractors (trafilatura, newspaper4k, readability), and the resulting selected_extractor tag.
  • Extractor Candidate: Evaluation model for an individual extractor containing raw text, error state, candidate usability state (Usable, Degraded, Unavailable), normalized token stream, 5-token shingles, and computed metrics (coverage, support, score, shingle_count).
  • Consensus Shingle Set: Set of unique 5-token shingles shared by 2 or more active extractor candidates.

Success Criteria (mandatory)

Measurable Outcomes

  • SC-001: 100% Selection Completeness: 100% of articles in the input collection receive a valid selected_extractor value from the closed set ['trafilatura', 'newspaper4k', 'readability'].
  • SC-002: 0% Ambiguity: Exactly 0 articles result in null, missing, empty, or ambiguous selection states.
  • SC-003: 100% Deterministic Reproducibility: 100% identical selected_extractor values when executing across multiple runs on identical input datasets.
  • SC-004: 100% Non-Destructive Integrity: 100% of pre-existing keys, nested objects, article counts, and article ordering are preserved identically in the output JSON.
  • SC-005: 100% Test Case Coverage: Passes 100% of defined mandatory test cases (CT-001 through CT-014).
  • SC-006: Atomic Operation: 0 partial or corrupted output files generated on process failure or invalid JSON inputs.

Assumptions

  • The input JSON is generated by the extraction pipeline and contains articles where each item may have trafilatura, newspaper4k, and readability sub-objects.
  • All three extraction libraries operated on the exact same HTML source document.
  • No external dependencies (LLM APIs, embedding services, or network calls) are permitted during the selection process.
  • Unicode NFKC normalization and standard tokenization cover multilingual article content (e.g. Portuguese, Spanish, English).
  • Default output file path naming convention <name>_selected.json is sufficient, with CLI support for optional custom destination.