Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Family-aware text splits and evaluation contamination

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Assign related text records to one evaluation partition and audit every outside corpus that a model may have seen.

Split at the right unit

A random row split is unsafe when one ticket appears as an original, a corrected export and a copied handoff. Assign the complete text family to training, validation or test. Preserve document-level labels and revision dates inside the group; sharing a family does not make every version interchangeable. Family review establishes the grouping evidence. Grouped temporal validation adds the harder requirement that a held-out event is later than the training data.

Freeze the evaluation boundary

Record split policy version, family registry version, corpus snapshot and model training sources. A fixed hash of family ID gives repeatable assignment, but it does not enforce time order, language balance or incident isolation by itself. Choose the unit from the question being tested: generalization to new tickets, new incidents, new organizations or future releases may require different groups. Report how many examples and families land in each partition; a tiny minority slice can disappear under an otherwise valid split.

Look beyond the labeled table

Contamination can arrive through pretraining, retrieval indexes, synthetic examples, prompt libraries or test examples used during tuning. Keep a manifest of accessible corpora and known benchmark exports. Search for exact and near-duplicate test spans in those sources, then classify exposure by where and when it occurred. A match is evidence for review, not automatic proof that the model memorized the answer. If training provenance is unavailable, state that limit in the evaluation report instead of claiming a clean split.

Gate future refreshes

When a new revision joins an existing family, inherit its partition or quarantine it until the registry is corrected. Recheck every training and retrieval source before adding an evaluation set. Track cross-partition family count, near-duplicate rate and untraceable source count. The leakage audit project includes a failed refresh that must be held before its score is reported.

Implementation

python
import hashlib

def partition_for_family(family_id):
    if not family_id:
        raise ValueError("family_id is required")
    digest = hashlib.sha256(family_id.encode("utf-8")).digest()
    bucket = int.from_bytes(digest[:8], "big") % 100
    if bucket < 70:
        return "train"
    if bucket < 85:
        return "validation"
    return "test"

records = [
    {"document_id": "ticket-47", "family_id": "incident-west-47"},
    {"document_id": "handoff-82", "family_id": "incident-west-47"},
]
assignments = {record["document_id"]: partition_for_family(record["family_id"])
               for record in records}
assert assignments["ticket-47"] == assignments["handoff-82"]
assert partition_for_family("incident-west-47") == assignments["ticket-47"]

Performance and operating cost

Hashing a family ID costs O(k) time for its k encoded bytes and O(1) additional application space, apart from the encoded input. Assigning d records costs O(total family-ID bytes) time and O(d) space if assignments are retained. A hash partition is deterministic, not a guarantee of temporal validity, slice balance or zero exposure to external corpora.

Common Mistakes

  • Splitting revisions of one ticket across training and test.
  • Calling a deterministic hash split temporally valid without checking dates.
  • Ignoring retrieval and synthetic corpora when auditing test exposure.
  • Moving an evaluation record into training after seeing its score.

Read next

ai-data
natural-language-processing
Storage details