Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit ticket families before publishing a text classifier score

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Find copied ticket families, hold a contaminated evaluation run and report a score whose split boundary can be inspected.

Define the release question

A support classifier is being scored on a new ticket export. The same incident appears as ticket 47, a later corrected reply and handoff 82. The evaluation should estimate performance on unseen incidents, so these records must share one family and one partition. Include an unrelated ticket that repeats the standard footer and a revision that reverses a rollback instruction. Near-duplicate review identifies candidates without equating their labels.

Build the manifest

Capture origin IDs, revision IDs, timestamps, access flags, export batches and source hashes. Generate lexical candidate edges, review each edge with source lineage and retain a reason for accepted or rejected family membership. Freeze a family registry and snapshot the training, validation and test manifests. A reviewer should be able to trace any scored document back to its raw record and its split decision without reading another team’s private tickets.

Detect contamination before scoring

Check whether a family occurs in more than one partition. Search the training corpus, retrieval index and synthetic example store for test-family matches. Quarantine uncertain matches rather than silently deleting them. The code below tests the partition invariant; a production audit also needs source manifests and a review queue for near-duplicate candidates. Family splitting prevents one known leakage path, while external-corpus checks cover the others.

Release a bounded claim

Report per-family and per-document metrics, unseen-incident slice size, contamination findings and any untraceable training source. If one family crosses partitions, hold the score, repair the split and retrain before retesting. Do not present a cleaned test score from the same model if it was tuned on the old test cases. Re-run slice and abstention evaluation after the boundary is repaired.

Implementation

python
def cross_partition_families(records):
    partitions_by_family = {}
    for record in records:
        family_id = record["family_id"]
        partition = record["partition"]
        partitions_by_family.setdefault(family_id, set()).add(partition)
    return {family_id: sorted(partitions)
            for family_id, partitions in partitions_by_family.items()
            if len(partitions) > 1}

manifest = [
    {"family_id": "incident-west-47", "partition": "train"},
    {"family_id": "incident-west-47", "partition": "test"},
    {"family_id": "billing-north-82", "partition": "test"},
]
assert cross_partition_families(manifest) == {
    "incident-west-47": ["test", "train"]}
manifest[1]["partition"] = "train"
assert cross_partition_families(manifest) == {}

Performance and operating cost

Scanning d records takes expected O(d) time and O(f) space for f families. Sorting partitions for conflicting families adds a small per-family cost; the full corpus search for unknown near-duplicates is separate and can dominate runtime. A passing family check says nothing about missing lineage, inaccessible pretraining data or label mistakes.

Common Mistakes

  • Publishing the headline score despite a cross-partition family.
  • Deleting a conflicting record without preserving its audit trail.
  • Using the same held-out cases to tune and then certify the model.
  • Reporting document count while hiding the number of independent incident families.

Read next

ai-data
natural-language-processing
Storage details