Changing the names or boundaries of target classes changes the learning problem and its evaluation contract.
Label taxonomy migrations: preserve meaning across retraining
Version meaning, not just spelling
A scanner defect classifier originally labels “damaged” or “clear.” The team now needs “torn label,” “blurred print” and “other damage.” This is not a cosmetic rename: an old positive example may fit several new classes and cannot always be mapped without reinspection. Create a taxonomy revision with definitions, examples, exclusion rules and effective date. Pin the revision to every annotation batch, training snapshot, model output and consumer contract. Label corrections modify judgments within a taxonomy; migration changes the taxonomy itself.
Build a mapping with an unknown state
Define which old labels map deterministically and which require adjudication. Preserve raw original labels and provenance, then write migrated labels as new records linked to the originals. Do not fill ambiguous cases by guessing from the old class or from model predictions. Track mapping coverage and unresolved counts by source and cohort. Source admission can hold batches whose taxonomy is missing or inconsistent.
Plan consumers and score outputs
The new model may emit three damage classes while older clients expect one Boolean. Decide whether a temporary compatibility projection is valid; “any damage” may be safe for manual review but loses the distinction a downstream repair workflow needs. Version the output schema and reject clients that interpret the new class index as the old Boolean. Client contract tests exercise the mixed-version period.
Rebuild evaluation evidence
Historical quality measured against the binary target cannot be directly compared to three-class quality. Maintain an old-label evaluation view for compatibility, create adjudicated examples for new classes and report both with separate denominators. Stage the migration by consumer readiness and outcome maturity. Dual evaluation protects the transition; the project catches an ambiguous legacy class that would otherwise be auto-mapped.
Implementation
def migrate_defect_label(old_label, mapping):
if old_label not in mapping:
return {"state": "hold:unknown-old-label", "new_label": None}
proposed = mapping[old_label]
if proposed is None:
return {"state": "review:ambiguous", "new_label": None}
return {"state": "mapped", "new_label": proposed}
mapping = {"clear": "clear", "damaged": None,
"unreadable": "blurred-print"}
assert migrate_defect_label("clear", mapping)["new_label"] == "clear"
assert migrate_defect_label("damaged", mapping)["state"] == "review:ambiguous"
assert migrate_defect_label("torn", mapping)["state"] == "hold:unknown-old-label"
Performance and operating cost
The mapping lookup is O(1) expected time and space per label. Adjudication cost scales with ambiguous legacy records; storing both revisions increases dataset and lineage storage. Automatic mapping is cheap but incorrect for split classes, so budget human review where old information cannot distinguish the new targets.
Common Mistakes
- Treating a class split as a string rename.
- Overwriting historical labels without preserving their taxonomy revision.
- Guessing a new subclass from an ambiguous old positive label.
- Sending new class indices to an old Boolean consumer.
Read next
- Dual-label evaluation: compare models across a target migration
- Project: split scanner defect labels without breaking old clients
- Label corrections: version outcomes before rebuilding quality metrics
- Training source admission: provenance, trust tiers and quarantine
- Migrate inference clients with compatibility tests and usage evidence
