Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: split scanner defect labels without breaking old clients

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Migrate a binary defect label into specific classes, review ambiguous history and stage a dual-output model.

Freeze old and new definitions

A scanner fleet currently marks shipping labels damaged or clear. Operations now distinguishes torn labels, blurred print and other damage because each routes to a different repair action. Publish exact class definitions and a new taxonomy revision. Preserve original labels and annotation source IDs. The old damaged class cannot be split from its name alone, so create an adjudication queue rather than an automatic remap. Migration rules keep unknown records explicit.

Build a clean transition snapshot

Take a pinned training snapshot, deterministically map clear and any independently verified unreadable cases, then relabel a sampled set of old damaged cases from retained source evidence. Exclude or hold records whose evidence is unavailable. Split by shipment and time, fitting preprocessing only inside training folds. Keep the new evaluation set independent of records used to settle migration choices. Snapshot replay lets an editor reconstruct both views.

Test both client generations

Train a four-class candidate and define a reviewed projection to old binary routes. Run native class quality, rare torn-label slices, confidence failures and the old route contract on frozen examples. A new label code that an old client interprets as clear must fail contract tests. Stage new and old clients separately, logging model digest, taxonomy revision and projection revision. Dual evaluation prevents a better four-class score from hiding a worse binary route.

Finish with a reversible cutover

Promote only after both quality and compatibility gates pass. Keep the prior model and binary projection until the consumer inventory shows all clients upgraded. Record remaining ambiguous history and exclude it from claims about new-class quality. If a new client fails, restore the prior complete route while preserving labels and lineage; do not rewrite historical examples to fit the rollback. Retirement checks determine when old output support can actually end.

Implementation

python
def taxonomy_cutover_gate(class_gate, legacy_gate,
                          unresolved_records, allowed_unresolved):
    if unresolved_records < 0 or allowed_unresolved < 0:
        raise ValueError("invalid unresolved count")
    if not class_gate:
        return "hold:new-classes"
    if not legacy_gate:
        return "hold:legacy-route"
    if unresolved_records > allowed_unresolved:
        return "hold:adjudication"
    return "stage:dual-output"

assert taxonomy_cutover_gate(True, True, 39, 47) == "stage:dual-output"
assert taxonomy_cutover_gate(True, False, 39, 47) == "hold:legacy-route"
assert taxonomy_cutover_gate(True, True, 82, 47) == "hold:adjudication"

Performance and operating cost

The gate is O(1) time and space after evaluation. Relabeling ambiguous history and running two client contracts cost time, but prevent an apparently successful new classifier from breaking old routing. Maintaining dual output temporarily adds code and monitoring; retire it only when actual consumer usage and contract evidence support removal.

Common Mistakes

  • Auto-splitting the old damaged class without image evidence.
  • Combining old and new taxonomies under one accuracy figure.
  • Deploying new class indices to older clients.
  • Deleting the compatibility route before all consumers migrate.

Read next

ai-data
mlops
Storage details