Contain an abnormal partner-label batch, test its influence and restore a receipt-risk model from accepted source lineage.
Project: investigate suspect receipt labels and rebuild safely
Freeze the evidence
A partner feed contributes 8,200 reviewed receipt labels. One afternoon, the share of “safe” labels rises sharply among high-value blurry scans. Preserve batch IDs, source identity, ingestion time, reviewer channel, deduplication fingerprints and the training runs that consumed them. The change may reflect a labeling UI bug or coordinated manipulation; record both hypotheses. Source admission blocks the batch from future snapshots while the team investigates. No raw receipt images belong in general incident dashboards.
Trace training and serving exposure
Build a lineage graph from the batch to the next training snapshot, candidate digest, promotion event and serving cohorts. The candidate shows unchanged global accuracy but more false negatives for high-value blurry scans. Identify decisions made by the candidate and keep their route and model digest for later adjudication. If that model is live, stage the previous unaffected digest, verify its runtime compatibility and compare routes before expanding rollback. The evidence timeline separates when labels arrived from when the model began serving.
Run a controlled comparison
Train one run with all previously accepted inputs and one run with the suspect batch, keeping code, preprocessing and split policy fixed. Evaluate on an independent adjudicated set that excludes the partner feed. Report uncertainty for the affected slice and inspect whether duplicate or conflicting labels cluster by submitter. A larger error count in the suspect run supports batch influence, not proof of hostile intent. The triage procedure defines that limit.
Recover and reopen carefully
Publish a clean snapshot manifest, replacement model digest, slice report, threshold crossing report and rollback observation. Keep the partner source paused until the labeling path is repaired and a sampled independent review passes. Record whether affected customer decisions need review under the product policy. The project is complete only when automatic retraining excludes the held batch, the clean candidate passes normal promotion gates and the source reopening decision has an owner. Retraining triggers resume after those controls are verified.
Implementation
def incident_batch_disposition(batch_id, quarantined, clean_reference,
clean_slice_errors, suspect_slice_errors):
if not batch_id:
raise ValueError("batch id required")
if not quarantined or not clean_reference:
return "hold:containment-or-reference"
if clean_slice_errors < 0 or suspect_slice_errors < 0:
raise ValueError("invalid errors")
if suspect_slice_errors > clean_slice_errors:
return "rebuild:accepted-snapshot"
return "investigate:other-cause"
assert incident_batch_disposition("partner-47", True, True, 4, 13) == "rebuild:accepted-snapshot"
assert incident_batch_disposition("partner-47", False, True, 4, 13) == "hold:containment-or-reference"
assert incident_batch_disposition("partner-47", True, True, 4, 4) == "investigate:other-cause"
Performance and operating cost
The disposition check is O(1), but lineage traversal grows with affected batches, runs, artifacts and cohorts. A controlled rebuild consumes another training run and independent adjudication. Quarantine delays model freshness; keep the known-good artifact serving while this cost is paid. Do not trade away clean-reference independence to shorten the incident.
Common Mistakes
- Deleting the suspicious batch before recording its lineage.
- Assuming unchanged global accuracy means the model is safe.
- Calling a partner malicious before ruling out a label workflow defect.
- Resuming automatic retraining while the held batch remains in an input view.
Read next
- Training source admission: provenance, trust tiers and quarantine
- Poisoning signal triage: isolate influence and recover a clean model
- Training replay: freeze the cohort, split and runtime
- Retraining decisions: require a reason and a challenger comparison
- Model incidents: build a release and evidence timeline before rollback
