Investigate suspicious training influence with comparison runs, clean reference data and a reversible production response.
Poisoning signal triage: isolate influence and recover a clean model
Start with observable symptoms
A rise in false negatives for high-value receipts after a partner batch arrived may indicate poisoned labels, a parser defect, delayed ground truth or a real behavior change. Align the first bad decision window with source ingestion, training run, evaluation snapshot and promotion timeline. Keep the hypothesis open. The incident timeline links those facts; source admission provides a safe hold while investigation continues.
Measure influence without leaking the suspect set
Compare the candidate trained with the suspect batch against a run that excludes only that batch, using the same code, seed policy, clean training base and independent frozen evaluation set. Look at route flips and slice deltas, not just global score. If the clean reference set shares the compromised source, the comparison cannot settle the question. Avoid “test until it passes” on the same held-out examples. Use a second review set or adjudicated outcomes when available. Uncertainty-aware slice gates keep small cohorts honest.
Contain the current serving effect
If a suspect trained model is live, identify every serving digest and cohort exposed to it. Repoint to the last approved unaffected artifact or apply a reviewed conservative route policy while that artifact loads. Preserve decision logs needed for later outcome review under the existing privacy policy. Quarantine the batch and pause automatic retraining triggers; otherwise the same data can re-enter a fresh candidate. Reversible pointers reduce response time when the prior package is still compatible.
Close with a reproducible recovery
Rebuild from an accepted snapshot whose source manifest excludes the suspect batch. Re-run quality, calibration and serving-contract checks, then stage promotion with an explicit incident link. Keep the quarantined evidence and exclusion decision in a restricted ledger; do not erase records so completely that the event cannot be explained. Reopen the source only after its admission controls change. The project includes a deceptive pass on global accuracy and a failure on a targeted cohort.
Implementation
def retrain_incident_gate(clean_slice_error, suspect_slice_error,
clean_reference_independent, suspected_batch_quarantined):
if min(clean_slice_error, suspect_slice_error) < 0:
raise ValueError("negative error count")
if not clean_reference_independent:
return "hold:reference-contaminated"
if not suspected_batch_quarantined:
return "hold:batch-still-admitted"
if suspect_slice_error > clean_slice_error:
return "investigate:batch-influence"
return "review:other-causes"
assert retrain_incident_gate(3, 12, True, True) == "investigate:batch-influence"
assert retrain_incident_gate(3, 12, False, True) == "hold:reference-contaminated"
assert retrain_incident_gate(3, 12, True, False) == "hold:batch-still-admitted"
Performance and operating cost
The toy gate is O(1) time and space, while the actual comparison may require two full training runs, fresh evaluation and reviewer time. Retaining clean reference sets costs storage and access control. Excluding one batch isolates a hypothesis only when all other training inputs and evaluation conditions are held fixed; it does not prove malicious intent.
Common Mistakes
- Declaring an attack from a single anomalous slice.
- Comparing runs that also changed code, hyperparameters or evaluation data.
- Rolling back a model while automatic retraining still reads the suspect batch.
- Using a contaminated validation set as proof of recovery.
Read next
- Training source admission: provenance, trust tiers and quarantine
- Project: investigate suspect receipt labels and rebuild safely
- Model incidents: build a release and evidence timeline before rollback
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Slice quality gates when labels are sparse or delayed
