Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Training source admission: provenance, trust tiers and quarantine

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A retraining feed needs source identity and a quarantine route before its examples can influence a production model.

Define who can contribute examples

A receipt classifier may train from reviewed customer submissions, partner imports and correction queues. Those sources have different controls and incentives. Assign each record an immutable source ID, collection window, ingestion batch, reviewer or automated label channel, consent and retention status. A row without origin should not enter an approved snapshot merely because its schema is valid. Feature admission checks shape and meaning; source admission checks whether the record is trusted for this training purpose.

Separate quality defects from hostile influence

A sudden burst of one merchant category, duplicated images with contradictory labels or labels that change only near the review threshold deserves investigation. It may be pipeline failure, coordinated abuse or a legitimate seasonal shift. Do not claim an attack from a statistical anomaly alone. Keep rate, duplicate and disagreement signals by source and time, then obtain independent evidence. The label ledger records corrections without silently rewriting the past.

Quarantine by lineage

When a batch is suspect, freeze its immutable identifier, prevent it from entering the next snapshot and record the reason and investigator. Quarantine is reversible only through a reviewed decision. Preserve a safe snapshot pointer and a list of affected training runs, challenger models and evaluation sets. A suspicious record can contaminate both model fitting and the apparent validation result, so test-set lineage also matters. Snapshot replay gives a route back to the exact accepted inputs.

Admit only after review

Review sampled raw evidence within privacy limits, compare against a trusted independent channel and decide whether to reject, relabel or accept the batch. If the source is compromised, pause future ingestion until its controls change. Document the evidence and all affected artifact digests; do not merely delete suspect rows from a mutable table. The triage playbook determines whether a retrain is needed, and the project walks through an ambiguous partner-label surge.

Implementation

python
def admit_training_batch(batch, approved_sources, quarantined_batches):
    if not batch.get("batch_id") or not batch.get("source_id"):
        return "reject:missing-lineage"
    if batch["batch_id"] in quarantined_batches:
        return "hold:quarantine"
    if batch["source_id"] not in approved_sources:
        return "hold:unapproved-source"
    if not batch.get("retention_approved"):
        return "reject:retention"
    return "admit"

batch = {"batch_id": "partner-47", "source_id": "partner-r8",
         "retention_approved": True}
assert admit_training_batch(batch, {"partner-r8"}, set()) == "admit"
assert admit_training_batch(batch, {"partner-r8"}, {"partner-47"}) == "hold:quarantine"
assert admit_training_batch({**batch, "source_id": "unknown"},
                            {"partner-r8"}, set()) == "hold:unapproved-source"

Performance and operating cost

Admission lookup is O(1) expected time and space for indexed source and quarantine sets. Persistent provenance adds storage per record or batch and review latency for held data. Overly broad quarantine slows legitimate retraining; overly narrow quarantine lets a contaminated snapshot through. Bound each hold to a concrete lineage scope and an accountable review.

Common Mistakes

  • Treating a schema-valid row as a trusted training example.
  • Calling every distribution shift a poisoning attack.
  • Quarantining mutable table rows without preserving batch identity.
  • Checking training rows while ignoring validation-set contamination.

Read next

Continue the workflow: Federated aggregation: dropout, contribution caps and privacy limits.

Continue the workflow: Inference-time evasion: threat model and input evidence.

ai-data
mlops
Storage details