Reproducing a model requires the same eligible population and split boundary, not merely the same training script.
Training replay: freeze the cohort, split and runtime
Freeze the population at a declared cutoff
A receipt model trained on settled outcomes must name the data snapshot and the latest outcome timestamp allowed into training. An updated table can acquire late labels or corrected rows after the first run. Querying it again by the same date filter may return a different cohort. Store immutable row identifiers or a content-addressed snapshot and record exclusions. Outcome maturity prevents recent unresolved transactions from becoming accidental negative examples.
Make the split reproducible and defensible
A random seed only repeats a random split for identical ordered rows. If rows arrive in a different order, assignments may change. Use stable entity IDs or persisted split membership, and apply time and group boundaries that match the expected production failure. Keep a final holdout separate from repeated candidate tuning. For receipt fraud, the same merchant or investigation can appear in many records; splitting related records across train and test can overstate performance. Challenger comparison needs the same frozen evaluation cohort.
Pin the execution environment
Record interpreter, dependency lock, container digest, hardware class and training configuration. A package update can change preprocessing or numerical behavior without any edit to model code. Capture random seeds and nondeterministic operations, then define what “reproduced” means: identical bytes, metrics within tolerance or matching ranking on a fixed cohort. Choose the claim before rerunning. Stage identity should include the environment inputs that affect output.
Verify replay with a manifest diff
Rerun from the frozen snapshot, compare row counts, split membership, feature summaries, evaluation results and artifact digests. If a result differs, identify the earliest stage whose output changed. A perfect metric match does not prove the same model bytes, while a changed digest does not always imply a meaningful quality change. The training project records both and refuses to call a run identical without evidence.
Implementation
from hashlib import sha256
def stable_split(entity_id, evaluation_fraction=0.23):
if not 0 < evaluation_fraction < 1:
raise ValueError("evaluation fraction must be between zero and one")
digest = sha256(entity_id.encode()).digest()
bucket = int.from_bytes(digest[:8], "big") / 2**64
return "evaluation" if bucket < evaluation_fraction else "training"
merchant_ids = ["merchant-47", "merchant-82", "merchant-129"]
original = {merchant: stable_split(merchant) for merchant in merchant_ids}
reordered = {merchant: stable_split(merchant) for merchant in reversed(merchant_ids)}
assert original == reordered
assert stable_split("merchant-47") == original["merchant-47"]
Performance and operating cost
Hashing one entity ID costs O(k) time for its encoded length k and O(1) fixed digest space. Assigning n rows is O(total ID bytes). This stable split prevents order-dependent reassignment, but it does not enforce temporal separation; keep a time cutoff and persist membership when evaluation design requires both.
Common Mistakes
- Assuming a date-filtered mutable table is a frozen snapshot.
- Using a random seed while row ordering changes.
- Letting records from the same merchant cross the evaluation boundary.
- Calling tolerance-level metric agreement byte-for-byte reproduction.
Read next
- Pipeline stages: cache by complete input identity
- Project: replay a receipt model from frozen training evidence
- Training manifests: link data, code, configuration and artifact
- Retraining decisions: require a reason and a challenger comparison
- Prediction-outcome joins: evaluate only mature, matched decisions
Continue the workflow: Training checkpoints: resume state after interruption.
Continue the workflow: Project: govern a distributed parcel-damage model search.
Continue the workflow: Evaluation-set governance: log access and protect the blind holdout.
