Checkpoint a receipt model, inject an interruption and show whether the resumed run stays within its declared resource and evidence limits.
Project: resume receipt-model training under a compute budget
Define the run envelope
The receipt model uses a frozen data snapshot, fixed split, code digest, optimizer revision and 23 GPU-hour budget. Measure pilot throughput and set a checkpoint interval from expected interruption and upload cost. Store complete training state and a manifest under a new immutable run ID. Resource admission decides whether the request can start; checkpoint rules decide whether it can resume.
Inject failure at awkward points
Interrupt once just before a checkpoint and once while its bytes are being written. Keep an older complete checkpoint available. Restart on a compatible environment and verify step, optimizer state and data identity. Then change the feature contract without changing the checkpoint path; the run must refuse to resume under the old ID. Include a budget breach after repeated interruptions so the job ends with an incomplete artifact rather than silently requesting unlimited capacity.
Verify the recovered run
Compare the resumed run with a clean run on the same snapshot: train/evaluation cohort counts, metric tolerance, final digest and elapsed spend. Exact bytes may differ under nondeterministic kernels; report that separately from metric agreement. Do not promote a partially trained model because it happens to pass one small smoke fixture. Replay evidence identifies the earliest changed stage when results diverge.
Deliver an operator record
Report interruption count, last valid checkpoint, discarded partial checkpoint, recovered step, wasted GPU hours, total cost and final promotion eligibility. A successful drill demonstrates an atomic checkpoint and a bounded recovery decision, not merely that a file existed. Link the final artifact to its complete manifest before any registry promotion.
Implementation
def recovery_decision(checkpoint, run, spent_gpu_hours, limit_gpu_hours):
if spent_gpu_hours > limit_gpu_hours:
return {"state": "stop", "reason": "budget-exhausted"}
required = ("data_digest", "code_digest", "feature_contract")
if not checkpoint.get("complete"):
return {"state": "restart", "reason": "partial-checkpoint"}
if any(checkpoint.get(field) != run.get(field) for field in required):
return {"state": "restart", "reason": "identity-changed"}
return {"state": "resume", "step": checkpoint["step"] + 1}
run = {"data_digest": "data-47", "code_digest": "code-82",
"feature_contract": "receipt-v4"}
saved = {**run, "complete": True, "step": 129}
assert recovery_decision(saved, run, 18, 23)["state"] == "resume"
assert recovery_decision({**saved, "complete": False}, run, 18, 23)[
"state"] == "restart"
Performance and operating cost
The recovery decision is O(f) for f identity fields and O(1) extra space. The actual cost is checkpoint I/O, lost work and repeated accelerator startup. Preserving one last known-good checkpoint increases storage but prevents a partial latest write from forcing a complete restart.
Common Mistakes
- Selecting a partial checkpoint by the largest step number.
- Resuming after a feature-contract change under the same run ID.
- Exceeding the declared budget because interruptions are treated as free.
- Promoting a budget-stopped, incomplete model artifact.
Read next
- Training checkpoints: resume state after interruption
- Training resource gates: bound cost before starting a run
- Project: replay a receipt model from frozen training evidence
- Training manifests: link data, code, configuration and artifact
- Model promotion: require evidence before changing the serving pointer
