A checkpoint must capture the training state required to continue, not merely a copy of the current model weights.
Training checkpoints: resume state after interruption
Save the state that affects the next update
A receipt classifier trained for many epochs can lose hours when a worker is preempted. A resume checkpoint should identify the model weights, optimizer state, learning-rate schedule, epoch and batch position, random generator state where relevant, and data snapshot. Saving only weights may restart with a different optimizer trajectory. Record code and environment digests to reject a checkpoint created by incompatible logic. Training replay supplies the stable data and split identities.
Publish checkpoints atomically
Write a new checkpoint to a temporary location, validate its manifest and checksum, then make it visible. Keep at least the last known-good checkpoint while a new write is in progress. A partially uploaded file should never be selected simply because its name has the latest step number. Checkpoint formats can carry executable deserialization risk, so load only trusted packages under the training environment policy. Artifact trust applies to training state as well as serving packages.
Resume only compatible runs
On restart, compare data snapshot, code digest, feature contract, optimizer configuration and world size before accepting a checkpoint. Distributed training may need per-worker state or a coordinated save point; one worker’s step count is not proof every worker committed the same update. If compatibility fails, begin a new run with a new manifest rather than quietly splicing histories. A resumed run should record interruption count and wasted work.
Choose frequency from failure cost
Checkpointing every step wastes I/O; checkpointing once a day risks losing most of a day. Estimate save time, expected interruption frequency and recovery cost, then test a forced interruption near a write. The recovery project verifies that an interrupted run resumes from a complete checkpoint and produces a defensible final artifact.
Implementation
def resume_checkpoint(checkpoint, requested):
identity = ("data_digest", "code_digest", "optimizer_revision",
"feature_contract")
mismatches = [field for field in identity
if checkpoint.get(field) != requested.get(field)]
if mismatches:
return {"state": "new-run", "mismatches": mismatches}
if not checkpoint.get("complete") or checkpoint.get("step", -1) < 0:
return {"state": "new-run", "mismatches": ["incomplete"]}
return {"state": "resume", "next_step": checkpoint["step"] + 1}
saved = {"data_digest": "data-47", "code_digest": "code-82",
"optimizer_revision": "adam-r3", "feature_contract": "receipt-v4",
"complete": True, "step": 129}
assert resume_checkpoint(saved, saved) == {"state": "resume", "next_step": 130}
assert resume_checkpoint(saved, {**saved, "data_digest": "data-48"})[
"state"] == "new-run"
Performance and operating cost
Comparing f manifest fields costs O(f) time and O(f) space for mismatch reporting. A real save copies O(B) bytes for B checkpoint state and incurs storage and upload time. Choose the save interval by measured I/O and expected lost work; a fast manifest check cannot prove that the checkpoint bytes are intact.
Common Mistakes
- Saving only weights and losing optimizer or scheduler state.
- Loading a half-written latest checkpoint.
- Resuming with a changed data snapshot under the old run ID.
- Assuming one distributed worker’s checkpoint describes the whole job.
Read next
- Training resource gates: bound cost before starting a run
- Project: resume receipt-model training under a compute budget
- Training replay: freeze the cohort, split and runtime
- Model artifacts: verify digest, origin and loading format
- Pipeline stages: cache by complete input identity
Continue the workflow: Incremental learning releases: checkpoint replay and forgetting gates.
