A resumable run records weights and optimizer state alongside epoch, class mapping, configuration and input identity.
Checkpoint recovery: save optimizer state and the run boundary
Distinguish deploy from resume
Weights may be enough for inference. They are not enough for a faithful training restart because momentum and adaptive optimizer statistics affect the next update. Save model and optimizer state, completed epoch, configuration, class mapping and input manifest. Save scheduler or precision-scaler state when used. Input snapshots] prevent a restart from silently changing data.
Write a complete artifact
Serialize to a temporary file in the same filesystem, then replace the published checkpoint. A process killed mid-write should leave the previous complete artifact readable. For remote object storage, use an immutable object plus a pointer rather than assuming rename is atomic. Keep best-validation and latest-recovery artifacts distinct.
Check before resuming
Compare architecture, class order, data manifest and configuration. A checkpoint may load while labels have changed order, returning plausible but wrong classifications. Restore random state if exact stochastic continuation matters; hardware and kernel behavior can still prevent bitwise replay. Document any deliberate learning-rate change rather than presenting it as an identical continuation.
Drill a failure
Train a few updates, interrupt during candidate checkpoint writing, then restart. Confirm the prior checkpoint remains readable and a truncated candidate is rejected. Compare the next fixed-batch loss under a controlled setup. Validation] determines which artifact may be promoted.
Implementation
checkpoint = {
"completed_epoch": completed_epoch,
"model_state": receipt_model.state_dict(),
"optimizer_state": optimizer.state_dict(),
"class_names": class_names,
"data_manifest": data_manifest_id,
}
temporary_path = checkpoint_path.with_suffix(".pending")
torch.save(checkpoint, temporary_path)
temporary_path.replace(checkpoint_path)Performance and operating cost
A save transfers O(P + S) bytes for P model parameters and S optimizer state. Frequent checkpoints cost I/O and storage; sparse saves increase lost work after failure.
Common Mistakes
- Do not claim a faithful optimizer resume from weights alone.
- Do not overwrite the only valid checkpoint during an incomplete write.
- Do not load without checking class order and input identity.
Read next
- Training and validation modes: measure the model you will serve
- Inference contracts: preserve preprocessing and measure tail latency
- Reproducible analysis snapshots: pin data, code and cutoff together
- Project: classify receipt image quality with a checked training contract
Continue the workflow: Staged fine-tuning and checkpoint selection.
Continue the workflow: Structured pruning and compute shape.
Continue the workflow: Mixed precision, loss scaling and gradient clipping.
Continue the workflow: Residual blocks and normalization state.
Continue the workflow: Recurrent hidden-state resets and truncated gradients.
Continue the workflow: Distributed checkpoint manifests and coherent resume.
Continue the workflow: Early stopping, best-state restoration and seed variance.
Continue the workflow: Activation checkpointing, recomputation and RNG parity.
Continue the workflow: Affine coupling layers, inverse checks and scale stability.
Continue the workflow: Project: audit a neural inventory dispatch policy before release.
