Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Checkpoint recovery: save optimizer state and the run boundary

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A resumable run records weights and optimizer state alongside epoch, class mapping, configuration and input identity.

Distinguish deploy from resume

Weights may be enough for inference. They are not enough for a faithful training restart because momentum and adaptive optimizer statistics affect the next update. Save model and optimizer state, completed epoch, configuration, class mapping and input manifest. Save scheduler or precision-scaler state when used. Input snapshots] prevent a restart from silently changing data.

Write a complete artifact

Serialize to a temporary file in the same filesystem, then replace the published checkpoint. A process killed mid-write should leave the previous complete artifact readable. For remote object storage, use an immutable object plus a pointer rather than assuming rename is atomic. Keep best-validation and latest-recovery artifacts distinct.

Check before resuming

Compare architecture, class order, data manifest and configuration. A checkpoint may load while labels have changed order, returning plausible but wrong classifications. Restore random state if exact stochastic continuation matters; hardware and kernel behavior can still prevent bitwise replay. Document any deliberate learning-rate change rather than presenting it as an identical continuation.

Drill a failure

Train a few updates, interrupt during candidate checkpoint writing, then restart. Confirm the prior checkpoint remains readable and a truncated candidate is rejected. Compare the next fixed-batch loss under a controlled setup. Validation] determines which artifact may be promoted.

Implementation

python
checkpoint = {
    "completed_epoch": completed_epoch,
    "model_state": receipt_model.state_dict(),
    "optimizer_state": optimizer.state_dict(),
    "class_names": class_names,
    "data_manifest": data_manifest_id,
}
temporary_path = checkpoint_path.with_suffix(".pending")
torch.save(checkpoint, temporary_path)
temporary_path.replace(checkpoint_path)

Performance and operating cost

A save transfers O(P + S) bytes for P model parameters and S optimizer state. Frequent checkpoints cost I/O and storage; sparse saves increase lost work after failure.

Common Mistakes

  • Do not claim a faithful optimizer resume from weights alone.
  • Do not overwrite the only valid checkpoint during an incomplete write.
  • Do not load without checking class order and input identity.

Read next

Continue the workflow: Staged fine-tuning and checkpoint selection.

Continue the workflow: Structured pruning and compute shape.

Continue the workflow: Mixed precision, loss scaling and gradient clipping.

Continue the workflow: Residual blocks and normalization state.

Continue the workflow: Recurrent hidden-state resets and truncated gradients.

Continue the workflow: Distributed checkpoint manifests and coherent resume.

Continue the workflow: Early stopping, best-state restoration and seed variance.

Continue the workflow: Activation checkpointing, recomputation and RNG parity.

Continue the workflow: Affine coupling layers, inverse checks and scale stability.

Continue the workflow: Project: audit a neural inventory dispatch policy before release.

ai-data
deep-learning
Storage details