Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: resume receipt-model training under a compute budget

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Checkpoint a receipt model, inject an interruption and show whether the resumed run stays within its declared resource and evidence limits.

Define the run envelope

The receipt model uses a frozen data snapshot, fixed split, code digest, optimizer revision and 23 GPU-hour budget. Measure pilot throughput and set a checkpoint interval from expected interruption and upload cost. Store complete training state and a manifest under a new immutable run ID. Resource admission decides whether the request can start; checkpoint rules decide whether it can resume.

Inject failure at awkward points

Interrupt once just before a checkpoint and once while its bytes are being written. Keep an older complete checkpoint available. Restart on a compatible environment and verify step, optimizer state and data identity. Then change the feature contract without changing the checkpoint path; the run must refuse to resume under the old ID. Include a budget breach after repeated interruptions so the job ends with an incomplete artifact rather than silently requesting unlimited capacity.

Verify the recovered run

Compare the resumed run with a clean run on the same snapshot: train/evaluation cohort counts, metric tolerance, final digest and elapsed spend. Exact bytes may differ under nondeterministic kernels; report that separately from metric agreement. Do not promote a partially trained model because it happens to pass one small smoke fixture. Replay evidence identifies the earliest changed stage when results diverge.

Deliver an operator record

Report interruption count, last valid checkpoint, discarded partial checkpoint, recovered step, wasted GPU hours, total cost and final promotion eligibility. A successful drill demonstrates an atomic checkpoint and a bounded recovery decision, not merely that a file existed. Link the final artifact to its complete manifest before any registry promotion.

Implementation

python
def recovery_decision(checkpoint, run, spent_gpu_hours, limit_gpu_hours):
    if spent_gpu_hours > limit_gpu_hours:
        return {"state": "stop", "reason": "budget-exhausted"}
    required = ("data_digest", "code_digest", "feature_contract")
    if not checkpoint.get("complete"):
        return {"state": "restart", "reason": "partial-checkpoint"}
    if any(checkpoint.get(field) != run.get(field) for field in required):
        return {"state": "restart", "reason": "identity-changed"}
    return {"state": "resume", "step": checkpoint["step"] + 1}

run = {"data_digest": "data-47", "code_digest": "code-82",
       "feature_contract": "receipt-v4"}
saved = {**run, "complete": True, "step": 129}
assert recovery_decision(saved, run, 18, 23)["state"] == "resume"
assert recovery_decision({**saved, "complete": False}, run, 18, 23)[
    "state"] == "restart"

Performance and operating cost

The recovery decision is O(f) for f identity fields and O(1) extra space. The actual cost is checkpoint I/O, lost work and repeated accelerator startup. Preserving one last known-good checkpoint increases storage but prevents a partial latest write from forcing a complete restart.

Common Mistakes

  • Selecting a partial checkpoint by the largest step number.
  • Resuming after a feature-contract change under the same run ID.
  • Exceeding the declared budget because interruptions are treated as free.
  • Promoting a budget-stopped, incomplete model artifact.

Read next

ai-data
mlops
Storage details