Compare constant and warmup-decay training under one receipt split, then restore a chosen checkpoint and audit stability across seeds before final release.
Project: select a receipt training schedule without test leakage
Lock the experiment manifest
Group receipt crops by physical receipt and reserve training, development, calibration and final-test populations. Fix model shape, optimizer family, effective batch, accepted-example budget and class weights. Define one constant-rate candidate and one five-update warmup followed by a 47-update cosine decay. Record per-parameter decay groups and whether the rate is applied before each parameter step. The schedule contract makes a resumed run comparable with an uninterrupted one.
Run one seed as a smoke test
Log the first, warmup-end, resumed and final planned learning rates, along with loss by class and accepted-example counts. Save a coherent checkpoint at update 23, resume it and compare the next fixed-batch parameter update with the uninterrupted reference. A mismatch can come from optimizer moments, scheduler index, sampler position or mixed-precision scaler. Do not proceed to expensive repeats until this parity check passes. The checkpoint lesson lists the state needed for replay.
Apply a predeclared selection rule
Evaluate on the development population at fixed update intervals and save the best qualifying checkpoint under one primary metric, with a declared minimum improvement and patience. Restore that checkpoint rather than the last one. If clipped-edge recall has a hard floor, reject any candidate below it before ranking validation loss. The code below checks a decision record over illustrative measured results; it does not train a model or claim these numbers were observed. The selection lesson sets the tie policy.
Repeat and inspect spread
Run both candidates over the same set of independent seeds, preserving split IDs and evaluation cadence. Report per-seed best checkpoint update, clipped-edge recall, false accepts and calibration error; show the spread, not only the mean. If a candidate wins one seed but loses most others, do not call it the stable schedule. A rare class with few receipts needs event counts alongside percentages. Review any data-source or device slice whose result moves in the opposite direction from the aggregate.
Open final test once
After choosing a schedule and checkpoint policy on development data, fit any probability temperature or review threshold on the calibration group. Freeze model artifact, preprocessing and decision policy. Run one final held-out physical-receipt evaluation and publish all release gates, including failures. Retain the prior model if the new plan misses a class-specific gate. Deliver the manifest, update-rate trace, resume-parity check, per-seed table and final decision, not merely the best validation scalar.
Implementation
from dataclasses import dataclass
@dataclass(frozen=True)
class RunResult:
seed: int
clipped_recall: float
false_accepts: int
best_validation_loss: float
restored_best: bool
def eligible_results(runs: list[RunResult], minimum_recall: float,
maximum_false_accepts: int) -> list[RunResult]:
return [run for run in runs if run.restored_best and
run.clipped_recall >= minimum_recall and
run.false_accepts <= maximum_false_accepts]
schedule_runs = [
RunResult(47, 0.94, 2, 0.31, True),
RunResult(71, 0.92, 3, 0.29, True),
RunResult(83, 0.91, 2, 0.33, False),
]
accepted = eligible_results(schedule_runs, 0.92, 3)
assert [run.seed for run in accepted] == [47, 71]
assert len(accepted) < len(schedule_runs)Performance and operating cost
A fair schedule comparison costs at least one full training run per candidate and seed. Validation cadence adds repeated forward passes; coherent checkpoints add disk I/O and optimizer-state storage. The gate code is O(R) for R run records, but preparing reliable records is the expensive part. Keep the same accepted-example and update budget when comparing schedules, and include time spent on resume and calibration checks in the project estimate.
Common Mistakes
- Do not use the final test to tune warmup length or patience.
- Do not rank a run whose best state was never restored.
- Do not report only the best seed when schedule choice is unstable across repeats.
Read next
- AdamW decay groups and update-indexed schedules
- Early stopping, best-state restoration and seed variance
- Checkpoint recovery: save optimizer state and the run boundary
- Project: audit receipt confidence and review handoff
- Project: classify receipt image quality with a checked training contract
Continue the workflow: Training memory profiling and allocator boundaries.
