Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: select a receipt training schedule without test leakage

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Compare constant and warmup-decay training under one receipt split, then restore a chosen checkpoint and audit stability across seeds before final release.

Lock the experiment manifest

Group receipt crops by physical receipt and reserve training, development, calibration and final-test populations. Fix model shape, optimizer family, effective batch, accepted-example budget and class weights. Define one constant-rate candidate and one five-update warmup followed by a 47-update cosine decay. Record per-parameter decay groups and whether the rate is applied before each parameter step. The schedule contract makes a resumed run comparable with an uninterrupted one.

Run one seed as a smoke test

Log the first, warmup-end, resumed and final planned learning rates, along with loss by class and accepted-example counts. Save a coherent checkpoint at update 23, resume it and compare the next fixed-batch parameter update with the uninterrupted reference. A mismatch can come from optimizer moments, scheduler index, sampler position or mixed-precision scaler. Do not proceed to expensive repeats until this parity check passes. The checkpoint lesson lists the state needed for replay.

Apply a predeclared selection rule

Evaluate on the development population at fixed update intervals and save the best qualifying checkpoint under one primary metric, with a declared minimum improvement and patience. Restore that checkpoint rather than the last one. If clipped-edge recall has a hard floor, reject any candidate below it before ranking validation loss. The code below checks a decision record over illustrative measured results; it does not train a model or claim these numbers were observed. The selection lesson sets the tie policy.

Repeat and inspect spread

Run both candidates over the same set of independent seeds, preserving split IDs and evaluation cadence. Report per-seed best checkpoint update, clipped-edge recall, false accepts and calibration error; show the spread, not only the mean. If a candidate wins one seed but loses most others, do not call it the stable schedule. A rare class with few receipts needs event counts alongside percentages. Review any data-source or device slice whose result moves in the opposite direction from the aggregate.

Open final test once

After choosing a schedule and checkpoint policy on development data, fit any probability temperature or review threshold on the calibration group. Freeze model artifact, preprocessing and decision policy. Run one final held-out physical-receipt evaluation and publish all release gates, including failures. Retain the prior model if the new plan misses a class-specific gate. Deliver the manifest, update-rate trace, resume-parity check, per-seed table and final decision, not merely the best validation scalar.

Implementation

python
from dataclasses import dataclass

@dataclass(frozen=True)
class RunResult:
    seed: int
    clipped_recall: float
    false_accepts: int
    best_validation_loss: float
    restored_best: bool

def eligible_results(runs: list[RunResult], minimum_recall: float,
                     maximum_false_accepts: int) -> list[RunResult]:
    return [run for run in runs if run.restored_best and
            run.clipped_recall >= minimum_recall and
            run.false_accepts <= maximum_false_accepts]

schedule_runs = [
    RunResult(47, 0.94, 2, 0.31, True),
    RunResult(71, 0.92, 3, 0.29, True),
    RunResult(83, 0.91, 2, 0.33, False),
]
accepted = eligible_results(schedule_runs, 0.92, 3)
assert [run.seed for run in accepted] == [47, 71]
assert len(accepted) < len(schedule_runs)

Performance and operating cost

A fair schedule comparison costs at least one full training run per candidate and seed. Validation cadence adds repeated forward passes; coherent checkpoints add disk I/O and optimizer-state storage. The gate code is O(R) for R run records, but preparing reliable records is the expensive part. Keep the same accepted-example and update budget when comparing schedules, and include time spent on resume and calibration checks in the project estimate.

Common Mistakes

  • Do not use the final test to tune warmup length or patience.
  • Do not rank a run whose best state was never restored.
  • Do not report only the best seed when schedule choice is unstable across repeats.

Read next

Continue the workflow: Training memory profiling and allocator boundaries.

ai-data
deep-learning
Storage details