Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Early stopping, best-state restoration and seed variance

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Stopping a run and choosing a checkpoint are separate actions; a patience counter is valid only with a fixed metric, cadence and tie policy.

Freeze the selection metric first

Decide whether validation loss, clipped-edge recall or a constrained composite selects a checkpoint before reading the run. Validation loss is continuous and often easier to rank, but the deployed decision may care about a rare class or review capacity. If several metrics matter, set hard gates and one tie-break metric. Keep threshold fitting and calibration on separate data if they influence a final decision. Confidence selection should not secretly reuse the checkpoint-selection population.

Define patience in evaluations

Patience counts validation checks, not necessarily epochs. An evaluation every 300 parameter updates with patience three allows 900 updates without a qualifying improvement. Define a minimum improvement, comparison direction and tie rule. In the example, a new loss must beat the best by more than 0.005; tiny oscillations do not reset patience. A run that stops on three misses may still have passed through a transient local fluctuation, so do not call patience a guarantee against overfitting. Record each check with completed update count.

Restore the best coherent generation

The last model in memory after early stopping is usually not the chosen model. Save a deep copy or durable checkpoint when a qualifying improvement occurs, and restore that generation for final evaluation. If training might resume from the best checkpoint, retain its optimizer, scheduler, scaler and sampler state too. For pure inference, model weights, buffers and preprocessing are sufficient, but that is a different artifact. Resume-state accounting avoids confusing the two. Verify the restored checkpoint reproduces saved validation logits.

Estimate run-to-run variation

A single seed is one realization of initialization, data shuffling and augmentation. Repeat a small set of defensible candidate configurations over the same declared seeds, then report the spread of rare-defect recall and false accepts. Do not cherry-pick the best seed after opening final test. If the winning configuration changes across seeds, the conclusion is uncertain even if one run has a high score. Split IDs and device slices should remain fixed so variance is not confounded with a changing evaluation population.

Protect the final test

Use the development set to select schedule, patience, objective weights and checkpoint. After that, lock the procedure and run the final physical-receipt holdout once for a release decision. Every repeated look at final test to adjust hyperparameters spends information from that holdout. If the result fails, document the failure and create a new evaluation population for another selection cycle. The project requires a predeclared selection record rather than a retrospective narrative.

Implementation

python
from dataclasses import dataclass

@dataclass
class SelectionState:
    best_loss: float = float("inf")
    best_check: int = -1
    missed_checks: int = 0

def observe_validation(state: SelectionState, check_index: int,
                       validation_loss: float, min_improvement: float,
                       patience: int) -> bool:
    if validation_loss < state.best_loss - min_improvement:
        state.best_loss = validation_loss
        state.best_check = check_index
        state.missed_checks = 0
    else:
        state.missed_checks += 1
    return state.missed_checks >= patience

loss_trace = [0.53, 0.47, 0.461, 0.460, 0.4605, 0.4603, 0.461]
selection = SelectionState()
stop_check = None
for check_index, observed_loss in enumerate(loss_trace):
    if observe_validation(selection, check_index, observed_loss, 0.005, 3):
        stop_check = check_index
        break
assert selection.best_check == 2
assert stop_check == 5

Performance and operating cost

The patience logic is O(1) per validation check, but validation itself costs a model forward over every selected example and checkpoint storage can be large. Saving every slight improvement may create excessive I/O; a declared minimum improvement controls selection frequency, not correctness. Multiple seeds multiply training cost roughly by the number of runs. That expense is useful when result variance is comparable to the improvement being claimed; report compute budget and uncertainty rather than relying on one lucky run.

Common Mistakes

  • Do not evaluate the final in-memory model when the best checkpoint occurred earlier.
  • Do not count optimizer updates as patience checks when validation cadence differs.
  • Do not choose a seed after examining the final test and report it as an unbiased result.

Read next

ai-data
deep-learning
Storage details