Stopping a run and choosing a checkpoint are separate actions; a patience counter is valid only with a fixed metric, cadence and tie policy.
Early stopping, best-state restoration and seed variance
Freeze the selection metric first
Decide whether validation loss, clipped-edge recall or a constrained composite selects a checkpoint before reading the run. Validation loss is continuous and often easier to rank, but the deployed decision may care about a rare class or review capacity. If several metrics matter, set hard gates and one tie-break metric. Keep threshold fitting and calibration on separate data if they influence a final decision. Confidence selection should not secretly reuse the checkpoint-selection population.
Define patience in evaluations
Patience counts validation checks, not necessarily epochs. An evaluation every 300 parameter updates with patience three allows 900 updates without a qualifying improvement. Define a minimum improvement, comparison direction and tie rule. In the example, a new loss must beat the best by more than 0.005; tiny oscillations do not reset patience. A run that stops on three misses may still have passed through a transient local fluctuation, so do not call patience a guarantee against overfitting. Record each check with completed update count.
Restore the best coherent generation
The last model in memory after early stopping is usually not the chosen model. Save a deep copy or durable checkpoint when a qualifying improvement occurs, and restore that generation for final evaluation. If training might resume from the best checkpoint, retain its optimizer, scheduler, scaler and sampler state too. For pure inference, model weights, buffers and preprocessing are sufficient, but that is a different artifact. Resume-state accounting avoids confusing the two. Verify the restored checkpoint reproduces saved validation logits.
Estimate run-to-run variation
A single seed is one realization of initialization, data shuffling and augmentation. Repeat a small set of defensible candidate configurations over the same declared seeds, then report the spread of rare-defect recall and false accepts. Do not cherry-pick the best seed after opening final test. If the winning configuration changes across seeds, the conclusion is uncertain even if one run has a high score. Split IDs and device slices should remain fixed so variance is not confounded with a changing evaluation population.
Protect the final test
Use the development set to select schedule, patience, objective weights and checkpoint. After that, lock the procedure and run the final physical-receipt holdout once for a release decision. Every repeated look at final test to adjust hyperparameters spends information from that holdout. If the result fails, document the failure and create a new evaluation population for another selection cycle. The project requires a predeclared selection record rather than a retrospective narrative.
Implementation
from dataclasses import dataclass
@dataclass
class SelectionState:
best_loss: float = float("inf")
best_check: int = -1
missed_checks: int = 0
def observe_validation(state: SelectionState, check_index: int,
validation_loss: float, min_improvement: float,
patience: int) -> bool:
if validation_loss < state.best_loss - min_improvement:
state.best_loss = validation_loss
state.best_check = check_index
state.missed_checks = 0
else:
state.missed_checks += 1
return state.missed_checks >= patience
loss_trace = [0.53, 0.47, 0.461, 0.460, 0.4605, 0.4603, 0.461]
selection = SelectionState()
stop_check = None
for check_index, observed_loss in enumerate(loss_trace):
if observe_validation(selection, check_index, observed_loss, 0.005, 3):
stop_check = check_index
break
assert selection.best_check == 2
assert stop_check == 5Performance and operating cost
The patience logic is O(1) per validation check, but validation itself costs a model forward over every selected example and checkpoint storage can be large. Saving every slight improvement may create excessive I/O; a declared minimum improvement controls selection frequency, not correctness. Multiple seeds multiply training cost roughly by the number of runs. That expense is useful when result variance is comparable to the improvement being claimed; report compute budget and uncertainty rather than relying on one lucky run.
Common Mistakes
- Do not evaluate the final in-memory model when the best checkpoint occurred earlier.
- Do not count optimizer updates as patience checks when validation cadence differs.
- Do not choose a seed after examining the final test and report it as an unbiased result.
