Hyperparameter search selects a model setting using development data; a final test estimates performance only after that selection and preprocessing are frozen.
Hyperparameter search and an untouched final test
Separate fitting, selection and evaluation
A shipment-clearance model has candidate ridge penalties of 0, 20 and 80. For each development fold, fit scaling and coefficients on the fold training rows, then score that fold’s held-out rows. Average those validation errors to choose a penalty. The future test period is not a fourth fold in the search; it is opened once after the penalty and feature set are fixed. The validation lesson explains why the fold calendar matters.
Count the search as part of the model
Trying dozens of features, loss functions and penalties and retaining the lowest validation error can overfit the development sample even if each individual fit is clean. Compare multiple candidates under the same folds and keep a record of all attempted settings. A nested outer validation loop can estimate the performance of the entire selection procedure when data support it. For a single final release, an untouched future holdout serves a similar practical purpose.
Inspect fold spread and failure modes
A mean MAE of 2.8 hours may conceal one route-month fold at 6.4 hours. Keep each fold score and support count, and investigate whether a deployment boundary or weather event explains the discrepancy. A setting that wins by 0.02 hours with much greater compute cost may not be worth adopting. Error slices should accompany the selected score.
Freeze all learned state
The selected setting is not just a penalty. Save feature order, missing-value policy, scaler parameters, training end date, target definition and model artifact. If preprocessing is fitted on the whole development dataset before cross-validation, validation rows influence it. The pipeline lesson keeps every learned transform inside the fold.
Use the final test once
The code chooses from already computed inner-fold scores and breaks a tie toward the smaller penalty. It does not fit models or estimate uncertainty. Once the setting is selected, fit it on all permitted development rows, then score a later test period. If the result disappoints and the model changes, acquire a new final period or label the old one as development data. The applied review records that boundary.
Implementation
def choose_penalty(inner_fold_mae):
if not inner_fold_mae:
raise ValueError("no candidates")
ranked = []
for penalty, fold_errors in inner_fold_mae.items():
if penalty < 0 or not fold_errors or any(error < 0 for error in fold_errors):
raise ValueError("invalid candidate scores")
ranked.append((sum(fold_errors) / len(fold_errors), penalty))
mean_error, selected_penalty = min(ranked)
return selected_penalty, mean_error
# These errors came from the same three time-aware development folds.
development_scores = {0: [3.3, 3.0, 4.2],
20: [2.9, 2.8, 3.6],
80: [3.1, 3.0, 3.5]}
penalty, validation_mae = choose_penalty(development_scores)
assert penalty == 20
assert abs(validation_mae - 3.1) < 1e-12Performance and operating cost
For H settings and F recorded fold scores, selection is O(H × F) time and O(H) ranking space. The actual cost is H × F model fits, multiplied again by outer folds if nested evaluation is used.
Common Mistakes
- Do not include final test scores in the hyperparameter table.
- Do not preprocess across development folds before fitting.
- Do not hide the number of attempted settings when interpreting a tiny validation win.
Read next
- Ridge regularization with training-only scaling
- Group and time validation: split by the failure you expect in production
- Permutation importance on held-out data
- Project: select a model and audit warehouse segments
Continue the workflow: Gradient boosting residuals and early stopping.
Continue the workflow: Pairwise ranking loss, ties and useful comparisons.
