Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Hyperparameter search and an untouched final test

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Hyperparameter search selects a model setting using development data; a final test estimates performance only after that selection and preprocessing are frozen.

Separate fitting, selection and evaluation

A shipment-clearance model has candidate ridge penalties of 0, 20 and 80. For each development fold, fit scaling and coefficients on the fold training rows, then score that fold’s held-out rows. Average those validation errors to choose a penalty. The future test period is not a fourth fold in the search; it is opened once after the penalty and feature set are fixed. The validation lesson explains why the fold calendar matters.

Count the search as part of the model

Trying dozens of features, loss functions and penalties and retaining the lowest validation error can overfit the development sample even if each individual fit is clean. Compare multiple candidates under the same folds and keep a record of all attempted settings. A nested outer validation loop can estimate the performance of the entire selection procedure when data support it. For a single final release, an untouched future holdout serves a similar practical purpose.

Inspect fold spread and failure modes

A mean MAE of 2.8 hours may conceal one route-month fold at 6.4 hours. Keep each fold score and support count, and investigate whether a deployment boundary or weather event explains the discrepancy. A setting that wins by 0.02 hours with much greater compute cost may not be worth adopting. Error slices should accompany the selected score.

Freeze all learned state

The selected setting is not just a penalty. Save feature order, missing-value policy, scaler parameters, training end date, target definition and model artifact. If preprocessing is fitted on the whole development dataset before cross-validation, validation rows influence it. The pipeline lesson keeps every learned transform inside the fold.

Use the final test once

The code chooses from already computed inner-fold scores and breaks a tie toward the smaller penalty. It does not fit models or estimate uncertainty. Once the setting is selected, fit it on all permitted development rows, then score a later test period. If the result disappoints and the model changes, acquire a new final period or label the old one as development data. The applied review records that boundary.

Implementation

python
def choose_penalty(inner_fold_mae):
    if not inner_fold_mae:
        raise ValueError("no candidates")
    ranked = []
    for penalty, fold_errors in inner_fold_mae.items():
        if penalty < 0 or not fold_errors or any(error < 0 for error in fold_errors):
            raise ValueError("invalid candidate scores")
        ranked.append((sum(fold_errors) / len(fold_errors), penalty))
    mean_error, selected_penalty = min(ranked)
    return selected_penalty, mean_error

# These errors came from the same three time-aware development folds.
development_scores = {0: [3.3, 3.0, 4.2],
                      20: [2.9, 2.8, 3.6],
                      80: [3.1, 3.0, 3.5]}
penalty, validation_mae = choose_penalty(development_scores)
assert penalty == 20
assert abs(validation_mae - 3.1) < 1e-12

Performance and operating cost

For H settings and F recorded fold scores, selection is O(H × F) time and O(H) ranking space. The actual cost is H × F model fits, multiplied again by outer folds if nested evaluation is used.

Common Mistakes

  • Do not include final test scores in the hyperparameter table.
  • Do not preprocess across development folds before fitting.
  • Do not hide the number of attempted settings when interpreting a tiny validation win.

Read next

Continue the workflow: Gradient boosting residuals and early stopping.

Continue the workflow: Pairwise ranking loss, ties and useful comparisons.

ai-data
machine-learning
Storage details