Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Regression baselines and honest holdout metrics

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A regression baseline predicts a simple training-derived value so a new model must demonstrate useful error reduction on unseen outcomes.

Choose a baseline that matches the loss

A logistics team predicts hours until a shipment clears a sorting center. If the decision is evaluated by absolute error, the training median is a strong constant benchmark; for squared error, the training mean is the corresponding constant benchmark. A previous-period forecast or a route-specific historical median can be harder to beat and may better reflect the operation. Define the baseline before fitting a larger model. The first-model lesson ties prediction to a specific action.

Keep the clock and partition honest

The baseline value must be estimated from training shipments only. The holdout consists of later shipments whose outcomes were unknown at prediction time. Recomputing the median after seeing the holdout leaks the answer into the benchmark and makes comparison meaningless. If the same shipment appears in multiple scans, keep all scans in one partition. Group and time validation explains that boundary.

Compare paired errors on the same rows

Predict every holdout shipment with both baseline and candidate model, then calculate each metric from the same labeled cases. Report the number of shipments, MAE, and a tail metric if late misses create operational harm. A model that lowers average error but produces rare severe underestimates may be worse for staffing. The code compares an explicit candidate prediction vector with a training-median baseline and refuses misaligned inputs.

Interpret what improvement means

A small MAE gain is not automatically a business gain. Include the dispatch decision, capacity limit and cost of late action. A constant baseline can look weak if routes have different service levels; compare against a simple route baseline before deploying a complex predictor. Conversely, a model that cannot beat a training-only median on a future holdout has not justified its added maintenance cost. Error slices locate where the gap appears.

Preserve the untouched final test

Use development folds to select features and settings. Evaluate the final choice once on the future test set; repeated tuning against that set turns it into development data. Store the row IDs, cutoff dates, target version and both prediction vectors. The teaching code calculates scores but does not create the partition; that design must be audited separately.

Implementation

python
from statistics import median

def mean_absolute_error(actual_hours, predicted_hours):
    if not actual_hours or len(actual_hours) != len(predicted_hours):
        raise ValueError("aligned nonempty outcomes required")
    return sum(abs(actual - predicted)
               for actual, predicted in zip(actual_hours, predicted_hours)) / len(actual_hours)

training_clearance_hours = [3, 4, 5, 8, 20]
future_actual_hours = [4, 6, 9, 7]
candidate_hours = [5, 6, 7, 8]
baseline_hour = median(training_clearance_hours)
baseline_hours = [baseline_hour] * len(future_actual_hours)
baseline_mae = mean_absolute_error(future_actual_hours, baseline_hours)
candidate_mae = mean_absolute_error(future_actual_hours, candidate_hours)
assert baseline_hour == 5
assert abs(baseline_mae - 2.0) < 1e-12
assert abs(candidate_mae - 1.0) < 1e-12

Performance and operating cost

Computing the training median by sorting N values is O(N log N) time; evaluating M holdout rows is O(M) time and O(M) storage for prediction vectors. Training and serving a candidate model add separate costs that the comparison must justify.

Common Mistakes

  • Do not compute the baseline from the final test outcomes.
  • Do not compare models on different holdout rows or target definitions.
  • Do not equate a small average-error gain with operational value.

Read next

Continue the workflow: Learning curves and training-size diagnosis.

ai-data
machine-learning
Storage details