A regression baseline predicts a simple training-derived value so a new model must demonstrate useful error reduction on unseen outcomes.
Regression baselines and honest holdout metrics
Choose a baseline that matches the loss
A logistics team predicts hours until a shipment clears a sorting center. If the decision is evaluated by absolute error, the training median is a strong constant benchmark; for squared error, the training mean is the corresponding constant benchmark. A previous-period forecast or a route-specific historical median can be harder to beat and may better reflect the operation. Define the baseline before fitting a larger model. The first-model lesson ties prediction to a specific action.
Keep the clock and partition honest
The baseline value must be estimated from training shipments only. The holdout consists of later shipments whose outcomes were unknown at prediction time. Recomputing the median after seeing the holdout leaks the answer into the benchmark and makes comparison meaningless. If the same shipment appears in multiple scans, keep all scans in one partition. Group and time validation explains that boundary.
Compare paired errors on the same rows
Predict every holdout shipment with both baseline and candidate model, then calculate each metric from the same labeled cases. Report the number of shipments, MAE, and a tail metric if late misses create operational harm. A model that lowers average error but produces rare severe underestimates may be worse for staffing. The code compares an explicit candidate prediction vector with a training-median baseline and refuses misaligned inputs.
Interpret what improvement means
A small MAE gain is not automatically a business gain. Include the dispatch decision, capacity limit and cost of late action. A constant baseline can look weak if routes have different service levels; compare against a simple route baseline before deploying a complex predictor. Conversely, a model that cannot beat a training-only median on a future holdout has not justified its added maintenance cost. Error slices locate where the gap appears.
Preserve the untouched final test
Use development folds to select features and settings. Evaluate the final choice once on the future test set; repeated tuning against that set turns it into development data. Store the row IDs, cutoff dates, target version and both prediction vectors. The teaching code calculates scores but does not create the partition; that design must be audited separately.
Implementation
from statistics import median
def mean_absolute_error(actual_hours, predicted_hours):
if not actual_hours or len(actual_hours) != len(predicted_hours):
raise ValueError("aligned nonempty outcomes required")
return sum(abs(actual - predicted)
for actual, predicted in zip(actual_hours, predicted_hours)) / len(actual_hours)
training_clearance_hours = [3, 4, 5, 8, 20]
future_actual_hours = [4, 6, 9, 7]
candidate_hours = [5, 6, 7, 8]
baseline_hour = median(training_clearance_hours)
baseline_hours = [baseline_hour] * len(future_actual_hours)
baseline_mae = mean_absolute_error(future_actual_hours, baseline_hours)
candidate_mae = mean_absolute_error(future_actual_hours, candidate_hours)
assert baseline_hour == 5
assert abs(baseline_mae - 2.0) < 1e-12
assert abs(candidate_mae - 1.0) < 1e-12Performance and operating cost
Computing the training median by sorting N values is O(N log N) time; evaluating M holdout rows is O(M) time and O(M) storage for prediction vectors. Training and serving a candidate model add separate costs that the comparison must justify.
Common Mistakes
- Do not compute the baseline from the final test outcomes.
- Do not compare models on different holdout rows or target definitions.
- Do not equate a small average-error gain with operational value.
Read next
- Regression error slices and costly tails
- Ridge regularization with training-only scaling
- Group and time validation: split by the failure you expect in production
- Project: review shipment-delay and claim-risk models
Continue the workflow: Learning curves and training-size diagnosis.
