Regression evaluation should show average and tail errors by operational slice because a single aggregate metric can hide expensive underestimation.
Regression error slices and costly tails
Start with the sign of the mistake
A shipment predicted to clear in four hours but clearing in nine has a five-hour underestimate; planning staff may miss a handoff. Predicting nine when it clears in four may reserve capacity unnecessarily. Absolute error treats both as five, while the operation may not. Keep signed error alongside MAE and RMSE. RMSE weights larger misses more heavily, but neither metric encodes the specific cost curve. The baseline lesson defines the comparison.
Slice by conditions known at prediction
Compare errors for overnight versus daytime arrivals, high-volume centers and routes with short service windows. Use labels available at prediction time for slicing; a post-outcome category such as actual delay severity can still be useful diagnostically but cannot define a deployable routing rule. Include support counts. A slice with two shipments should not receive the same confidence as one with two thousand.
Find tails without erasing the distribution
Report the maximum and a high quantile of underestimation, plus the fraction exceeding a policy-relevant threshold. The code calculates MAE, RMSE and underestimates of at least three hours per center. It does not estimate uncertainty or optimize a loss function. Keep the raw residual sequence for debugging; a single tail count cannot reveal whether a clock change, missing scan or rare route caused the error.
Separate model error from target error
If actual clearance time is recorded from the wrong scan, both baseline and candidate may appear wrong. Audit target construction before changing features. If the model errors spike only after a new scanning system launches, evaluate label lineage and distribution shift. Feature availability and outcome measurement are independent checks.
Use the slice report to choose a next experiment
A large overnight tail might justify an overnight-specific model, an additional early feature or a capacity buffer. Do not immediately branch the model on every weak slice; this can overfit and increase maintenance. Reserve a new holdout period to verify the proposed change. The shipment project turns the error report into an explicit decision.
Implementation
from math import sqrt
def center_error_report(shipments, costly_underestimate_hours):
groups = {}
for center, actual_hours, predicted_hours in shipments:
if actual_hours < 0 or predicted_hours < 0:
raise ValueError("negative clearance time")
groups.setdefault(center, []).append(actual_hours - predicted_hours)
report = {}
for center, signed_errors in groups.items():
count = len(signed_errors)
report[center] = {
"shipments": count,
"mae": sum(abs(error) for error in signed_errors) / count,
"rmse": sqrt(sum(error ** 2 for error in signed_errors) / count),
"costly_underestimates": sum(
error >= costly_underestimate_hours for error in signed_errors),
}
return report
future_shipments = [("Harbor", 9, 4), ("Harbor", 6, 7),
("Inland", 4, 5), ("Inland", 8, 7)]
report = center_error_report(future_shipments, 3)
assert report["Harbor"]["costly_underestimates"] == 1
assert report["Inland"]["costly_underestimates"] == 0
assert report["Harbor"]["mae"] == 3Performance and operating cost
For N scored shipments, group accumulation and metric calculation cost O(N) expected time and O(N) space for residuals. Quantiles require sorting within groups, up to O(N log N). Manual inspection of tail cases is often the expensive step.
Common Mistakes
- Do not rely on aggregate MAE when the operational cost is asymmetric.
- Do not report a slice rate without its shipment count.
- Do not treat corrupted target timestamps as a feature-engineering problem.
Read next
- Regression baselines and honest holdout metrics
- Ridge regularization with training-only scaling
- Prediction-time feature availability: reject future information before training
- Project: review shipment-delay and claim-risk models
Continue the workflow: Paired bootstrap intervals for model gain.
Continue the workflow: Split conformal intervals for clearance forecasts.
