Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Regression error slices and costly tails

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Regression evaluation should show average and tail errors by operational slice because a single aggregate metric can hide expensive underestimation.

Start with the sign of the mistake

A shipment predicted to clear in four hours but clearing in nine has a five-hour underestimate; planning staff may miss a handoff. Predicting nine when it clears in four may reserve capacity unnecessarily. Absolute error treats both as five, while the operation may not. Keep signed error alongside MAE and RMSE. RMSE weights larger misses more heavily, but neither metric encodes the specific cost curve. The baseline lesson defines the comparison.

Slice by conditions known at prediction

Compare errors for overnight versus daytime arrivals, high-volume centers and routes with short service windows. Use labels available at prediction time for slicing; a post-outcome category such as actual delay severity can still be useful diagnostically but cannot define a deployable routing rule. Include support counts. A slice with two shipments should not receive the same confidence as one with two thousand.

Find tails without erasing the distribution

Report the maximum and a high quantile of underestimation, plus the fraction exceeding a policy-relevant threshold. The code calculates MAE, RMSE and underestimates of at least three hours per center. It does not estimate uncertainty or optimize a loss function. Keep the raw residual sequence for debugging; a single tail count cannot reveal whether a clock change, missing scan or rare route caused the error.

Separate model error from target error

If actual clearance time is recorded from the wrong scan, both baseline and candidate may appear wrong. Audit target construction before changing features. If the model errors spike only after a new scanning system launches, evaluate label lineage and distribution shift. Feature availability and outcome measurement are independent checks.

Use the slice report to choose a next experiment

A large overnight tail might justify an overnight-specific model, an additional early feature or a capacity buffer. Do not immediately branch the model on every weak slice; this can overfit and increase maintenance. Reserve a new holdout period to verify the proposed change. The shipment project turns the error report into an explicit decision.

Implementation

python
from math import sqrt

def center_error_report(shipments, costly_underestimate_hours):
    groups = {}
    for center, actual_hours, predicted_hours in shipments:
        if actual_hours < 0 or predicted_hours < 0:
            raise ValueError("negative clearance time")
        groups.setdefault(center, []).append(actual_hours - predicted_hours)
    report = {}
    for center, signed_errors in groups.items():
        count = len(signed_errors)
        report[center] = {
            "shipments": count,
            "mae": sum(abs(error) for error in signed_errors) / count,
            "rmse": sqrt(sum(error ** 2 for error in signed_errors) / count),
            "costly_underestimates": sum(
                error >= costly_underestimate_hours for error in signed_errors),
        }
    return report

future_shipments = [("Harbor", 9, 4), ("Harbor", 6, 7),
                    ("Inland", 4, 5), ("Inland", 8, 7)]
report = center_error_report(future_shipments, 3)
assert report["Harbor"]["costly_underestimates"] == 1
assert report["Inland"]["costly_underestimates"] == 0
assert report["Harbor"]["mae"] == 3

Performance and operating cost

For N scored shipments, group accumulation and metric calculation cost O(N) expected time and O(N) space for residuals. Quantiles require sorting within groups, up to O(N log N). Manual inspection of tail cases is often the expensive step.

Common Mistakes

  • Do not rely on aggregate MAE when the operational cost is asymmetric.
  • Do not report a slice rate without its shipment count.
  • Do not treat corrupted target timestamps as a feature-engineering problem.

Read next

Continue the workflow: Paired bootstrap intervals for model gain.

Continue the workflow: Split conformal intervals for clearance forecasts.

ai-data
machine-learning
Storage details