Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit a route-time regression before promising arrivals

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Fit a route-time association, inspect long-trip errors and separate a mean estimate from an individual ETA.

Define the release question

A delivery team wants one number for arrival estimates. State two targets: average route minutes at a chosen distance and an interval for a future route. Extract routes from a declared period, preserve depot IDs, and record the observed distance range. Keep future weeks out of fitting. A coefficient is not the effect of forcing a courier to travel farther; route selection and traffic differ. The coefficient lesson fixes the estimand.

Inspect the candidate

Fit a single-distance line and compare it with a constant-time baseline. Inspect residuals by distance band, depot and week. One unusually long route has both high influence and a possible scan error. Verify the original record before comparing fits with and without it; retain both results in the review packet. A widening far-route error makes pooled mean absolute error an unsafe sole gate. Diagnostics determine whether the line is adequate.

Set the interval contract

Use a mean-response interval when planning average depot capacity and a prediction interval for an individual driver commitment. Calculate both at a distance supported by the sample. Check prediction coverage in held-out weeks and long-trip slices; a narrow formula interval with poor realized coverage is not ready for customers. If the long-trip data are thin, collect more routes or publish a broader message. Interval targets must not be swapped to meet a display requirement.

Deliver a decision packet

Report rows by depot, distance support, baseline and candidate error, influential-record investigation, interval assumptions and held-out coverage. A failing long-distance gate holds customer ETA use even if the average line is useful for planning. Name the owner of the missing-data and route-quality follow-up. Link this review to tail cost and depot clustering before a later rollout.

Implementation

python
def route_eta_release(report, limits):
    if report["future_week_used_in_fit"]:
        return "hold:time-leakage"
    if report["far_route_mae"] > limits["far_route_mae"]:
        return "hold:far-route-error"
    if report["heldout_coverage"] < limits["minimum_coverage"]:
        return "hold:interval-coverage"
    if not report["influential_route_checked"]:
        return "hold:route-provenance"
    return "release:individual-eta"

limits = {"far_route_mae": 7.5, "minimum_coverage": 0.87}
report = {"future_week_used_in_fit": False, "far_route_mae": 8.3,
          "heldout_coverage": 0.91, "influential_route_checked": True}
assert route_eta_release(report, limits) == "hold:far-route-error"
assert route_eta_release({**report, "far_route_mae": 6.8}, limits) == "release:individual-eta"

Performance and operating cost

The gate is O(1) time and space after report preparation. Fitting and stratified evaluation cost O(n) or more depending on the model and diagnostic refits. Investigating an influential scan, collecting held-out weeks and labeling depot-specific failures are the main operational costs; a fast gate cannot replace that evidence.

Common Mistakes

  • Shipping a customer ETA from an interval for the mean route.
  • Using future weeks in fitting and then calling them a holdout.
  • Ignoring far-route error because pooled error improved.
  • Removing an influential route before checking whether its source record is valid.

Read next

Continue the workflow: Project: review binary claim risk and duplicate-count rates.

Continue the workflow: Project: rank depot handling time with group-aware uncertainty.

Continue the workflow: Project: audit a route-time regression before claiming a policy gain.

ai-data
applied-statistics
Storage details