Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Mean-response and prediction intervals: identify whose uncertainty is covered

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An interval for a fitted average route time is narrower than an interval for one future route under the same error model.

Name the interval target

The dispatcher asks for the mean time of routes at seven distance units; a driver asks what one future route might take. The first is a confidence interval for a conditional mean. The second is a prediction interval and includes both uncertainty in the estimated mean and variability of individual routes. Report the target, distance, data range and probability procedure in the same sentence. Coverage concerns repeated use of a method, not a probability that a fixed observed interval contains its fixed parameter.

Inspect the width terms

In a simple-line model, uncertainty about the fitted mean grows away from the sample mean distance. For a new route, add an individual-error term under the square root, so the prediction interval is wider at the same distance. The formula below accepts a critical multiplier chosen for the degrees of freedom and desired coverage; it does not estimate that multiplier. It assumes a straight conditional mean and an appropriate error model. Residual diagnostics decide whether the assumptions deserve use.

Guard the deployment boundary

Do not use a mean interval as a customer-facing arrival window. In sparse long-distance bands, uncertainty can be wider than the pooled result suggests, and extrapolating beyond observed distances is a different risk. Compare realized prediction coverage on held-out routes, including distance bands and depots. If empirical coverage fails, investigate conditional error spread, dependence and shift before widening a number mechanically. Slice coverage checks complement the statistical formula.

Record the decision cost

A prediction interval may be operationally too wide to promise a delivery time. That is a finding, not a reason to relabel it as a confidence interval. Route more uncertain cases to a broad ETA message or collect more representative routes. The project distinguishes a precise average estimate from a safe individual commitment; the population frame decides which future routes the check may represent.

Implementation

python
from math import sqrt

def route_interval_halfwidths(distances, residual_sd, target_distance, critical):
    if len(distances) < 3 or residual_sd < 0 or critical <= 0:
        raise ValueError("invalid line interval inputs")
    center = sum(distances) / len(distances)
    spread = sum((distance - center) ** 2 for distance in distances)
    if spread == 0:
        raise ValueError("no distance variation")
    mean_factor = 1 / len(distances) + (target_distance - center) ** 2 / spread
    return (critical * residual_sd * sqrt(mean_factor),
            critical * residual_sd * sqrt(1 + mean_factor))

mean_width, route_width = route_interval_halfwidths(
    [2, 4, 6, 8, 10], 1.2, 7, 2.78)
assert route_width > mean_width > 0

Performance and operating cost

Computing centered spread takes O(n) time and O(1) extra space; a fitted model can reuse this statistic. A model-based interval is cheap to calculate but can be misleading when residual spread changes with distance or route observations are dependent. Out-of-period coverage checks require enough fresh, labeled routes in each operating slice.

Common Mistakes

  • Presenting a mean-response interval as a range for one new route.
  • Interpreting an interval’s confidence level as a posterior probability about a fixed parameter.
  • Using the line formula well outside the sampled distance range.
  • Reporting nominal coverage without checking held-out route slices.

Read next

Continue the workflow: Quantile crossing and tail coverage: audit ordered forecasts by group.

Continue the workflow: Forecast intervals: check coverage and width at each planning horizon.

ai-data
applied-statistics
Storage details