A forecast interval describes plausible future observations under an explicit model and evaluation process.
Forecast intervals: measure future-outcome coverage by horizon
Separate two uncertainties
An interval around an estimated historical mean is not the same as an interval for tomorrow’s actual receipt count. Future observations include process noise and possible shifts. Construct a prediction interval for the quantity that operations needs. Confidence-interval coverage is related, but its target can be a fixed population parameter rather than a future outcome.
Check empirical coverage
On held-out backtest origins, count how often actual demand falls between the stated lower and upper bounds. Report this by forecast horizon and volume slice. A nominal 90 percent interval that covers only 62 percent of peak days is not operationally calibrated for staffing. Also report width; a needlessly huge interval can cover everything while remaining useless.
Preserve causal calibration
If residual quantiles are used to form bounds, estimate them from earlier backtest errors and evaluate on later periods. Reusing the final test errors to tune their own interval width leaks outcome information. Intervals for counts should respect the nonnegative domain, but clipping a lower bound can alter achieved coverage and should be rechecked.
Exercise a fixture
Use four held-out actual counts and fixed interval bounds. Verify that values exactly on endpoints count as covered under the documented convention. Reject inverted bounds and negative actual counts. A later release should compare measured coverage with the earlier baseline under the same horizon and calendar rules.
Implementation
def empirical_interval_coverage(actual_counts, bounds):
if len(actual_counts) != len(bounds) or not actual_counts:
raise ValueError("paired intervals required")
covered = 0
for actual, (lower, upper) in zip(actual_counts, bounds):
if actual < 0 or lower > upper:
raise ValueError("invalid count or interval")
covered += lower <= actual <= upper
return covered / len(actual_counts)Performance and operating cost
Coverage over N paired records is O(N) time and O(1) extra space. Constructing model intervals may cost far more, especially with simulation or repeated fitting. Cache intervals with model, origin, horizon and source-snapshot IDs.
Common Mistakes
- Do not label a mean-estimate interval as a future-outcome interval.
- Do not tune bounds on the final period used to report coverage.
- Do not report coverage without interval width and horizon.
