A forecast interval should cover future outcomes at its declared rate on comparable future periods without becoming uselessly wide.
Forecast intervals: check coverage and width at each planning horizon
Attach uncertainty to the right target
A staffing forecast for next week and one for six weeks ahead have different information sets. Keep their origin and horizon labels with the point prediction and lower and upper bounds. A confidence interval for an average trend is not a prediction interval for a new weekly case count. An interval around the point forecast should account for future outcome variation as well as estimation uncertainty under the chosen model. The interval distinction carries over from regression.
Measure empirical coverage
On archived rolling-origin forecasts with mature outcomes, count the fraction of realized weekly volumes within the stated bounds separately for each horizon. A nominal 90-percent interval should have roughly that coverage across comparable future periods if it is calibrated, but small samples are noisy. Also report average width and how often the actual count exceeds the upper staffing limit. The code audits supplied records; it does not construct or recalibrate an interval. The backtest lesson ensures the records were genuine forecasts.
Watch sharpness and regime changes
A huge interval can cover nearly everything while offering little staffing value. Compare width and operational undercapacity cost with a frozen baseline. Coverage may look right overall while holiday weeks, new service lines or one region under-cover. Group reports need denominators, and adjacent rolling forecasts can share errors. If the outcome definition or case intake policy changes, the old interval calibration may no longer transport.
Report asymmetric risk
Understaffing from an upper-tail miss can cost more than overstaffing. A central interval may not match that loss; a high predictive quantile can be a better staffing threshold. Show the cost choice and avoid shifting it after seeing the results. The quantile lesson gives an asymmetric score. The project requires horizon coverage, width, peak-period checks and upper-tail misses before approving staffing recommendations.
Implementation
def interval_coverage_by_horizon(forecast_rows):
summary = {}
for horizon, actual, lower, upper in forecast_rows:
if horizon <= 0 or lower > upper:
raise ValueError("valid horizon and bounds required")
covered, count, total_width = summary.get(horizon, (0, 0, 0.0))
summary[horizon] = (covered + (lower <= actual <= upper),
count + 1, total_width + upper - lower)
return {horizon: {"coverage": covered / count,
"mean_width": width / count, "rows": count}
for horizon, (covered, count, width) in summary.items()}
report = interval_coverage_by_horizon([(1, 29, 25, 34), (1, 38, 27, 35),
(3, 42, 30, 48)])
assert report[1] == {"coverage": 0.5, "mean_width": 8.5, "rows": 2}
Performance and operating cost
Auditing n saved forecasts is O(n) time and O(h) space for h distinct horizons. Generating interval forecasts can be much costlier and must be repeated as the model updates. A cheap global coverage percentage can conceal expensive upper-tail misses at the horizon used for staffing.
Common Mistakes
- Calling an interval for the mean a range for next week’s actual count.
- Pooling one-week and six-week forecasts into one coverage figure.
- Making intervals wider until coverage looks good without reporting width.
- Checking only periods whose outcome has already arrived and dropping delayed peaks.
Read next
- Rolling-origin backtests: make every forecast from information available then
- Project: validate a case-volume forecast before it sets staffing
- Mean-response and prediction intervals: identify whose uncertainty is covered
- Conditional quantiles: optimize tail predictions with pinball loss
- Serial correlation: daily rows are not daily independent evidence
