Forecast quality changes with how far ahead a prediction is made, so a pooled score can conceal operational failures.
Forecast metrics by horizon: keep errors, zeros and denominators explicit
Score each horizon
A one-day forecast and a seven-day forecast support different staffing decisions. Calculate absolute error at each horizon and report origin count and target count. A mean across all scored cells can overweight horizons with more available observations. Define a fixed evaluation matrix or present separate horizon metrics. Rolling origins provide the forecast records.
Handle zero demand
Percentage error divides by actual demand and becomes undefined at zero. Dropping zero days can make a service with intermittent receipts look much better than it is. MAE stays defined, while a scaled error needs a training-only scale that is nonzero. Show the number of zero-demand targets and how each metric treats them.
Make bias visible
Mean absolute error hides the direction of error. A model that underforecasts every peak day may have acceptable MAE yet repeatedly understaff the queue. Report signed error, peak-day slices and interval coverage if intervals are supplied. Interval evaluation addresses the uncertainty of future counts, not just point accuracy.
Recompute an example
For actual counts 47, 52 and 0, with forecasts 43, 58 and 6, absolute errors are 4, 6 and 6; MAE is 16 divided by 3. Signed forecast-minus-actual errors are -4, 6 and 6. A percentage metric cannot consume the final actual value without an explicit zero policy.
Implementation
def forecast_error_summary(actual_counts, predicted_counts):
if len(actual_counts) != len(predicted_counts) or not actual_counts:
raise ValueError("paired nonempty forecasts required")
errors = [forecast - actual for actual, forecast in zip(actual_counts, predicted_counts)]
return {"mae": sum(abs(error) for error in errors) / len(errors),
"bias": sum(errors) / len(errors),
"zero_actuals": sum(actual == 0 for actual in actual_counts)}Performance and operating cost
Calculating metrics for N paired forecasts costs O(N) time and O(N) memory here; streaming sums reduce memory to O(1). Grouping by H horizons needs O(H) aggregates. Preserve the origin and target identifiers so paired scores remain auditable.
Common Mistakes
- Do not silently drop zero-demand targets from a percentage metric.
- Do not combine horizons with unequal record counts without a weighting rule.
- Do not report only unsigned error for an asymmetric staffing cost.
