Forecast operations need a versioned prediction ledger and an outcome-availability rule before error alerts can be trusted.
Forecast monitoring: separate data delay, demand shift and model failure
Log each issue
Store forecast origin, target period, model version, source snapshot, point estimate, interval and issued time. When actuals arrive, match by target identifier and outcome version. Without that ledger, a later revision can make old forecasts appear to have used information they never had. Watermarks explain why recent actuals may still be provisional.
Diagnose the alert
A sudden run of zero actuals can mean an ingestion outage rather than a genuine collapse in demand. Check source freshness and completeness before calculating drift or launching retraining. Compare error by horizon, channel and peak day. Directional error can reveal persistent underforecasting while average magnitude looks stable.
Control retraining
A scheduled refit is not automatically an improvement. Rebuild features under the same availability contract, compare with the incumbent and seasonal naive baseline on later origins, then publish atomically with rollback metadata. Keep an untouched evaluation period for release decisions and avoid automatically training on provisional labels.
Prove reversibility
Simulate a bad release that underforecasts one channel. The monitor should identify the affected horizon and source version, and the serving layer should restore the prior model without rewriting historical forecast records. Record whether the correction changes only future forecasts or also revises dashboard history.
Implementation
def eligible_forecast_errors(forecast_rows, final_actuals):
scored = []
for forecast in forecast_rows:
target_key = forecast["target_key"]
if target_key not in final_actuals:
continue
scored.append((target_key, forecast["prediction"] - final_actuals[target_key]))
return scoredPerformance and operating cost
Matching F forecasts to final actuals in a dictionary costs O(F) expected time and O(F) output space. Retraining cost depends on the model, but gate checks should stay cheap enough to run for every candidate release.
Common Mistakes
- Do not alert on error before the outcome is final.
- Do not treat an ingestion outage as demand drift.
- Do not overwrite historical predictions when a model is replaced.
Read next
- Forecast intervals: measure future-outcome coverage by horizon
- Project: forecast receipt volume with a reproducible rolling backtest
- Shadow and canary rollout: compare a candidate without losing a rollback
- Time-series calendar: distinguish missing periods from measured zeros
Continue the workflow: Anomaly operations: distinguish drift, incidents and broken inputs.
