A probabilistic forecast needs a declared horizon, quantile order and calibration check; low point error alone cannot justify an uncertainty interval.
Quantile forecast loss and crossing policy
Define the unit and horizon
A queue forecast might estimate completed orders thirty minutes ahead for each warehouse. Record whether the target is an integer count, rate or time-to-clear, and whether it covers a point at minute 77 or the sum over minutes 48 through 77. Those are different labels. Build features only from values available at issue time. A quantile model predicts conditional cutoffs, such as lower, median and upper outcomes, not three independent accuracy scores. The causal lesson fixes the input clock.
Use asymmetric pinball cost
At quantile q, underpredicting an observation by five contributes q times five; overpredicting by five contributes one-minus-q times five. Thus a lower quantile is punished less for being below the observation and more for being above it. The code checks both directions with fresh values rather than relying on a library helper. Aggregate the loss at matched horizons and nonmissing labels. Point forecast metrics can accompany it, but mean absolute error does not assess the promised interval.
Prevent or audit crossing
Separately trained outputs can produce an estimated lower quantile above the upper quantile. Reject crossing predictions in a release audit, or parameterize ordered outputs with positive increments. Sorting after prediction may hide a modeling issue and can change which head receives training gradients. The project uses a median with positive distances to lower and upper bounds. This enforces order but not empirical coverage. The project audits both.
Measure interval coverage and width together
For a proposed lower–upper interval, coverage is the fraction of realized demand values inside it; width measures how informative that interval is. A huge interval can attain high coverage without helping staffing. Evaluate coverage by warehouse, traffic regime and horizon on rolling held-out periods. An interval selected on the same final period used to report coverage leaks tuning into evaluation. Count missing labels and avoid reporting confidence from a tiny number of independent days.
Make the decision cost visible
If staffing undercapacity is more costly than idle staff, an upper quantile may feed the decision rule, but the ratio of costs needs operational validation. Keep the forecast distribution and the staffing policy separate so either can be changed without mislabeling model accuracy. Report labor or service-level outcomes under replay with capacity constraints, rather than assuming a lower pinball loss guarantees a better operational decision. Retain a seasonal baseline for rollback.
Implementation
from math import isclose
def pinball(observed, predicted, quantile):
if not 0 < quantile < 1:
raise ValueError("quantile must be between zero and one")
error = observed - predicted
return max(quantile * error, (quantile - 1) * error)
observed_orders = 47
lower_estimate = 42
upper_estimate = 52
assert lower_estimate <= upper_estimate
assert isclose(pinball(observed_orders, lower_estimate, 0.2), 1.0)
assert isclose(pinball(observed_orders, upper_estimate, 0.8), 1.0)
observed_interval_hit = lower_estimate <= observed_orders <= upper_estimate
assert observed_interval_hitPerformance and operating cost
The scalar pinball calculation is O(1); over N forecast rows and Q quantiles it costs O(NQ) time with streaming O(1) auxiliary space per row. Training cost is dominated by the forecasting network, while storing every horizon and quantile needs O(NHQ) output space for H horizons. Coverage estimates need enough independent issue times, not merely many overlapping windows from one day. The numeric check is illustrative and does not establish calibrated intervals.
Common Mistakes
- Do not score a point target against an interval-sum forecast.
- Do not claim calibration because lower and upper outputs are ordered.
- Do not tune quantiles or interval width on the final reporting period.
