Variation among independently refitted models signals sensitivity to the training sample, while repeated outcome variability at similar inputs signals a different source of uncertainty.
Ensemble disagreement and outcome noise
Do not call every spread the same uncertainty
Four shipment models trained on different valid samples may issue clearance forecasts of six, seven, nine and ten hours for one intake record. Their spread shows model sensitivity under that training protocol. It does not measure all future outcome variation, because every model can agree and still miss a sudden outage. The code reports between-model spread and the observed squared error of their mean on a small mature cohort.
Construct refits without leakage
Train each model on a declared resample or fold within the permitted development window. Keep feature transformations inside each refit and do not let calibration or final-test outcomes train any member. Correlated models may agree for the wrong reason, so low disagreement is not a guarantee of reliability. The bagging lesson shows how training samples vary; the spread here has a different diagnostic purpose.
Inspect residual noise separately
A center may have clearance time that varies because of random dock congestion not captured at intake. Even if refitted models agree on a seven-hour mean, individual outcomes can range widely. Compare issued forecasts with mature outcomes across similar operating conditions. That residual variation mixes unobserved process noise, measurement error and model misspecification; it is not a clean mathematical decomposition from a small table. Error slices identify where to investigate.
Avoid false precision from a tiny ensemble
A standard deviation across four highly correlated models is not a calibrated prediction interval. Increasing members may stabilize the estimate of disagreement, but it cannot make a shifted future cohort exchangeable. If dispatch requires an interval with an explicit coverage target, use a separately calibrated procedure and audit it later. The conformal lesson states the relevant boundary and coverage scope.
Turn uncertainty into a decision test
Compare whether high disagreement predicts future errors or only flags rare but easy routes. Record support, interval width, review volume and cost of unnecessary escalation. A review rule is useful only when it improves policy value under capacity constraints. Selective prediction closes that loop.
Implementation
from math import sqrt
# Each tuple holds forecasts from independently refitted development models.
issued_forecasts = [(6, 7, 9, 10), (5, 5, 6, 6), (8, 8, 9, 9)]
mature_outcomes = [12, 5, 8]
def mean_and_spread(member_forecasts):
center = sum(member_forecasts) / len(member_forecasts)
spread = sqrt(sum((forecast - center) ** 2 for forecast in member_forecasts)
/ len(member_forecasts))
return center, spread
diagnostics = [mean_and_spread(members) for members in issued_forecasts]
mean_squared_error = sum((actual - center) ** 2
for actual, (center, _) in zip(mature_outcomes, diagnostics))
mean_squared_error /= len(mature_outcomes)
assert diagnostics[0][1] > diagnostics[1][1]
assert mean_squared_error > 0
assert diagnostics[0][0] == 8Performance and operating cost
For N cases and B model members, this diagnostic costs O(NB) time and O(N) summary storage when forecast vectors are streamed. Training B independent models and serving all B predictions can multiply compute and latency. Store only the detail needed for a reviewed decision.
Common Mistakes
- Do not label ensemble spread a calibrated interval.
- Do not infer irreducible noise from residuals without checking misspecification and label error.
- Do not treat agreement among correlated models as certainty.
