Review small depots, a new-site prediction and a mid-pilot equipment change before publishing performance ranks.
Project: rank depot handling time with group-aware uncertainty
Freeze the group and time ledger
A logistics team ranks depots by dispatch handling time. Define each eligible shift, depot, route mix, equipment state and measurement clock, then record which depots existed during the training period. Some depots contribute many shifts while newer ones contribute very few. Separate the target for comparing currently operating depots from the target for forecasting a future site. Partial pooling can stabilize small-group means, but its network center must refer to comparable sites.
Find the broken pooling assumption
One depot replaced its scanner midway through the pilot. Its older shifts are much slower and belong to a different operating regime, yet the first model pools them as one stationary depot. Inspect timestamps, scanner versions, staffing and route mix; either model that change explicitly or limit comparison to a common regime. Keep raw and pooled estimates together. A smooth-looking ranking can conceal a real equipment problem. Residual patterns are the warning, not a nuisance to average away.
Evaluate two uncertainty targets
Estimate within- and between-depot variation, then report uncertainty for each existing depot and for a hypothetical newly opened depot. Predict on held-out later shifts by depot size; if pooled estimates fail on the changed-scanner depot, the exchangeability claim is weak. Compare ranking changes when that site is modeled separately. Do not call a new-site center a guaranteed operational baseline, because it has no local observations. The variance lesson distinguishes known and unseen group risk.
Publish a decision with limits attached
Deliver group counts, raw and pooled means, variance components, held-out error, equipment-change ledger and uncertainty intervals. The gate below blocks ranks if the group map or change record is incomplete. A high-rank depot with few shifts should be flagged for more observation rather than declared certainly best. A changed operating regime may require a second model and an explicit scope statement. Shared-prior limits and correct resampling units inform follow-up.
Implementation
def depot_ranking_gate(audit, limits):
if audit["missing_depot_ids"] or audit["unlogged_equipment_changes"]:
return "hold:group-contract"
if audit["depots_with_too_few_shifts"] > limits["maximum_sparse_depots"]:
return "hold:small-group-support"
if audit["late_holdout_error"] > limits["maximum_holdout_error"]:
return "hold:future-fit"
return "review:scoped-ranking"
limits = {"maximum_sparse_depots": 2, "maximum_holdout_error": 4.5}
audit = {"missing_depot_ids": 0, "unlogged_equipment_changes": 1,
"depots_with_too_few_shifts": 1, "late_holdout_error": 3.2}
assert depot_ranking_gate(audit, limits) == "hold:group-contract"
assert depot_ranking_gate({**audit, "unlogged_equipment_changes": 0},
limits) == "review:scoped-ranking"
Performance and operating cost
The gate is O(1), while the group ledger and held-out evaluation require at least O(n) work across shifts. Fitting hierarchical components is more expensive than raw averages, but the cheap rank can make a small depot look exceptional from a handful of noisy shifts.
Common Mistakes
- Pooling old and new scanner regimes as one stable depot.
- Publishing a rank without shift counts behind each estimate.
- Using the same interval for observed and unseen depots.
- Treating one low held-out error as proof that every depot is comparable.
Read next
- Partial pooling: stabilize small-group estimates without hiding variation
- Multilevel variance: separate within-group noise from new-group risk
- Standard error and cluster bootstrap: resample the independent unit
- Regression diagnostics: residual pattern, scale and influential routes
- Shared priors, shrinkage and the limits of fixed pooling
