A paired bootstrap resamples the evaluation units while keeping baseline and challenger predictions together, showing how unstable the observed error gain may be.
Paired bootstrap intervals for model gain
Define the gain on aligned cases
A route model and a proposed ensemble predict clearance time for the same future shipments. For each shipment, calculate baseline absolute error minus challenger absolute error; a positive mean favors the challenger. The pairing matters because both models faced the same difficult cases. Resampling their errors independently destroys that relationship and invents variability. The baseline guide establishes the common holdout.
Resample independent units
Several shipments from one warehouse may share a staffing outage or scan clock. The code samples warehouse groups with replacement and carries every shipment in a selected group together. This protects within-site dependence better than drawing individual shipments, though six sites still give a fragile interval. If deployment must generalize to unseen regions, regions rather than warehouses may be the right unit. The resampling unit must match the claim.
Treat percentiles as a diagnostic
The code reports the 5th and 95th percentiles of 1,600 bootstrap mean gains as a simple uncertainty band. It is not a guaranteed 90-percent frequentist confidence interval under arbitrary dependence or dataset selection. Small group counts, rare tail events and model tuning on the holdout can make it misleading. Record the number of independent groups and the fraction of draws favoring the challenger alongside the endpoints.
Do not confuse sampling uncertainty with shift
Bootstrap draws rearrange observed sites; they cannot create a future labor strike, a new carrier contract or a changed target definition. Use a later untouched period and inspect costly error tails before release. A positive interval on one historical cohort does not prove improvement everywhere. Where the operational policy has asymmetric costs, bootstrap the policy value itself instead of relying on MAE.
Make the release rule explicit
Decide in advance whether a gain must exceed a practical minimum and whether the lower interval endpoint must clear zero. If the result is uncertain, collect more independent labeled sites or run a limited pilot rather than reporting a false precision. The release checklist combines paired gain, validity and serving cost in one decision.
Implementation
from random import Random
site_results = {
"East": [(7, 5, 6), (9, 5, 8)],
"West": [(4, 7, 5), (8, 6, 8)],
"Harbor": [(11, 7, 10), (6, 4, 6)],
"North": [(5, 8, 5), (9, 6, 8)],
"Inland": [(10, 7, 9), (7, 5, 7)],
"South": [(6, 8, 6), (12, 8, 11)],
}
def mean_gain(site_names):
gains = [abs(actual - baseline) - abs(actual - challenger)
for site in site_names
for actual, baseline, challenger in site_results[site]]
return sum(gains) / len(gains)
sites = list(site_results)
observed_gain = mean_gain(sites)
randomizer = Random(47)
draws = sorted(mean_gain([randomizer.choice(sites) for _ in sites])
for _ in range(1600))
interval = (draws[80], draws[1520])
assert observed_gain > 0
assert interval[0] <= interval[1]
assert len(draws) == 1600Performance and operating cost
With B bootstrap replicates and N shipments, direct resampling and scoring costs O(BN) time; sorting B estimates costs O(B log B), with O(B + N) retained data. For large holdouts, cache per-group error sums and counts to avoid revisiting every shipment on each draw.
Common Mistakes
- Do not resample baseline and challenger errors separately.
- Do not bootstrap individual rows when deployment units are dependent groups.
- Do not describe a percentile band as protection against future distribution shift.
Read next
- Regression baselines and honest holdout metrics
- Feature drift and delayed-label monitoring
- Group and time validation: split by the failure you expect in production
- Ensemble release review project
Continue the workflow: Champion–challenger shadow comparison.
Continue the workflow: Augmented propensity policy value.
