Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Paired bootstrap intervals for model gain

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A paired bootstrap resamples the evaluation units while keeping baseline and challenger predictions together, showing how unstable the observed error gain may be.

Define the gain on aligned cases

A route model and a proposed ensemble predict clearance time for the same future shipments. For each shipment, calculate baseline absolute error minus challenger absolute error; a positive mean favors the challenger. The pairing matters because both models faced the same difficult cases. Resampling their errors independently destroys that relationship and invents variability. The baseline guide establishes the common holdout.

Resample independent units

Several shipments from one warehouse may share a staffing outage or scan clock. The code samples warehouse groups with replacement and carries every shipment in a selected group together. This protects within-site dependence better than drawing individual shipments, though six sites still give a fragile interval. If deployment must generalize to unseen regions, regions rather than warehouses may be the right unit. The resampling unit must match the claim.

Treat percentiles as a diagnostic

The code reports the 5th and 95th percentiles of 1,600 bootstrap mean gains as a simple uncertainty band. It is not a guaranteed 90-percent frequentist confidence interval under arbitrary dependence or dataset selection. Small group counts, rare tail events and model tuning on the holdout can make it misleading. Record the number of independent groups and the fraction of draws favoring the challenger alongside the endpoints.

Do not confuse sampling uncertainty with shift

Bootstrap draws rearrange observed sites; they cannot create a future labor strike, a new carrier contract or a changed target definition. Use a later untouched period and inspect costly error tails before release. A positive interval on one historical cohort does not prove improvement everywhere. Where the operational policy has asymmetric costs, bootstrap the policy value itself instead of relying on MAE.

Make the release rule explicit

Decide in advance whether a gain must exceed a practical minimum and whether the lower interval endpoint must clear zero. If the result is uncertain, collect more independent labeled sites or run a limited pilot rather than reporting a false precision. The release checklist combines paired gain, validity and serving cost in one decision.

Implementation

python
from random import Random

site_results = {
    "East": [(7, 5, 6), (9, 5, 8)],
    "West": [(4, 7, 5), (8, 6, 8)],
    "Harbor": [(11, 7, 10), (6, 4, 6)],
    "North": [(5, 8, 5), (9, 6, 8)],
    "Inland": [(10, 7, 9), (7, 5, 7)],
    "South": [(6, 8, 6), (12, 8, 11)],
}

def mean_gain(site_names):
    gains = [abs(actual - baseline) - abs(actual - challenger)
             for site in site_names
             for actual, baseline, challenger in site_results[site]]
    return sum(gains) / len(gains)

sites = list(site_results)
observed_gain = mean_gain(sites)
randomizer = Random(47)
draws = sorted(mean_gain([randomizer.choice(sites) for _ in sites])
               for _ in range(1600))
interval = (draws[80], draws[1520])
assert observed_gain > 0
assert interval[0] <= interval[1]
assert len(draws) == 1600

Performance and operating cost

With B bootstrap replicates and N shipments, direct resampling and scoring costs O(BN) time; sorting B estimates costs O(B log B), with O(B + N) retained data. For large holdouts, cache per-group error sums and counts to avoid revisiting every shipment on each draw.

Common Mistakes

  • Do not resample baseline and challenger errors separately.
  • Do not bootstrap individual rows when deployment units are dependent groups.
  • Do not describe a percentile band as protection against future distribution shift.

Read next

Continue the workflow: Champion–challenger shadow comparison.

Continue the workflow: Augmented propensity policy value.

ai-data
machine-learning
Storage details