A bootstrap interval resamples independent units to estimate how a statistic might vary under repeated sampling from a comparable population.
Bootstrap intervals: estimate uncertainty at the right sampling unit
Choose the resampling unit
A receipt-service study measures median review delay across stores. If several receipts come from one store, individual rows are not independent observations for a store-level question. Resample stores and carry each store’s receipts together, or use a design suited to the original sampling plan. A row bootstrap would understate variation when one store contributes many similar receipts. The unit of analysis] must match the inference unit.
Compute and read the interval
For an individual-receipt median under an independent-sampling assumption, draw N receipts with replacement for each replicate, calculate the median, and take suitable quantiles of the replicate statistics. State the sample size, replicate count, random seed and interval method. A percentile interval is easy to implement but can perform poorly for strongly biased or discrete statistics; use a more suitable method or simulation study when coverage matters.
Do not overclaim
An interval describes sampling uncertainty under its resampling assumptions. It does not include missing records, faulty labels, a changed metric definition or selection bias. A narrow interval around a biased sample is still misleading. Compare the interval against a practical decision threshold and show the estimate itself. If a cohort contains only a handful of independent units, report the raw values and treat the interval as fragile.
Implementation
import numpy as np
rng = np.random.default_rng(47)
delays = np.asarray(review_delay_hours, dtype=float)
assert len(delays) >= 30 and np.isfinite(delays).all()
replicates = np.empty(4000)
for draw in range(len(replicates)):
sample = rng.choice(delays, size=len(delays), replace=True)
replicates[draw] = np.median(sample)
estimate = float(np.median(delays))
lower, upper = np.quantile(replicates, [0.025, 0.975])Performance and operating cost
With B replicates of N rows, drawing and partition-based medians take about O(BN) expected time and O(B + N) working space when draws are processed one at a time.
Common Mistakes
- Do not resample receipts independently when the sampling unit is a store.
- Do not interpret the interval as protection against selection bias.
- Do not publish only interval endpoints without the estimate and sampling assumptions.
Read next
- Metric denominators and cohorts: make a rate reproducible
- Exploratory analysis without peeking: inspect the data and preserve the test
- Missing data policy: distinguish absence from a measured zero
- Group and time validation: split by the failure you expect in production
Continue the workflow: Standard error and cluster bootstrap: resample the independent unit.
Continue the workflow: Precision budgets and effective sample size for weighted analyses.
Continue the workflow: Right censoring and time-to-event analysis for open cases.
Continue the workflow: Derived metrics: carry input bounds and shared error into the result.
Continue the workflow: Posterior intervals and threshold decisions.
Continue the workflow: Branch-cluster bootstrap for a rollout contrast.
