Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Standard error and cluster bootstrap: resample the independent unit

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Uncertainty depends on the unit that was sampled; repeated receipts from one customer should move together in a bootstrap resample.

Find the independent unit

If customers generate multiple receipts, the rows are correlated. Resampling individual receipts pretends every row arrived independently and can make an interval too narrow. Resample customers, carrying all their eligible receipts into each draw. If assignment or sampling happened at store level, the store may be the correct cluster instead. The estimand] determines how clusters contribute to the final rate.

Implement a paired statistic

For each bootstrap draw, sample clusters with replacement and recompute the statistic from all selected records. A customer drawn twice contributes twice. Keep a fixed seed and record the number of draws. A percentile interval describes the resampling distribution under this design; it does not repair a biased sampling frame or missing outcomes.

Watch small-cluster limits

A dataset with thousands of receipts but only six stores has little independent store-level information. More bootstrap draws do not create more stores. Report cluster count and size distribution. For very few clusters, use a design-specific method or mark the interval unstable. Coverage interpretation] depends on the repeated-sampling procedure, not on a single computed interval.

Validate with a fixture

Create four customers with two receipts each, and make one customer’s outcomes unusually poor. Compare row and customer resampling. The customer-based distribution should reflect that customer entering zero, one or multiple times. Test a customer with no valid outcome and state whether it stays in the denominator.

Implementation

python
import random

def customer_bootstrap_rate(rows_by_customer, draws=1500, seed=47):
    customer_ids = list(rows_by_customer)
    if len(customer_ids) < 2 or any(not values for values in rows_by_customer.values()):
        raise ValueError("need two customers with observed outcomes")
    generator = random.Random(seed)
    estimates = []
    for _ in range(draws):
        sampled = [generator.choice(customer_ids) for _ in customer_ids]
        outcomes = [outcome for customer_id in sampled
                    for outcome in rows_by_customer[customer_id]]
        estimates.append(sum(outcomes) / len(outcomes))
    return sorted(estimates)

Performance and operating cost

With B draws and N total rows per draw, a direct implementation costs O(BN) time plus O(B) stored estimates. Aggregate per-cluster successes and counts first to reduce repeated row traversal.

Common Mistakes

  • Do not resample rows when customers were the independent units.
  • Do not confuse more bootstrap draws with more independent data.
  • Do not ignore clusters with missing outcomes silently.

Read next

ai-data
applied-statistics
Storage details