Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Paired bootstrap uncertainty for prediction-score differences

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A case-level paired bootstrap resamples whole cases to assess how much a score difference between two frozen prediction methods varies with the validation cohort.

Compare on identical cases

Suppose a baseline and a revised model each predict nine-day resolution for the same mature cases. Compute each case’s squared-error difference, baseline minus revised; a positive mean favors the revision under Brier loss. Resample case IDs with replacement and carry both predictions and the outcome together. Resampling separate model tables would break the pairing and exaggerate or distort uncertainty. Fixed-horizon scoring defines the outcome.

Match the sampling unit

If cases are independently sampled, case-level resampling is a useful first approximation. If one customer can open many related cases or a policy rollout occurs by branch, independent case resampling understates uncertainty. Use customer or rollout cluster as the resampling unit and retain every member of a selected cluster. A month-level process shift is not represented by repeatedly drawing cases from only one month; a later cohort tests transfer. Temporal validation addresses that distinct question.

Read the interval as a teaching diagnostic

The code uses a fixed random seed, 600 resamples and empirical percentile endpoints. It returns the observed average improvement and an interval computed from case resampling. With a tiny fixture or a skewed score difference, percentile endpoints can be unreliable. A production report should increase resamples, inspect the distribution, choose a suitable interval method, and explain the validation sampling design. A narrow interval around zero is different from evidence of a useful change.

Do not bootstrap away missing follow-up

The fixture contains only known nine-day outcomes. If a competing-risk Brier score uses censoring weights, a resample should repeat the relevant censoring estimation and scoring procedure at the correct sampling unit. Holding estimated weights fixed can understate uncertainty. If the goal is uncertainty of a complete fitted pipeline, model training and tuning must also be repeated inside the resample; the simple code instead conditions on frozen predictions.

Connect precision to action

A score difference is one part of release review. Report sample size, competing-event share, subgroup support, horizon and the absolute score of each model. An apparent improvement in overall Brier score may hide worse calibration in the queue where decisions are made. Use the interval to express sampling variability, not a guarantee of future deployment performance. The project records both claims separately.

Implementation

python
import random

def paired_brier_improvement(case_scores, repetitions=600, seed=47):
    if not case_scores or repetitions < 100:
        raise ValueError("insufficient cohort or resamples")
    if len({case_id for case_id, _, _, _ in case_scores}) != len(case_scores):
        raise ValueError("duplicate case ID")
    differences = []
    for _, target, baseline_risk, revised_risk in case_scores:
        if target not in (0, 1) or not 0 <= baseline_risk <= 1 or not 0 <= revised_risk <= 1:
            raise ValueError("invalid frozen prediction")
        differences.append((baseline_risk - target) ** 2 -
                           (revised_risk - target) ** 2)
    observed = sum(differences) / len(differences)
    generator = random.Random(seed)
    replicates = sorted(sum(generator.choices(differences, k=len(differences))) /
                        len(differences) for _ in range(repetitions))
    lower = replicates[int(.025 * repetitions)]
    upper = replicates[int(.975 * repetitions)]
    return observed, (lower, upper)

review_cases = [("E501", 1, .55, .72), ("E502", 0, .48, .31),
                ("E503", 0, .39, .27), ("E504", 1, .62, .68),
                ("E505", 0, .44, .29), ("E506", 1, .59, .66)]
improvement, interval = paired_brier_improvement(review_cases)
assert improvement > 0
assert interval[0] <= improvement <= interval[1]

Performance and operating cost

With N cases and B resamples, direct case resampling is O(B × N + B log B) time and O(N + B) space. For clustered data, sampling and rebuilding all members of a chosen cluster adds bookkeeping. The calculation does not cover model-refit or censor-model uncertainty.

Common Mistakes

  • Do not bootstrap the two models independently when predictions are paired by case.
  • Do not call an interval from frozen predictions uncertainty for a retrained pipeline.
  • Do not treat a bootstrap interval as protection against calendar drift.

Read next

ai-data
data-science
Storage details