Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multiple comparisons and peeking: protect a predeclared decision rule

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Testing many outcomes, segments or interim snapshots raises false-positive risk unless the comparison family and stopping rule are planned.

Name the comparison family

A team may inspect review rate, delay, cost and 20 product slices, then highlight the one favorable result. Each test offers another chance to see noise. Choose a primary outcome and planned guardrails before launch. For a family of confirmatory comparisons, use a stated correction or joint decision rule. Exploratory slices are useful for diagnosis but should be labeled as exploratory.

Choose when to stop

Repeatedly checking an ordinary fixed-sample p-value and stopping at the first low value breaks its nominal error interpretation. Predeclare a sample size or use a valid sequential design. An operational safety stop is separate: a severe incident can halt a test, but should not be rebranded as statistical success. Experiment design] fixes the assignment clock.

Retain effect size

A correction can make significance harder to reach; it does not change the observed effect. Report estimated changes and intervals alongside adjusted decisions. A practically harmful guardrail should matter even if it narrowly misses a significance threshold. Effect-size reporting] keeps the interpretation anchored in work impact.

Audit the analysis plan

Store primary metric, family of tests, correction, planned end condition and exclusion policy under a versioned experiment ID. Run a fixture with several identical null segments and ensure the report does not elevate the smallest p-value as if it had been the sole planned test.

Implementation

python
def bonferroni_decisions(p_values_by_metric, family_alpha=0.05):
    if not p_values_by_metric or not 0 < family_alpha < 1:
        raise ValueError("invalid comparison family or alpha")
    if any(not 0 <= value <= 1 for value in p_values_by_metric.values()):
        raise ValueError("invalid p-value")
    per_test_limit = family_alpha / len(p_values_by_metric)
    return {metric: value <= per_test_limit
            for metric, value in p_values_by_metric.items()}

Performance and operating cost

A simple correction is O(M) for M planned tests. Its conservatism can increase required sample size; unplanned repeated looks create interpretation cost that extra compute cannot fix.

Common Mistakes

  • Do not select the best-looking slice after seeing results and call it confirmatory.
  • Do not stop at the first ordinary p-value below a threshold.
  • Do not hide effect sizes behind adjusted significance labels.

Read next

Continue the workflow: Sensitivity and placebo checks: state how the causal claim could fail.

Continue the workflow: Experiment analysis: repeated looks, precision and pre-period covariates.

ai-data
applied-statistics
Storage details