Testing many outcomes, segments or interim snapshots raises false-positive risk unless the comparison family and stopping rule are planned.
Multiple comparisons and peeking: protect a predeclared decision rule
Name the comparison family
A team may inspect review rate, delay, cost and 20 product slices, then highlight the one favorable result. Each test offers another chance to see noise. Choose a primary outcome and planned guardrails before launch. For a family of confirmatory comparisons, use a stated correction or joint decision rule. Exploratory slices are useful for diagnosis but should be labeled as exploratory.
Choose when to stop
Repeatedly checking an ordinary fixed-sample p-value and stopping at the first low value breaks its nominal error interpretation. Predeclare a sample size or use a valid sequential design. An operational safety stop is separate: a severe incident can halt a test, but should not be rebranded as statistical success. Experiment design] fixes the assignment clock.
Retain effect size
A correction can make significance harder to reach; it does not change the observed effect. Report estimated changes and intervals alongside adjusted decisions. A practically harmful guardrail should matter even if it narrowly misses a significance threshold. Effect-size reporting] keeps the interpretation anchored in work impact.
Audit the analysis plan
Store primary metric, family of tests, correction, planned end condition and exclusion policy under a versioned experiment ID. Run a fixture with several identical null segments and ensure the report does not elevate the smallest p-value as if it had been the sole planned test.
Implementation
def bonferroni_decisions(p_values_by_metric, family_alpha=0.05):
if not p_values_by_metric or not 0 < family_alpha < 1:
raise ValueError("invalid comparison family or alpha")
if any(not 0 <= value <= 1 for value in p_values_by_metric.values()):
raise ValueError("invalid p-value")
per_test_limit = family_alpha / len(p_values_by_metric)
return {metric: value <= per_test_limit
for metric, value in p_values_by_metric.items()}Performance and operating cost
A simple correction is O(M) for M planned tests. Its conservatism can increase required sample size; unplanned repeated looks create interpretation cost that extra compute cannot fix.
Common Mistakes
- Do not select the best-looking slice after seeing results and call it confirmatory.
- Do not stop at the first ordinary p-value below a threshold.
- Do not hide effect sizes behind adjusted significance labels.
Read next
- Hypothesis tests: pair the decision rule with an effect size
- Experiment design: assign the right unit and guard against interference
- Project: evaluate a receipt-review workflow without changing the question midstream
- Exploratory analysis without peeking: inspect the data and preserve the test
Continue the workflow: Sensitivity and placebo checks: state how the causal claim could fail.
Continue the workflow: Experiment analysis: repeated looks, precision and pre-period covariates.
