A planned series of hypothesis tests needs a declared error budget across all opportunities to stop for success.
Interim analyses: allocate false-positive risk before the first look
List every decision opportunity
A checkout team plans to inspect an experiment after 36, 72, and 120 independent store-days. If it runs an ordinary five-percent test at each look and stops at the first favorable result, its chance of a false success exceeds the one-look target. Freeze the primary outcome, direction, assignment unit, look schedule, and total type-I-error budget before exposure begins. The stopping-rule lesson explains why a dashboard refresh is an analysis when it can change the decision.
Spend a simple conservative budget
One transparent design assigns a small local significance level to each planned look, with their sum no greater than the total budget. The union bound then controls the chance of any false rejection even though the looks reuse data, but it can sacrifice power. This is a deliberately simple allocation, not an implementation of a specialized group-sequential boundary. The code checks planned local thresholds and returns the first declared crossing; it assumes each supplied p-value is valid for its prespecified look under the null.
Distinguish efficacy from futility
Stopping because evidence supports benefit and stopping because further recruitment is not useful are different decisions. A futility rule may be nonbinding for type-I-error control under a stated design, yet it still affects expected sample size and interpretation. A harm monitor also needs its own outcome and operating rule. Do not select the number or timing of looks after watching the effect estimate, then pretend those looks were prespecified. The power lesson links early boundaries to required information.
Report the full path
For every planned and actual look, report information accrued, p-value or test statistic, local boundary, and whether enrollment continued. If the trial stops early, disclose that the point estimate may be exaggerated by selection of the first boundary crossing. A confidence interval copied from a fixed-sample analysis may have the wrong repeated-use properties. The next lesson separates the decision boundary from effect-size estimation; the project records the audit trail.
Implementation
def first_planned_crossing(p_values, local_alpha, total_alpha):
if not p_values or len(p_values) != len(local_alpha):
raise ValueError("one threshold required for each planned look")
if not 0 < total_alpha < 1 or any(not 0 <= value <= 1 for value in p_values):
raise ValueError("p-values and budget must be probabilities")
if any(threshold < 0 for threshold in local_alpha) or sum(local_alpha) > total_alpha + 1e-12:
raise ValueError("planned thresholds exceed total alpha")
for look, (p_value, threshold) in enumerate(zip(p_values, local_alpha), 1):
if p_value <= threshold:
return look
return None
assert first_planned_crossing([0.021, 0.009, 0.035],
[0.005, 0.015, 0.03], 0.05) == 2
Performance and operating cost
Checking k planned looks is O(k) time and O(1) extra space. Specialized boundaries can preserve more power but need calibrated design software and simulation. The cheap act of reusing a fixed-sample threshold at every look can invalidate the stated false-positive rate.
Common Mistakes
- Inspecting extra unrecorded looks and claiming the original budget still applies.
- Confusing a conservative local-alpha allocation with an optimized group-sequential boundary.
- Changing the primary metric after seeing an interim result.
- Reporting only the final favorable look and hiding earlier decisions.
Read next
- Early stopping: separate the success decision from effect estimation
- Project: monitor a checkout experiment with declared interim looks
- Multiple comparisons and peeking: protect a predeclared decision rule
- Minimum detectable effect: plan a study around a useful change
- Experiment design: assign the right unit and guard against interference
