Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Interim analyses: allocate false-positive risk before the first look

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A planned series of hypothesis tests needs a declared error budget across all opportunities to stop for success.

List every decision opportunity

A checkout team plans to inspect an experiment after 36, 72, and 120 independent store-days. If it runs an ordinary five-percent test at each look and stops at the first favorable result, its chance of a false success exceeds the one-look target. Freeze the primary outcome, direction, assignment unit, look schedule, and total type-I-error budget before exposure begins. The stopping-rule lesson explains why a dashboard refresh is an analysis when it can change the decision.

Spend a simple conservative budget

One transparent design assigns a small local significance level to each planned look, with their sum no greater than the total budget. The union bound then controls the chance of any false rejection even though the looks reuse data, but it can sacrifice power. This is a deliberately simple allocation, not an implementation of a specialized group-sequential boundary. The code checks planned local thresholds and returns the first declared crossing; it assumes each supplied p-value is valid for its prespecified look under the null.

Distinguish efficacy from futility

Stopping because evidence supports benefit and stopping because further recruitment is not useful are different decisions. A futility rule may be nonbinding for type-I-error control under a stated design, yet it still affects expected sample size and interpretation. A harm monitor also needs its own outcome and operating rule. Do not select the number or timing of looks after watching the effect estimate, then pretend those looks were prespecified. The power lesson links early boundaries to required information.

Report the full path

For every planned and actual look, report information accrued, p-value or test statistic, local boundary, and whether enrollment continued. If the trial stops early, disclose that the point estimate may be exaggerated by selection of the first boundary crossing. A confidence interval copied from a fixed-sample analysis may have the wrong repeated-use properties. The next lesson separates the decision boundary from effect-size estimation; the project records the audit trail.

Implementation

python
def first_planned_crossing(p_values, local_alpha, total_alpha):
    if not p_values or len(p_values) != len(local_alpha):
        raise ValueError("one threshold required for each planned look")
    if not 0 < total_alpha < 1 or any(not 0 <= value <= 1 for value in p_values):
        raise ValueError("p-values and budget must be probabilities")
    if any(threshold < 0 for threshold in local_alpha) or        sum(local_alpha) > total_alpha + 1e-12:
        raise ValueError("planned thresholds exceed total alpha")
    for look, (p_value, threshold) in enumerate(zip(p_values, local_alpha), 1):
        if p_value <= threshold:
            return look
    return None

assert first_planned_crossing([0.021, 0.009, 0.035],
                              [0.005, 0.015, 0.03], 0.05) == 2

Performance and operating cost

Checking k planned looks is O(k) time and O(1) extra space. Specialized boundaries can preserve more power but need calibrated design software and simulation. The cheap act of reusing a fixed-sample threshold at every look can invalidate the stated false-positive rate.

Common Mistakes

  • Inspecting extra unrecorded looks and claiming the original budget still applies.
  • Confusing a conservative local-alpha allocation with an optimized group-sequential boundary.
  • Changing the primary metric after seeing an interim result.
  • Reporting only the final favorable look and hiding earlier decisions.

Read next

ai-data
applied-statistics
Storage details