A trial stopped on favorable interim evidence selects unusually strong early estimates, so its effect report needs the stopping history and mature outcomes.
Early stopping: separate the success decision from effect estimation
Recognize selection at the boundary
Suppose a queue experiment stops as soon as an improvement statistic crosses a high early threshold. Among trials that stop, sampling noise often helped push the observed contrast across that threshold. The observed effect can therefore overstate the underlying effect even when the stopping decision controlled its false-positive probability. A valid rejection rule is not automatically an unbiased point estimator or a compatible confidence interval. The interim-look lesson defines the decision rule.
Freeze outcome maturity
At an early look, long-running cases may still be unresolved. Comparing completed cases only can select fast cases differently across arms. Set a maturity window before each look, account for missing outcomes, and record the eligible number per randomized unit. If maturity rules change midtrial, the nominal p-value may be testing a different endpoint from the planned one. The missing-outcome ledger provides the counts and reasons.
Report an adjusted or qualified effect
A fully adjusted estimate or repeated confidence interval depends on the planned sequential design and should be calculated with a validated method. If that method is unavailable, report the raw estimate as descriptive, label the uncertainty limitation, and collect a prespecified follow-up sample for a separate estimate. Do not treat later nonrandom adopters as an unbiased continuation of the trial. The code below audits maturity and assignment counts before any effect calculation; it deliberately does not manufacture a sequential interval.
Keep safety and deployment distinct
A trial may cross an efficacy boundary while a cost or safety guardrail remains uncertain. Record every boundary and the eventual adoption rule before reading interim results. An early stop may leave little data on rare harms and on segments enrolled later. The project requires the exact stop reason, maturity ledger, and effect-report status before making a rollout recommendation. Rare-event intervals illustrate why a small safety sample is weak evidence of absence.
Implementation
def interim_maturity_gate(look):
if look["eligible"] != look["mature"] + look["not_yet_mature"]:
return "hold:cohort-ledger"
if look["randomized_units"] != look["assigned_units"]:
return "hold:assignment-ledger"
if look["not_yet_mature"] and not look["maturity_rule_applied"]:
return "hold:outcome-maturity"
if not look["boundary_version_frozen"]:
return "hold:unfrozen-boundary"
return "review:sequential-result"
look = {"eligible": 127, "mature": 119, "not_yet_mature": 8,
"randomized_units": 44, "assigned_units": 44,
"maturity_rule_applied": False, "boundary_version_frozen": True}
assert interim_maturity_gate(look) == "hold:outcome-maturity"
assert interim_maturity_gate({**look, "maturity_rule_applied": True}) == "review:sequential-result"
Performance and operating cost
The gate is O(1). Reconstructing event maturity is O(n) over cases, and estimating a stopping-adjusted effect requires specialized analysis of the full design. Reporting a fixed-sample interval is computationally cheap but can be misleading after result-dependent stopping.
Common Mistakes
- Calling the first boundary-crossing effect unbiased.
- Treating unresolved outcomes as failures or dropping them without a frozen rule.
- Using a fixed-sample interval as if repeated looks had not occurred.
- Claiming a rare safety event is absent after a small early-stop sample.
Read next
- Interim analyses: allocate false-positive risk before the first look
- Project: monitor a checkout experiment with declared interim looks
- Missing outcomes: count absence before choosing an estimator
- Confidence intervals: interpret coverage and precision honestly
- Rare proportions: keep interval uncertainty visible at zero and one
