Freeze a store-day experiment, allocate its success-test budget, and document maturity and safety before a rollout call.
Project: monitor a checkout experiment with declared interim looks
Write the trial contract
A retailer tests a changed checkout prompt by assigning store-days within regions. The target is a predeclared change in successful checkouts per eligible visit, with a refund complaint rate as a safety outcome. Fix the assignment unit, region blocks, eligibility rules, event lag, maximum study length, and planned look dates. An analyst can refresh a dashboard whenever needed for operations, but a refreshed significance test becomes another decision opportunity if it can stop the trial. The design lesson defines the independent unit.
Freeze success thresholds
Allocate the total false-positive budget across the scheduled looks before viewing outcomes. A simple conservative allocation is auditable; a less conservative design needs calibrated boundaries and simulation of its operating properties. Store versioned thresholds alongside each look. Check that the observed count of decision looks matches the plan. The alpha-allocation lesson explains the error accounting, and multiplicity also applies if several primary metrics are promoted after launch.
Audit maturity and guardrails
At each look, reconcile eligible, mature, pending and missing outcomes for both arms. A checkout visit whose refund window has not closed cannot yet contribute to the complaint endpoint as a known noncomplaint. Publish the exposure and assignment counts by store-day, inspect cross-store traffic where relevant, and keep the complaint guardrail separate from the success boundary. The early-stop lesson warns that a selected winning effect can be exaggerated.
Deliver a scoped decision
The packet includes randomization audit, all scheduled and actual looks, effect and safety estimates, boundary decisions, maturity rules, and the method used for uncertainty after stopping. The gate below blocks a rollout claim when an undocumented look occurred or immature outcomes were silently counted. Passing the gate permits a statistical and business review; it does not promise the effect will transfer to stores outside the trial. A monitored rollout should keep the same outcome definitions.
Implementation
def checkout_interim_gate(trial):
if trial["actual_decision_looks"] != trial["logged_decision_looks"]:
return "hold:unlogged-look"
if not trial["alpha_plan_frozen"]:
return "hold:alpha-plan"
if trial["immature_counted_as_known"]:
return "hold:outcome-maturity"
if not trial["safety_reported"]:
return "hold:safety"
return "review:rollout-evidence"
trial = {"actual_decision_looks": 3, "logged_decision_looks": 2,
"alpha_plan_frozen": True, "immature_counted_as_known": False,
"safety_reported": True}
assert checkout_interim_gate(trial) == "hold:unlogged-look"
assert checkout_interim_gate({**trial, "logged_decision_looks": 3}) == "review:rollout-evidence"
Performance and operating cost
The gate is O(1). Joining store-day assignments to visit and refund events is O(n) expected time with indexed IDs. Statistical simulation for a richer sequential design costs more, but the expensive audit is usually discovering undocumented looks and outcome delays.
Common Mistakes
- Running extra success tests while retaining the planned error claim.
- Counting pending refund windows as clean outcomes.
- Reporting a selected early estimate as a guaranteed rollout effect.
- Omitting the safety endpoint because the primary metric crossed a boundary.
Read next
- Interim analyses: allocate false-positive risk before the first look
- Early stopping: separate the success decision from effect estimation
- Experiment design: assign the right unit and guard against interference
- Multiple comparisons and peeking: protect a predeclared decision rule
- Rare proportions: keep interval uncertainty visible at zero and one
