An experiment ends with a reproducible launch decision that includes telemetry integrity, effect size and release risk.
Experiment rollout: ramp gates, rollback and a reviewable decision
Ramp under gates
Start with a small allocation that can reveal failures, then expand only after assignment logs, exposure funnel, latency and severe-error guardrails pass. Every ramp stage has its own allocation and time boundary. Mixing ramp stages without recording them can obscure day-of-week and novelty effects. A policy may be safe for a small group yet fail under full candidate volume.
Protect continuity
Persist learner assignment across ramps unless the design explicitly permits reassignment. If the new policy is rolled back, stop new treatment exposure and record the rollback time while retaining the event stream for outcome maturation. Canary and rollback practices help separate monitoring from the final causal estimate.
Write the decision packet
Include protocol version, eligibility rule, assignment salt, event definitions, sample-ratio diagnostics, exposure counts, primary effect with interval, guardrails, subgroup counts, missing-label status and deviations. State a ship, revise or stop decision against thresholds chosen before analysis. A reviewer should reconstruct the denominator and one learner’s variant from packet data.
Rehearse a failure
During a 15 percent ramp, suppose treatment feed errors rise and exposure falls while assignments remain balanced. Roll back based on the predeclared error gate, even if early completions look favorable. Preserve the unexposed assigned learners in the analysis. If the issue is fixed, start a separately versioned run rather than silently continuing the old result.
Implementation
def ramp_gate(error_rate, exposure_gap, max_error_rate, max_exposure_gap):
return error_rate <= max_error_rate and abs(exposure_gap) <= max_exposure_gapPerformance and operating cost
A gate evaluation is O(1) after counters are computed. Computing counters over N requests is O(N) time with bounded aggregate state; keep raw event lineage because a broken counter can otherwise make the rollback decision unauditable.
Common Mistakes
- Do not continue a changed treatment under the old experiment version.
- Do not ship on one favorable metric when a predeclared guardrail fails.
- Do not erase rollback-era outcomes from the assigned population.
Read next
- Experiment integrity: sample-ratio mismatch and A/A checks
- Project: run a guarded lesson-feed experiment
- Shadow and canary rollout: compare a candidate without losing a rollback
- Offline recommendation evaluation: replay catalog and learner state in time
Continue the workflow: Project: release review for maintenance-search ranking.
Continue the workflow: Project: release an evidence-backed lesson-reminder policy.
