Design and replay an online comparison of two lesson-feed policies, including assignment, failure handling and a written launch decision.
Project: run a guarded lesson-feed experiment
Create the protocol
Declare active learners as the population, learner ID as the assignment unit and qualifying seven-day lesson completion as the primary outcome. Set a fixed evaluation horizon, minimum useful gain and latency and empty-slate guardrails. Record an immutable salt and eligibility rule. The experiment contract should let another engineer assign the same learner twice and get the same variant.
Build event fixtures
Generate assignment, render, completion and error events for both policies, including repeated delivery, one learner switching devices and a treatment-only render failure. Add an outcome arriving after the provisional report. Deduplicate events and produce daily counts for eligible, assigned, exposed and completed learners. Confirm the assigned-population denominator retains the render failure.
Analyze and decide
Run an A/A integrity check, then the treatment comparison. Report sample-ratio diagnostics, effect and interval, guardrail values, pending outcomes and daily ticket-worthy failures. Test a small ramp gate and a rollback event. Do not tune the metric or fixed-horizon stopping rule after viewing treatment data. Keep exploratory device slices labeled as such.
Deliver evidence
Submit protocol, assignment function, event schema, replay fixture, analysis code, output packet and decision memo. Include one line-level trace from raw completion event to aggregate numerator and one learner who was assigned but never exposed. The decision memo must say why the launch gate passed or failed, not merely quote a p-value.
Implementation
from hashlib import sha256
def stable_learner_arm(learner_id, salt):
digest = sha256(f"{salt}:{learner_id}".encode()).digest()
return "B" if int.from_bytes(digest[:8], "big") % 2 else "A"
assert stable_learner_arm("learner-47", "lesson-feed-v2") == stable_learner_arm("learner-47", "lesson-feed-v2")Performance and operating cost
Assignment costs O(L) time for an identifier of length L and constant retained state. End-to-end replay costs O(E) expected time for E events plus dedup storage; keep the event fixture small but preserve the same counting and correction rules used at scale.
Common Mistakes
- Do not change learner assignment between devices.
- Do not exclude failed treatment renders from the primary estimate.
- Do not treat a provisional late-outcome report as final.
