An allocation mismatch or telemetry asymmetry can invalidate an effect estimate before statistical significance becomes relevant.
Experiment integrity: sample-ratio mismatch and A/A checks
Compare assigned counts
For a 50/50 allocation, assignment counts should fluctuate around equal shares. A large persistent departure can signal a broken hash, a changed eligibility gate or missing logs. Compare observed counts with configured allocation using an appropriate diagnostic and inspect segments. Do not apply the check only to exposed users; variant-dependent rendering can change exposure without changing assignment.
Trace the funnel
Count eligible, assigned, rendered and outcome-bearing learners by variant and day. A treatment-specific fall from assignment to exposure may indicate page failure rather than reduced interest. Join by stable IDs and investigate data loss before analyzing completions. The event contract keeps assignment, exposure and outcome roles distinct.
Run an A/A rehearsal
Before shipping a changed policy, assign two groups that receive identical behavior. Check that event collection, deduplication and metrics behave symmetrically. An A/A run does not prove the product metric will move as expected, but it can expose broken joins, skewed assignment and reporting logic. Keep its analysis procedure fixed before the live comparison.
Use a concrete fixture
Assign 2,350 learners to control and 2,350 to treatment, then remove 80 treatment exposure events while retaining assignments. The assignment ratio remains equal; the exposure funnel no longer does. An analyst should not repair the exposure gap by excluding the affected treatment learners from the primary effect estimate. Record the missing-render cause.
Implementation
def allocation_gap(assignment_rows, expected_treatment_share):
total = len(assignment_rows)
if total == 0:
return None
observed = sum(row["variant"] == "treatment" for row in assignment_rows) / total
return observed - expected_treatment_sharePerformance and operating cost
A count pass is O(N) time and O(1) auxiliary space. A full integrity report adds group-by state for days and segments; every diagnostic needs a known expected allocation and an investigation threshold set before reading results.
Common Mistakes
- Do not declare a clean experiment from a single pooled assignment count.
- Do not run the allocation diagnostic only on exposed users.
- Do not ignore a broken A/A run because the treatment result looks attractive.
