Analysis timing and variance reduction are design choices; changing them after reading results can turn noise into a launch claim.
Experiment analysis: repeated looks, precision and pre-period covariates
Fix the read schedule
A daily dashboard is useful for harm monitoring, but repeatedly stopping at the first small p-value inflates false-positive risk under a fixed-horizon test. Choose a fixed final horizon or a valid sequential procedure in advance. Keep separate emergency guardrail rules for clear harm. The stopping rule belongs in the experiment manifest.
Use pre-treatment information
A learner’s prior completion rate may reduce variance when included as a covariate measured before assignment. Post-assignment activity is affected by treatment and cannot be used as an ordinary control variable. Freeze the covariate extraction and missing-value rule before the run. Test that the pre-period feature is balanced across assignment arms.
Report uncertainty and size
Show the estimated difference in completion probability with an interval, sample counts and the smallest effect worth shipping. A statistically detectable gain can still be too small to pay for higher latency. Conversely, a wide interval crossing zero does not prove no effect. Present guardrail uncertainty and subgroup sizes without declaring a winner from every exploratory slice.
Run a dry analysis
Give each learner a pre-period count and a seven-day outcome, with one missing pre-period value. Fit the planned adjustment on the allowed pre-period signal only. Compare raw and adjusted estimates, then confirm the assigned population and outcome window are identical. If the adjustment greatly changes the sign, inspect missingness and extreme covariate values before trusting it.
Implementation
def centered_preperiod_values(values_by_learner):
observed = [value for value in values_by_learner.values() if value is not None]
if not observed:
return {}
center = sum(observed) / len(observed)
return {learner_id: value - center for learner_id, value in values_by_learner.items()
if value is not None}Performance and operating cost
Centering P observed values costs O(P) time and O(P) space. An adjusted estimator has additional fitting cost and assumptions; evaluate its precision on a rehearsal before adding it to a production decision packet.
Common Mistakes
- Do not peek each day and apply a fixed-horizon significance rule at every look.
- Do not adjust for a covariate measured after assignment.
- Do not call a wide inconclusive interval evidence of no effect.
