A credible causal analysis records assumptions and probes them with checks that could contradict the proposed explanation.
Sensitivity and placebo checks: state how the causal claim could fail
List threats explicitly
Unmeasured motivation can confound panel adoption; missing outcomes can differ by assignment; one academy can change curriculum during the test. A single p-value does not answer these threats. Write a threat-to-check table before analysis and keep the primary effect estimate separate from diagnostic results. The causal graph identifies where adjustment helps and where it cannot.
Use negative controls
A panel shown today should not affect a completion outcome recorded last month. If the analysis reports an effect on that pre-treatment outcome, it may be detecting group differences or timestamp errors. Choose a negative-control outcome only when there is a defensible reason treatment cannot affect it. Its failure is informative; its success does not establish the main effect.
Bound missing outcomes
Suppose 6 of 47 treated learners and 2 of 53 controls have unknown seven-day status. Compute best- and worst-case completion rates under plausible missing-status assignments, or use a justified missingness model. Report the range alongside the complete-case estimate. Assigned-unit accounting makes such bounds possible.
Report decision sensitivity
Vary a predeclared overlap cutoff or observation window and show whether the sign and practical conclusion persist. A result that appears only after one favorable slice deserves a narrower claim. If a placebo launch also yields a similar effect, hold the release decision and inspect data and concurrent interventions before asserting causality.
Implementation
def binary_rate_bounds(observed_successes, observed_count, assigned_count):
if not 0 <= observed_successes <= observed_count <= assigned_count or assigned_count == 0:
raise ValueError("invalid assigned-unit counts")
missing = assigned_count - observed_count
return observed_successes / assigned_count, (observed_successes + missing) / assigned_countPerformance and operating cost
Binary missing-outcome bounds cost O(1) after counts are available. Repeating an estimator over S defensible specifications costs roughly S times one fit; preregister the set so compute is not used to search for a preferred answer.
Common Mistakes
- Do not present a passed placebo as proof of no confounding.
- Do not discard missing assigned units from the study population.
- Do not choose sensitivity settings after seeing which produce significance.
Read next
- Observational overlap and weighting: know where the data can compare
- Difference-in-differences: compare changes under a defended trend assumption
- Project: review the effect of a new lesson recommendation panel
- Multiple comparisons and peeking: protect a predeclared decision rule
Continue the workflow: Event-time pre-trends and support for a staggered rollout.
Continue the workflow: Negative-control outcomes and placebo rollout checks.
