A falsification check repeats a prespecified comparison where the intervention should have no effect, looking for evidence of a broken design or measurement process.
Negative-control outcomes and placebo rollout checks
Choose a control outcome for a reason
A support triage screen may affect breach rate but should not change an archived hardware inventory count recorded through a separate system. If both jump at rollout, an organization-wide reporting change or another concurrent policy may be present. An outcome is a valid negative control only when there is a defensible reason the triage screen cannot affect it, while the suspected bias would affect both. Do not choose an arbitrary unrelated metric merely because it returns a reassuring zero.
Apply the same contrast
The fixture computes a two-group, two-period difference-in-differences contrast for the breach rate and for a separately specified control outcome. The breach contrast is negative; the negative-control contrast is zero. This is a calculation check. A zero on one small control metric does not prove parallel trends, no spillovers or accurate treatment dates. The rollout contrast explains the group-time target.
Placebo dates need clean pre-periods
A second test can pretend rollout happened earlier, using only periods before any actual adoption or anticipation. A nonzero placebo contrast may reveal preexisting divergence, but the chosen fake date and outcome window must be fixed before searching many alternatives. Using a period when staff were already training on the new screen is not a clean placebo. Event-time leads offer a related diagnostic with support counts.
Interpret a failed check carefully
A negative-control effect may mean confounding, a measurement change, an unexpected real pathway or random variation. Audit the control metric’s data pipeline and outcome eligibility before declaring the primary result invalid. Conversely, repeated null placebo findings do not supply the missing counterfactual; they eliminate some specific failure stories. Keep all planned tests in the review packet, including those that do not flatter the effect estimate.
Keep scope and uncertainty aligned
With few branches, a control-outcome estimate near zero can still have a wide interval. Show group counts, actual contrasts and data lineage rather than turning a nonsignificant p-value into evidence of no bias. If branches share routing, a supposedly unaffected outcome could move through spillovers. The exposure map helps identify that possibility.
Implementation
def two_group_change(treated_before, treated_after,
control_before, control_after):
return ((treated_after - treated_before) -
(control_after - control_before))
service_breach = two_group_change(.20, .13, .18, .17)
hardware_inventory = two_group_change(.30, .31, .29, .30)
assert abs(service_breach - (-.06)) < 1e-12
assert abs(hardware_inventory) < 1e-12
planned_checks = {"service_breach": service_breach,
"hardware_inventory_control": hardware_inventory}
assert set(planned_checks) == {"service_breach", "hardware_inventory_control"}Performance and operating cost
Each two-period contrast is O(1) time and space after cohort rates are computed. With M planned outcomes, calculation is O(M); data lineage and dependence-aware uncertainty are the expensive parts. Searching many controls or dates increases false reassurance unless the plan is frozen.
Common Mistakes
- Do not pick a control outcome solely because its estimated effect is zero.
- Do not use post-training periods as a pre-rollout placebo.
- Do not interpret a nonsignificant falsification test as proof of validity.
Read next
- Assignment-aware permutation tests for a few branches
- Leave-one-branch-out stability for rollout effects
- Direct and indirect exposure maps for branch rollouts
- A spillover difference-in-differences contrast
- Project: review a small branch rollout with spillovers
Continue the workflow: Synthetic-control placebos and fit ratios.
Continue the workflow: Transition windows and lag sensitivity for policy series.
