A placebo-in-space check repeats a fixed synthetic-control procedure on untreated units and compares post-period departures only among units with usable pre-period fit.
Synthetic-control placebos and fit ratios
Repeat the procedure, not only the arithmetic
For each eligible donor branch, pretend it received the policy at the real intervention date. Remove that branch from its own donor pool, refit weights using only its pre-policy data and calculate its post-policy gap. The originally treated branch should not be smuggled into placebo donor pools after treatment. This turns the placebo into a test of whether the observed gap is unusual under the same workflow, rather than a comparison with convenient static lines. The fit procedure must be identical across runs.
Compare pre and post error on the same scale
Root mean squared prediction error, or RMSPE, summarizes the size of a gap sequence. The post/pre RMSPE ratio highlights units that track well before the policy date and depart afterward. A branch with almost zero pre-error can have an enormous ratio from a small absolute post shift, so inspect both errors and the actual outcome scale. Also compare the treated branch’s pre-fit to placebo pre-fits before ranking ratios.
Gate poor matches before ranking
The teaching code accepts precomputed observed and predicted paths from independently refitted runs. It filters placebos whose pre-RMSPE exceeds twice the treated unit’s pre-RMSPE. That threshold is a disclosed diagnostic rule, not a universal scientific cutoff. If few placebos remain, the rank is coarse and sensitive to exclusions. Report the total candidate count, accepted count and every exclusion, including treatment or spillover in a placebo branch.
Read the rank with restraint
An observed ratio larger than every accepted placebo ratio suggests the treated branch is unusual in this donor set. It is not automatically a randomization p-value: branches may not have been exchangeably assigned treatment. Geography, volume and management selection can affect which branch received the policy and how it evolves. A placebo ranking cannot rescue a poor donor design. Assignment-aware tests need an actual assignment mechanism.
Inspect patterns, not just one statistic
Draw or list the treated and placebo gaps over time. A sudden treated departure at launch supports a different story from a divergence already underway during training. Check a fake intervention date in the true pre-period as another diagnostic, but keep that date outside anticipation. Negative controls probe a different failure route; neither check validates every assumption.
Implementation
from math import sqrt
def rmspe(observed, predicted):
if not observed or len(observed) != len(predicted):
raise ValueError("aligned nonempty paths required")
return sqrt(sum((actual - expected) ** 2
for actual, expected in zip(observed, predicted)) / len(observed))
def comparable_fit_ratios(refitted_paths, treated_branch, fit_limit=2.0):
if treated_branch not in refitted_paths or fit_limit <= 0:
raise ValueError("invalid treated branch or limit")
errors = {}
for branch, paths in refitted_paths.items():
pre = rmspe(paths["pre_observed"], paths["pre_predicted"])
post = rmspe(paths["post_observed"], paths["post_predicted"])
if pre == 0:
raise ValueError("zero pre-fit error needs separate handling")
errors[branch] = (pre, post)
treated_pre = errors[treated_branch][0]
accepted = {branch: post / pre for branch, (pre, post) in errors.items()
if branch == treated_branch or pre <= fit_limit * treated_pre}
rejected = sorted(set(errors) - set(accepted))
return accepted, rejected
paths = {"North": {"pre_observed": [.20, .19], "pre_predicted": [.19, .18],
"post_observed": [.12, .11], "post_predicted": [.19, .18]},
"West": {"pre_observed": [.16, .17], "pre_predicted": [.15, .16],
"post_observed": [.18, .19], "post_predicted": [.17, .18]},
"Harbor": {"pre_observed": [.22, .23], "pre_predicted": [.18, .19],
"post_observed": [.24, .25], "post_predicted": [.20, .21]}}
ratios, rejected = comparable_fit_ratios(paths, "North")
assert set(ratios) == {"North", "West"} and rejected == ["Harbor"]
assert ratios["North"] > ratios["West"]Performance and operating cost
Given U independently refitted units and T observations each, computing diagnostics is O(U × T) time and O(U) space. Re-fitting each unit is much more costly and can change the donor pool; the displayed function deliberately starts after those fits and cannot substitute for them.
Common Mistakes
- Do not call an observational placebo rank an exact randomization p-value.
- Do not compare a poorly fitted placebo with a well-fitted treated unit without disclosing pre-fit.
- Do not reuse treated post-policy outcomes as placebo donors.
