A sensitivity audit repeats the full pre-period fit after defensible donor or time-window changes and reports how much the estimated post-policy gap moves.
Synthetic-control donor and window sensitivity
Prespecify plausible perturbations
One donor may have changed staffing, another may be unusually large, and the earliest months may use a different ticket classification. Freeze a short list of defensible reruns before reading their post-policy effects: drop each major donor in turn, start the fit two months later, and exclude a known outage month. Record the reason for each change. A post hoc search for the most favorable gap is selection, not sensitivity analysis. The donor gate supplies documented exclusion reasons.
Refit weights on every rerun
Dropping a donor while retaining the old weights leaves a weight vector that no longer sums to one. Changing the pre-period while retaining weights ignores the alternative design. Rebuild the eligible donor set, refit under the same constraint and recompute both pre-error and post gap. For a two-donor teaching model, deleting one donor leaves a single-donor comparator; that result can be informative but may fit poorly.
Report fit beside effect
A gap can grow when pre-fit deteriorates. The code below records one candidate result per declared design and rejects any design lacking a fitted pre-error or post gap. The displayed range is a design-sensitivity range, not a confidence interval and not a probability statement. If only the poorly fitting designs show a strong effect, the operational conclusion should be narrower. The placebo lesson discusses comparable fit across units.
Distinguish instability from invalidity
A donor deletion that changes the sign may identify a donor-specific shock, a narrow donor envelope or a branch that carries most of the pre-fit. Investigate operations and measurement history; do not automatically discard the offending design. Conversely, stable signs across a small handpicked set do not establish causality. Shared shocks, anticipation and interference may affect every rerun in the same direction.
Tie the window to the rollout process
A branch may prepare staff before the formal launch. Treating those preparation months as untreated fitting data can bend weights toward an already affected trajectory. Prefer the last clearly unexposed month as the cutoff, then test reasonable earlier cutoffs. An unusually short pre-period reduces the evidence that a donor mix tracks the branch. Event-time support helps locate such problems.
Implementation
def summarize_design_sensitivity(design_results, maximum_pre_error):
if maximum_pre_error <= 0 or not design_results:
raise ValueError("invalid design audit")
accepted = {}
rejected = {}
for design_name, result in design_results.items():
pre_error = result["pre_rmspe"]
post_gap = result["mean_post_gap"]
if pre_error < 0:
raise ValueError("negative fit error")
if pre_error > maximum_pre_error:
rejected[design_name] = "pre-fit exceeds declared limit"
else:
accepted[design_name] = post_gap
if not accepted:
raise ValueError("no design has acceptable pre-fit")
return min(accepted.values()), max(accepted.values()), rejected
# Each row is the output of a complete, separate constrained refit.
reruns = {"all_eligible": {"pre_rmspe": .008, "mean_post_gap": -.052},
"drop_west": {"pre_rmspe": .011, "mean_post_gap": -.041},
"later_start": {"pre_rmspe": .009, "mean_post_gap": -.047},
"drop_harbor": {"pre_rmspe": .026, "mean_post_gap": -.018}}
low, high, excluded = summarize_design_sensitivity(reruns, .015)
assert (low, high) == (-.052, -.041)
assert excluded == {"drop_harbor": "pre-fit exceeds declared limit"}Performance and operating cost
Summarizing R completed reruns is O(R) time and O(R) output space. The expensive operation is R independent refits, each over its donor count and pre-period length. Predefining a small audit set controls both compute and the risk of selecting a pleasing result.
Common Mistakes
- Do not remove a donor and reuse weights fitted with that donor present.
- Do not describe the rerun range as a confidence interval.
- Do not hide designs that failed a declared pre-fit gate.
