Synthetic-control weights are nonnegative and sum to one; they are fitted on pre-intervention data and then held fixed to predict the untreated post-intervention path.
Synthetic-control weight search and prediction
Fit a path, not an impressive result
A branch starts the routing policy after June. The treated branch has a six-month breach-rate series, and two eligible donors have the same six months plus three later months. For each candidate weight, form a weighted donor rate in every pre-policy month and score its squared tracking error. The code searches weights from zero to one in steps of 0.01. It never consults the treated branch’s post-policy values while selecting weights. Donor eligibility must be settled first.
Understand the convex constraint
Two donor weights that are nonnegative and sum to one keep the synthetic path inside the donors’ monthly range. That avoids extrapolation through negative weights, but it may leave poor fit when the treated branch sits outside the donor envelope. Adding more donors can improve apparent fit while increasing design sensitivity. A large donor pool with a short pre-period can produce a fragile match even under the convex rule.
Separate the gap from the causal claim
Once weights are frozen, predict each post-policy month from donor outcomes and calculate observed minus synthetic. A negative gap means fewer breaches than predicted under the chosen construction. It is not automatically a policy effect: an unrelated staffing change or a donor-side shock could create the same pattern. Retain the full monthly gap sequence and the pre-fit error. A single averaged number hides reversals and delayed effects.
Keep preprocessing outside the outcome window
If seasonality, ticket mix or a reporting change matters, define transformations using pre-policy information and apply them consistently. Do not normalize by the full-series mean; that leaks post-policy information into the fit. Do not fill missing pre-period rates with post-period values. Where donor outcomes are counts and denominators differ, decide whether the target is a rate or a count before fitting.
Treat the grid search as a teaching implementation
The two-donor search below makes the constraint and held-out prediction visible. A larger production design needs a constrained optimizer, a clearly specified feature or outcome weighting rule and a reproducible tie-break. Document the software inputs and the fit window. Placebo comparisons and sensitivity reruns can then probe whether the gap is distinctive or design-dependent.
Implementation
def fit_two_donor_control(treated_pre, donor_pre, donor_post):
if len(donor_pre) != 2 or set(donor_pre) != set(donor_post):
raise ValueError("exactly two aligned donors required")
names = sorted(donor_pre)
if not treated_pre or any(len(donor_pre[name]) != len(treated_pre)
for name in names):
raise ValueError("incomplete pre-period")
if len({len(donor_post[name]) for name in names}) != 1:
raise ValueError("misaligned post-period")
candidates = []
for step in range(101):
first_weight = step / 100
weights = {names[0]: first_weight, names[1]: 1 - first_weight}
predicted_pre = [sum(weights[name] * donor_pre[name][month]
for name in names) for month in range(len(treated_pre))]
mse = sum((observed - predicted) ** 2
for observed, predicted in zip(treated_pre, predicted_pre)) / len(treated_pre)
candidates.append((mse, step, weights))
mse, _, weights = min(candidates)
predicted_post = [sum(weights[name] * donor_post[name][month]
for name in names)
for month in range(len(donor_post[names[0]]))]
return weights, mse, predicted_post
treated_pre = [.21, .20, .18, .19]
donor_pre = {"West": [.18, .18, .17, .17],
"South": [.24, .22, .19, .21]}
donor_post = {"West": [.19, .18], "South": [.22, .21]}
weights, pre_mse, expected_without_policy = fit_two_donor_control(
treated_pre, donor_pre, donor_post)
post_gaps = [observed - expected for observed, expected
in zip([.13, .12], expected_without_policy)]
assert abs(sum(weights.values()) - 1) < 1e-12
assert all(weight >= 0 for weight in weights.values())
assert len(post_gaps) == 2 and all(gap < 0 for gap in post_gaps)Performance and operating cost
With G candidate weights, P pre-period months, Q post-period months and two donors, grid search costs O(G × P + Q) time and O(G + Q) space here. Full constrained optimization with D donors has a different solver cost; storing and checking the panel can dominate either algorithm.
Common Mistakes
- Do not use treated post-policy outcomes to choose weights or scaling.
- Do not interpret a negative gap as causal without reviewing concurrent shocks.
- Do not hide poor pre-fit behind a post-period average.
