Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Synthetic-control donor eligibility and pre-fit

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A synthetic control builds a comparison trajectory from untreated units, but the donor set and the quality of pre-intervention fit must be audited before a post-intervention gap is interpreted.

Define the treated unit and clock

Suppose one service branch begins a routing policy in July. The outcome is the monthly share of tickets that miss a service deadline, with an unchanged denominator rule. Freeze January through June as the fitting window and July through September as the evaluation window. The treatment unit is the entire branch. A neighboring branch that starts training in June is already contaminated even if its feature flag remains off. The rollout clock should record training, launch and outcome dates separately.

Make donor eligibility a design decision

Candidate branches must remain untreated throughout the evaluation window and be outside plausible spillover paths. They also need the same outcome definition, complete pre-period records and no branch-specific reporting migration. Similar-looking outcome lines alone do not make a donor valid. A branch with a common queue shared with the treated branch may inherit the policy effect; using it as a donor would pull the counterfactual toward the observed outcome. The exposure screen gives a practical exclusion rule.

Look at the trajectory before optimization

Plot or inspect each candidate and the treated branch over the full fitting window. If the treated branch has a sharp pre-policy jump that no convex combination of donors can reproduce, a tiny post gap is not reassuring. A stable level match with a different slope is also weak. The pre-period should contain enough operational variation to test tracking, not merely two convenient points.

Use explicit support checks

The implementation below checks a frozen eligibility record, a complete monthly grid and a common cutoff. It refuses a donor with treatment, known indirect exposure or missing months. It does not decide whether unmeasured shocks are shared. Preserve the exclusion reasons; donor selection after seeing the post-policy result turns a design choice into an outcome search. Weight fitting begins only after this gate.

Know what the design can and cannot answer

The post-policy difference compares one branch with a modeled untreated path. A causal claim still needs no anticipation, no interference into donors and a credible claim that the donor combination would have tracked this branch without the policy. The fitted pre-period is evidence about that claim, not proof. With a single treated branch, the result concerns that branch and time span; moving to a new branch requires a separate transfer argument.

Implementation

python
def eligible_donors(monthly_rates, donor_status, pre_months, post_months):
    if not pre_months or not post_months or max(pre_months) >= min(post_months):
        raise ValueError("invalid intervention clock")
    required = set(pre_months) | set(post_months)
    approved = []
    rejected = {}
    for branch, status in donor_status.items():
        reasons = []
        if status["treated"] or status["indirect_exposure"]:
            reasons.append("policy exposure")
        if status["definition_changed"]:
            reasons.append("outcome definition changed")
        if any((branch, month) not in monthly_rates for month in required):
            reasons.append("incomplete monthly grid")
        if reasons:
            rejected[branch] = reasons
        else:
            approved.append(branch)
    return sorted(approved), rejected

rates = {(branch, month): .14 + month / 100
         for branch in ("West", "South", "Harbor") for month in range(1, 10)}
status = {"West": {"treated": False, "indirect_exposure": False,
                   "definition_changed": False},
          "South": {"treated": False, "indirect_exposure": True,
                    "definition_changed": False},
          "Harbor": {"treated": False, "indirect_exposure": False,
                     "definition_changed": False}}
approved, rejected = eligible_donors(rates, status, range(1, 7), range(7, 10))
assert approved == ["Harbor", "West"]
assert rejected["South"] == ["policy exposure"]

Performance and operating cost

For D candidates and T months, membership checks cost O(D × T) expected time and O(D) output space beyond the input panel. The real cost is collecting treatment, queue and measurement histories; a fast scan cannot validate an undocumented donor.

Common Mistakes

  • Do not choose donors after seeing the estimated post-policy gap.
  • Do not use a branch exposed through a shared queue as an untreated donor.
  • Do not treat close pre-period fit as proof of the missing counterfactual.

Read next

ai-data
data-science
Storage details