A branch-level randomization test compares an observed treatment contrast with the contrasts possible under the actual assignment rule and a sharp no-effect null.
Assignment-aware permutation tests for a few branches
Recover the assignment mechanism
Suppose five eligible branches entered a lottery that selected exactly two to receive a triage screen in April. Under a sharp null that the screen changed no branch outcome, the observed before-to-after change for each branch would be the same under every allowed assignment. Enumerate the ten possible treated pairs and calculate the treated-minus-control change for each. The observed pair is one member of that set. If executives selected branches because their breach rates were rising, uniform reassignment is fictional and the resulting number is not an exact randomization p-value. The rollout log should retain assignment rules.
Keep the unit intact
The lottery acts on branches. Do not shuffle monthly rows or cases independently; that breaks the treatment unit and repeated-outcome structure. The fixture uses one change per branch after the branch-month panel has been audited. Its observed contrast is the strongest negative contrast among the ten allowed pairs, so a two-sided tail contains one assignment and has probability 0.10. A small assignment space limits attainable p-values regardless of how impressive the raw rate difference looks.
State the tested null
The code checks a sharp null of no effect on any branch. It does not estimate a confidence interval for an average effect or test a vague claim that the average effect is zero. Under heterogeneous effects those are different questions. A two-sided statistic also needs a prespecified direction or absolute-value rule. Do not adjust the tail after seeing the sign. For staggered or stratified lotteries, enumerate only assignments the design actually allowed, preserving cohort sizes and restrictions.
Do not turn observational placebo assignments into a lottery
In a nonrandom rollout, recomputing the statistic over other branches can reveal whether the observed branch is unusually extreme, but exchangeability requires an argument. Different baseline trends, branch sizes and rollout priorities can destroy it. Report such a calculation as a sensitivity diagnostic, not design-based exact inference. Branch influence is another diagnostic when a few units drive the result.
Tie the output to action
The operational packet should show the branch changes, observed contrast, allowed assignment count, null, statistic and smallest attainable tail probability. If there are too few independent assignments to distinguish a policy-relevant effect, say so. No method manufactures information that was never in the rollout. The review project keeps this limitation visible.
Implementation
from itertools import combinations
def assignment_test(branch_changes, treated_branches):
branches = sorted(branch_changes)
treated = set(treated_branches)
if not treated or treated == set(branches) or not treated <= set(branches):
raise ValueError("invalid treated set")
def contrast(selected):
selected = set(selected)
control = set(branches) - selected
treated_mean = sum(branch_changes[branch] for branch in selected) / len(selected)
control_mean = sum(branch_changes[branch] for branch in control) / len(control)
return treated_mean - control_mean
observed = contrast(treated)
possible = [contrast(candidate) for candidate in combinations(branches, len(treated))]
extreme = sum(abs(value) >= abs(observed) - 1e-12 for value in possible)
return observed, extreme / len(possible), len(possible)
branch_changes = {"North": -.07, "East": -.05, "West": .01,
"South": .00, "Central": .02}
observed, tail_probability, assignments = assignment_test(
branch_changes, {"North", "East"})
assert assignments == 10
assert abs(observed - (-.07)) < 1e-12
assert abs(tail_probability - .1) < 1e-12Performance and operating cost
Enumerating every assignment costs O(C(B,T) × B) time and O(C(B,T)) space for B branches and T treated branches. Larger lotteries require a design-preserving sample of assignments. The limiting factor with few branches is assignment support, not computation.
Common Mistakes
- Do not call a permutation p-value exact when rollout assignment was not exchangeable.
- Do not shuffle cases when branches received treatment.
- Do not confuse a sharp no-effect test with uncertainty for an average effect.
Read next
- Leave-one-branch-out stability for rollout effects
- Direct and indirect exposure maps for branch rollouts
- Negative-control outcomes and placebo rollout checks
- Project: review a small branch rollout with spillovers
Continue the workflow: Treatment-effect bounds under differential attrition.
