An exact randomization test reproduces the actual assignment mechanism under a sharp no-effect null instead of shuffling labels across incomparable units.
Blocked randomization inference: permute labels only inside assignment blocks
Match the test to assignment
A warehouse assigns a new scan procedure to half the shifts within each depot. Depot is an assignment block because route volume and scanner setup differ. Under the sharp null that the procedure changes no shift outcome, observed outcomes remain fixed while labels are reallocated according to the same within-depot randomization rule. A global permutation that moves assignments between depots was never possible in this experiment. The design lesson identifies the shift as the unit.
Enumerate a small design exactly
For each block, choose the same number of treated shifts as originally assigned; take the Cartesian product of those block-specific assignments. Calculate a prespecified statistic under each possible assignment. The code uses the mean of block treatment-minus-control differences and returns the upper-tail randomization p-value including the observed assignment. It is intentionally for tiny experiments: combinations multiply quickly. Ties count as at least as extreme in this one-sided test, and the test should not be mislabeled two-sided.
Understand what the null covers
A sharp null says every shift would have the same outcome under either assignment, allowing all counterfactual outcomes to be imputed from observed outcomes. Failure to reject is not proof of no useful effect; a confidence interval or effect estimate is needed for that. If outcomes are missing after assignment, permutation of only observed shifts breaks the design unless the missingness rule is justified. Stopping and multiplicity still matter if many endpoints were searched.
Keep interference on the table
A control shift sharing staff and a scanner with a treated shift may receive part of the intervention. Then the potential outcome of one shift depends on assignments around it, violating the simple no-interference reading. Record adjacency and operational overlap before interpreting a shift-level contrast. The interference lesson discusses exposure mapping; the project checks it before accepting the test.
Implementation
from itertools import combinations, product
def blocked_upper_tail(blocks, observed_statistic):
assignment_options = []
for outcomes, treated_count in blocks:
if not 0 < treated_count < len(outcomes):
raise ValueError("each block needs treated and control shifts")
assignment_options.append(list(combinations(range(len(outcomes)),
treated_count)))
extreme = total = 0
for assignment in product(*assignment_options):
block_differences = []
for (outcomes, _), treated_indices in zip(blocks, assignment):
treated = set(treated_indices)
treated_mean = sum(outcomes[i] for i in treated) / len(treated)
control_mean = sum(value for i, value in enumerate(outcomes)
if i not in treated) / (len(outcomes) - len(treated))
block_differences.append(treated_mean - control_mean)
statistic = sum(block_differences) / len(block_differences)
extreme += statistic >= observed_statistic - 1e-12
total += 1
return extreme / total
assert blocked_upper_tail([([21, 29], 1), ([33, 41], 1)], 8) == 0.25
Performance and operating cost
With b blocks, the exact enumeration count is the product of combinations of block size and treated count; runtime grows at least that fast. Memory stores each block assignment list. Larger experiments need a prespecified Monte Carlo draw from the same design, with its simulation error reported.
Common Mistakes
- Shuffling labels across blocks that never shared an assignment lottery.
- Calling a one-sided upper-tail value two-sided.
- Treating nonrejection as evidence of equivalence.
- Ignoring shared staff or other interference between assigned shifts.
Read next
- Interference exposure: decide whether the assigned unit can be isolated
- Project: audit a depot scan trial with blocked assignment and shared staff
- Experiment design: assign the right unit and guard against interference
- Paired randomization checks: enumerate sign assignments under a sharp null
- Multiple comparisons and peeking: protect a predeclared decision rule
