Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Blocked randomization inference: permute labels only inside assignment blocks

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An exact randomization test reproduces the actual assignment mechanism under a sharp no-effect null instead of shuffling labels across incomparable units.

Match the test to assignment

A warehouse assigns a new scan procedure to half the shifts within each depot. Depot is an assignment block because route volume and scanner setup differ. Under the sharp null that the procedure changes no shift outcome, observed outcomes remain fixed while labels are reallocated according to the same within-depot randomization rule. A global permutation that moves assignments between depots was never possible in this experiment. The design lesson identifies the shift as the unit.

Enumerate a small design exactly

For each block, choose the same number of treated shifts as originally assigned; take the Cartesian product of those block-specific assignments. Calculate a prespecified statistic under each possible assignment. The code uses the mean of block treatment-minus-control differences and returns the upper-tail randomization p-value including the observed assignment. It is intentionally for tiny experiments: combinations multiply quickly. Ties count as at least as extreme in this one-sided test, and the test should not be mislabeled two-sided.

Understand what the null covers

A sharp null says every shift would have the same outcome under either assignment, allowing all counterfactual outcomes to be imputed from observed outcomes. Failure to reject is not proof of no useful effect; a confidence interval or effect estimate is needed for that. If outcomes are missing after assignment, permutation of only observed shifts breaks the design unless the missingness rule is justified. Stopping and multiplicity still matter if many endpoints were searched.

Keep interference on the table

A control shift sharing staff and a scanner with a treated shift may receive part of the intervention. Then the potential outcome of one shift depends on assignments around it, violating the simple no-interference reading. Record adjacency and operational overlap before interpreting a shift-level contrast. The interference lesson discusses exposure mapping; the project checks it before accepting the test.

Implementation

python
from itertools import combinations, product

def blocked_upper_tail(blocks, observed_statistic):
    assignment_options = []
    for outcomes, treated_count in blocks:
        if not 0 < treated_count < len(outcomes):
            raise ValueError("each block needs treated and control shifts")
        assignment_options.append(list(combinations(range(len(outcomes)),
                                                   treated_count)))
    extreme = total = 0
    for assignment in product(*assignment_options):
        block_differences = []
        for (outcomes, _), treated_indices in zip(blocks, assignment):
            treated = set(treated_indices)
            treated_mean = sum(outcomes[i] for i in treated) / len(treated)
            control_mean = sum(value for i, value in enumerate(outcomes)
                               if i not in treated) / (len(outcomes) - len(treated))
            block_differences.append(treated_mean - control_mean)
        statistic = sum(block_differences) / len(block_differences)
        extreme += statistic >= observed_statistic - 1e-12
        total += 1
    return extreme / total

assert blocked_upper_tail([([21, 29], 1), ([33, 41], 1)], 8) == 0.25

Performance and operating cost

With b blocks, the exact enumeration count is the product of combinations of block size and treated count; runtime grows at least that fast. Memory stores each block assignment list. Larger experiments need a prespecified Monte Carlo draw from the same design, with its simulation error reported.

Common Mistakes

  • Shuffling labels across blocks that never shared an assignment lottery.
  • Calling a one-sided upper-tail value two-sided.
  • Treating nonrejection as evidence of equivalence.
  • Ignoring shared staff or other interference between assigned shifts.

Read next

ai-data
applied-statistics
Storage details