An exact sign-flip check compares a predeclared paired statistic with assignments permitted by the study design.
Paired randomization checks: enumerate sign assignments under a sharp null
State the null and assignment
Suppose each depot alternates two ticket interfaces across matched shifts and the interface order is randomized within each pair. Under a sharp null of no effect on any paired shift, the observed magnitude of each difference is fixed while its sign follows the randomized assignment. Enumerating those sign patterns gives a finite reference distribution. Without randomized sign assignment, treating the same enumeration as a design-based causal p-value needs an additional symmetry assumption. Pair construction is a prerequisite.
Choose a statistic in advance
Use the absolute mean paired difference when a two-sided change matters. Count all assignments whose absolute statistic is at least as large as the observed one, including the observed assignment. Divide by the number of permitted assignments. With four nonzero pairs, the smallest two-sided p-value from this simple sign-flip scheme is two out of sixteen, so a conventional five-percent decision is impossible even if every difference points the same way. That resolution is a design limit, not weak code.
Respect dependence and interference
If several shift pairs come from one depot and depot conditions move together, independently flipping every shift is not the assignment the experiment used. Flip or permute at the randomized unit. Spillover between interfaces can also break the sharp-null comparison. Inspect order, eligibility, exclusions and outcome timing before generating permutations. The unit lesson and the stopping rule prevent an exact computation from being attached to a changed design.
Report size as well as evidence
A p-value does not say how many minutes were saved. Report the paired mean or median difference, pair count, the complete distribution of differences and the predeclared statistic. A small design may yield an inconclusive test while still showing a range of practically important effects. Effect-size reporting keeps that uncertainty visible; the project uses the check only after duplicate pairs are fixed.
Implementation
from itertools import product
def paired_sign_flip_pvalue(differences):
if not differences:
raise ValueError("at least one pair required")
observed = abs(sum(differences) / len(differences))
extreme = 0
total = 0
for signs in product((-1, 1), repeat=len(differences)):
total += 1
permuted = abs(sum(sign * change for sign, change in
zip(signs, differences)) / len(differences))
extreme += permuted >= observed - 1e-12
return extreme / total
assert paired_sign_flip_pvalue([4, 3, 5, 2]) == 0.125
assert paired_sign_flip_pvalue([4, -3, 5, -2]) >= 0.125
Performance and operating cost
Exact enumeration takes O(n × 2^n) time and O(n) space for n pairs. It is sensible for a small design but not hundreds of pairs; a prespecified Monte Carlo approximation can then estimate the tail probability. The computation does not repair wrong pairing, nonrandom assignment or hidden dependence.
Common Mistakes
- Calling sign flips an exact randomized test when order was never randomized.
- Flipping shifts independently when depots were the assigned units.
- Choosing a one-sided direction after inspecting the differences.
- Interpreting a p-value as the probability that the null is true.
Read next
- Paired comparisons: analyze within-unit changes and preserve the match
- Project: compare support interfaces with complete agent-shift pairs
- Hypothesis tests: pair the decision rule with an effect size
- Experiment design: assign the right unit and guard against interference
- Multiple comparisons and peeking: protect a predeclared decision rule
