A paired design compares two observations from each independent unit, so the differences—not the raw rows—form the analysis sample.
Paired comparisons: analyze within-unit changes and preserve the match
Define the pair
A support team measures ticket resolution minutes under two interface variants for the same agent on matched shifts. The estimand is the mean within-agent change over the eligible agents and shift types. Join on a stable agent and shift key, then reject duplicate or incomplete pairs under a declared policy. Pairing removes stable agent differences from the comparison, but it does not erase a calendar shock that affects only one variant. Assignment design determines what the contrast can claim.
Analyze differences, not pooled records
Compute candidate minus incumbent time for each complete pair. Then summarize the distribution of those differences, its average and its standard error over independent pairs. A negative difference favors the candidate when shorter resolution is desired. If one agent contributes many shifts, the shifts are not automatically independent; aggregate by agent or use a dependence-aware method. Cluster resampling handles repeated agents better than treating every shift as a new person.
Understand what matching cannot fix
If agent behavior adapts after trying one interface, the second period can carry over the first. Counterbalance order when possible and record which variant came first. A matched before/after design without random assignment remains vulnerable to learning and workload changes. Inspect unmatched agents and why their pairs failed; excluding only difficult shifts changes the target population. The sampling frame should state who is represented.
Use a design-compatible test
A paired sign-flip randomization check is appropriate only when the assignment mechanism or sharp-null symmetry supports exchangeability of signs. A paired t procedure asks a different distributional question about mean differences. Choose before inspecting favorable results. The randomization lesson gives the exact small-sample calculation, and the project catches a duplicated agent key.
Implementation
from math import sqrt
def paired_shift_summary(incumbent_minutes, candidate_minutes):
if len(incumbent_minutes) != len(candidate_minutes) or len(incumbent_minutes) < 2:
raise ValueError("at least two complete pairs required")
changes = [candidate - incumbent for incumbent, candidate in
zip(incumbent_minutes, candidate_minutes)]
average = sum(changes) / len(changes)
variance = sum((change - average) ** 2 for change in changes) / (len(changes) - 1)
return {"pairs": len(changes), "mean_change": average,
"standard_error": sqrt(variance / len(changes))}
summary = paired_shift_summary([31, 29, 35, 33], [27, 28, 30, 31])
assert summary["pairs"] == 4
assert summary["mean_change"] == -3
assert summary["standard_error"] > 0
Performance and operating cost
Creating n pair differences is O(n) time and O(n) space; streaming mean and variance can use O(1) extra space. Matching and validating unique keys often cost more than the arithmetic. A false match can move the effect estimate and make its uncertainty appear much smaller than the design supports.
Common Mistakes
- Running an unpaired test on rows that were deliberately matched.
- Treating repeated shifts from one agent as independent agents.
- Dropping incomplete pairs without reporting who was lost.
- Calling a before/after difference causal without addressing period and order effects.
Read next
- Paired randomization checks: enumerate sign assignments under a sharp null
- Project: compare support interfaces with complete agent-shift pairs
- Experiment design: assign the right unit and guard against interference
- Standard error and cluster bootstrap: resample the independent unit
- Population, estimand and sampling frame: name the quantity before calculating
Continue the workflow: Paired sign checks: treat zeros and direction as design decisions.
Continue the workflow: Paired device agreement: bias, limits and decision tolerance.
