Validate matched shifts, calculate within-agent changes and reject an invalid independent-row significance claim.
Project: compare support interfaces with complete agent-shift pairs
Freeze the paired design
A support desk tests two ticket interfaces on matched agent shifts. Assign interface order within each eligible pair, keep the ticket mix and timing rules, and predeclare resolution minutes as the outcome. The analysis unit is the matched agent-shift pair; if an agent contributes repeated shifts, dependence must be handled at the agent level. Save pair IDs and assignment order before outcomes arrive. The pairing contract defines which records belong together.
Audit the join
The first export contains the same agent-shift key twice for the candidate interface and lacks an incumbent record for a night shift. A naïve row join multiplies one observation and makes the sample size look larger. Quarantine duplicate keys, report missing pairs and inspect whether missingness tracks night work. Do not silently keep whichever duplicate appears first. The complete-case estimate represents only retained pairs unless a defensible missing-data plan expands the target. Missing-data policy should be set before comparing outcomes.
Compute a design-compatible comparison
Take candidate minus incumbent time for each validated pair; retain the full distribution, mean, standard error and pair count. If assignment was randomized within pairs and the sharp null is relevant, use the exact sign-flip check on independent assigned pairs. If the desk assigned one interface only after a training week, do not call that same calculation a design-based randomized test. The permutation calculation cannot supply randomization that never happened.
Decide and hand off
Report order balance, duplicate and missing-key counts, mean minutes saved, uncertainty, ticket-type slices and a practical capacity threshold. A negative average can still hide long resolution times for complex tickets. Hold rollout if the night-shift loss is material or if the pair join remains ambiguous. Deliver the frozen export, matching rule and decision log; link follow-up to effect-size interpretation and the stopping rule.
Implementation
def validate_shift_pairs(records):
by_pair = {}
for record in records:
key = (record["agent_id"], record["shift_id"])
slot = by_pair.setdefault(key, {})
variant = record["variant"]
if variant in slot:
return "hold:duplicate-variant"
slot[variant] = record["minutes"]
if any(set(slot) != {"incumbent", "candidate"} for slot in by_pair.values()):
return "hold:incomplete-pair"
changes = [slot["candidate"] - slot["incumbent"]
for slot in by_pair.values()]
return sum(changes) / len(changes)
shifts = [
{"agent_id": "agent-47", "shift_id": "am", "variant": "incumbent", "minutes": 31},
{"agent_id": "agent-47", "shift_id": "am", "variant": "candidate", "minutes": 27},
{"agent_id": "agent-82", "shift_id": "pm", "variant": "incumbent", "minutes": 29},
{"agent_id": "agent-82", "shift_id": "pm", "variant": "candidate", "minutes": 28},
]
assert validate_shift_pairs(shifts) == -2.5
assert validate_shift_pairs(shifts + [shifts[-1]]) == "hold:duplicate-variant"
Performance and operating cost
The dictionary join is O(n) expected time and O(n) space for n shift records. A production review must also check repeated-agent dependence and the reason pairs are missing. Cleaning the join and obtaining comparable shifts cost more than computing a mean; skipping them can turn a duplicate into false precision.
Common Mistakes
- Joining on agent alone and creating many-to-many matches.
- Keeping a duplicate variant because its value looks plausible.
- Treating repeated agent shifts as independent people.
- Using a randomized-test label for a nonrandom before/after rollout.
Read next
- Paired comparisons: analyze within-unit changes and preserve the match
- Paired randomization checks: enumerate sign assignments under a sharp null
- Experiment design: assign the right unit and guard against interference
- Hypothesis tests: pair the decision rule with an effect size
- Missing data policy: distinguish absence from a measured zero
Continue the workflow: Project: review routing-time differences without losing the unit.
