Audit independent shipments and matched driver shifts before choosing a rank or paired comparison.
Project: review routing-time differences without losing the unit
Separate two study designs
A warehouse pilot contains independent shipment samples from two routing policies and a smaller set of drivers who tried both policies on matched shifts. These are not interchangeable rows. Define the dispatch-to-scan clock, eligible routes, driver and shift keys, assignment rules and missing scans for each design. The cross-group pair statistic applies to the independent cohort; within-driver changes apply to matched shifts.
Find the false sample expansion
The first export joins on driver ID alone. Drivers with several shifts appear repeatedly, creating a many-to-many table and a deceptively precise comparison. Rebuild one record per declared unit, quarantine duplicate shifts and report incomplete pairs. For independent shipments, keep driver clusters visible rather than pretending every parcel is unrelated. A single combined rank-test p-value from the malformed export has no defensible interpretation. Cluster-aware uncertainty may be needed after the join is repaired.
Compare effects and tails
For independent shipments, report the probability that a random policy-A shipment scans faster than a random policy-B shipment, including ties. For matched shifts, report positive, negative and zero change counts, plus a mean or median time difference. Inspect late-scan quantiles in both designs. The candidate may win many pairs yet make a small set of long routes unacceptable. Tail reporting is a release requirement rather than an optional plot.
Deliver a limited conclusion
Freeze the corrected extraction, unit inventory, effect summaries and uncertainty method. If matched order was randomized, retain the assignment record for a design-based check; if it was not, describe the contrast as observational. Hold a rollout if missing scans or late-route failures exceed the agreed limit. The sign lesson requires an explicit zero policy; the design lesson controls causal language.
Implementation
def routing_review_gate(report, limits):
if report["duplicate_unit_keys"]:
return "hold:unit-join"
if report["unmatched_shift_share"] > limits["maximum_unmatched_share"]:
return "hold:pair-coverage"
if report["late_scan_rate"] > limits["maximum_late_rate"]:
return "hold:tail-service"
if report["assignment_record_missing"]:
return "publish:observational-contrast"
return "review:design-based-contrast"
limits = {"maximum_unmatched_share": 0.07, "maximum_late_rate": 0.04}
report = {"duplicate_unit_keys": 3, "unmatched_shift_share": 0.03,
"late_scan_rate": 0.02, "assignment_record_missing": False}
assert routing_review_gate(report, limits) == "hold:unit-join"
assert routing_review_gate({**report, "duplicate_unit_keys": 0}, limits) == "review:design-based-contrast"
Performance and operating cost
The gate is O(1) time and space after the data audit. Rebuilding a unique-unit extraction is O(n) expected time with a keyed index, while credible uncertainty needs route or driver grouping. A fast rank calculation on a many-to-many join is cheaper only because it skips the actual study design.
Common Mistakes
- Joining matched shifts by driver alone.
- Pooling independent shipments with repeated driver shifts.
- Interpreting a rank statistic as a median difference.
- Ignoring a worse late-scan tail because most cross-group pairs favor the candidate.
Read next
- Independent rank comparisons: report the pair probability, not a median claim
- Paired sign checks: treat zeros and direction as design decisions
- Paired comparisons: analyze within-unit changes and preserve the match
- Standard error and cluster bootstrap: resample the independent unit
- Distribution summaries: report tails and define the outlier policy
Continue the workflow: Project: combine warehouse pilots without double-counting evidence.
