A rank comparison can measure how often an observation from one group is below one from another, including ties.
Independent rank comparisons: report the pair probability, not a median claim
Start with independent units
Two warehouse routing policies produce different dispatch-to-scan times. Compare one shipment from policy A with one independently sampled shipment from policy B. If many shipments share a route or driver, their rows are not independent units; group-aware assignment and uncertainty are needed. Freeze the eligible shipment frame and clock before calculating a rank statistic. The population definition decides which future shipments the result describes.
Use a quantity people can read
Count each cross-group pair where A is faster than B, give a tie half credit, and divide by the number of pairs. A value above one half favors A on this probabilistic ordering. It is related to a Mann–Whitney statistic, but a distributional test does not automatically test equality of medians: groups can differ in spread or shape. Report medians and tail quantiles separately. The distribution lesson prevents a single rank summary from hiding a long late-delivery tail.
Choose a valid reference distribution
Under randomized policy assignment, a permutation of policy labels at the assigned unit can test a sharp no-effect null. With observational groups, label permutation needs an exchangeability assumption that may fail when route mix differs. Ties and small samples affect exact test calculations; the pair-probability code below is an effect summary, not a p-value. If the same driver tried both policies on matched shifts, use a paired analysis instead of an independent-rank test.
State the operating consequence
A policy can win most random pairs while making the worst few shipments much slower. Review median, high quantile, missing scans and service-limit failures before a rollout. Keep the statistic, sample counts, assignment rule and uncertainty together in the decision packet. The paired sign lesson covers a different design; the project catches a mixed design whose first analysis treats matched shifts as independent.
Implementation
def faster_pair_probability(policy_a_minutes, policy_b_minutes):
if not policy_a_minutes or not policy_b_minutes:
raise ValueError("both routing groups need observations")
favorable = sum(a_time < b_time for a_time in policy_a_minutes
for b_time in policy_b_minutes)
ties = sum(a_time == b_time for a_time in policy_a_minutes
for b_time in policy_b_minutes)
comparisons = len(policy_a_minutes) * len(policy_b_minutes)
return (favorable + 0.5 * ties) / comparisons
probability = faster_pair_probability([4, 6, 8], [5, 7, 9])
assert round(probability, 6) == round(2 / 3, 6)
Performance and operating cost
The direct cross-pair calculation takes O(a × b) time and O(1) extra space for group sizes a and b. Sorting and rank aggregation can reduce cost on larger samples. Calculation speed is rarely the main risk: dependence, route imbalance and unobserved scan times can dominate inference.
Common Mistakes
- Calling a rank-test result a test of median equality without equal-shape assumptions.
- Treating shipments from the same route as independent units.
- Using an independent test for matched driver shifts.
- Reporting a favorable pair probability while hiding the late-delivery tail.
Read next
- Paired sign checks: treat zeros and direction as design decisions
- Project: review routing-time differences without losing the unit
- Distribution summaries: report tails and define the outlier policy
- Paired comparisons: analyze within-unit changes and preserve the match
- Experiment design: assign the right unit and guard against interference
