Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Independent rank comparisons: report the pair probability, not a median claim

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A rank comparison can measure how often an observation from one group is below one from another, including ties.

Start with independent units

Two warehouse routing policies produce different dispatch-to-scan times. Compare one shipment from policy A with one independently sampled shipment from policy B. If many shipments share a route or driver, their rows are not independent units; group-aware assignment and uncertainty are needed. Freeze the eligible shipment frame and clock before calculating a rank statistic. The population definition decides which future shipments the result describes.

Use a quantity people can read

Count each cross-group pair where A is faster than B, give a tie half credit, and divide by the number of pairs. A value above one half favors A on this probabilistic ordering. It is related to a Mann–Whitney statistic, but a distributional test does not automatically test equality of medians: groups can differ in spread or shape. Report medians and tail quantiles separately. The distribution lesson prevents a single rank summary from hiding a long late-delivery tail.

Choose a valid reference distribution

Under randomized policy assignment, a permutation of policy labels at the assigned unit can test a sharp no-effect null. With observational groups, label permutation needs an exchangeability assumption that may fail when route mix differs. Ties and small samples affect exact test calculations; the pair-probability code below is an effect summary, not a p-value. If the same driver tried both policies on matched shifts, use a paired analysis instead of an independent-rank test.

State the operating consequence

A policy can win most random pairs while making the worst few shipments much slower. Review median, high quantile, missing scans and service-limit failures before a rollout. Keep the statistic, sample counts, assignment rule and uncertainty together in the decision packet. The paired sign lesson covers a different design; the project catches a mixed design whose first analysis treats matched shifts as independent.

Implementation

python
def faster_pair_probability(policy_a_minutes, policy_b_minutes):
    if not policy_a_minutes or not policy_b_minutes:
        raise ValueError("both routing groups need observations")
    favorable = sum(a_time < b_time for a_time in policy_a_minutes
                    for b_time in policy_b_minutes)
    ties = sum(a_time == b_time for a_time in policy_a_minutes
               for b_time in policy_b_minutes)
    comparisons = len(policy_a_minutes) * len(policy_b_minutes)
    return (favorable + 0.5 * ties) / comparisons

probability = faster_pair_probability([4, 6, 8], [5, 7, 9])
assert round(probability, 6) == round(2 / 3, 6)

Performance and operating cost

The direct cross-pair calculation takes O(a × b) time and O(1) extra space for group sizes a and b. Sorting and rank aggregation can reduce cost on larger samples. Calculation speed is rarely the main risk: dependence, route imbalance and unobserved scan times can dominate inference.

Common Mistakes

  • Calling a rank-test result a test of median equality without equal-shape assumptions.
  • Treating shipments from the same route as independent units.
  • Using an independent test for matched driver shifts.
  • Reporting a favorable pair probability while hiding the late-delivery tail.

Read next

ai-data
applied-statistics
Storage details