An aggregate rate can reverse the ordering seen inside every recorded stratum when groups have different compositions.
Composition reversal: reconcile crude and within-stratum comparisons
Read the cells before the headline
A candidate claims workflow can have lower escalation rates than an incumbent workflow for both routine and complex cases, yet a higher overall escalation rate if it receives many more complex cases. The crude summary is not an arithmetic error; it averages different stratum rates with different mixes. Ask whether the target is each workflow’s actual workload or a comparison under the same workload before choosing a headline. Direct standardization expresses the second target.
Calculate a concrete reversal
In the illustrative code, the candidate rates are two percent for routine and twelve percent for complex cases; the incumbent rates are three percent and fourteen percent. Candidate work is sixty percent complex, while incumbent work is twenty percent complex. Its crude rate is therefore higher despite being lower in each stratum. Under one shared half-routine, half-complex reference mix, the ordering follows the stratum-specific pattern. The numbers are invented for this lesson and do not identify a causal workflow effect. The estimand decides which comparison is relevant.
Investigate why composition differs
Severity may have been measured before assignment, routed by staff, or recorded differently after the workflow change. Those histories have different implications. A pre-assignment severity variable can help describe like-with-like groups, but a variable changed by the workflow should not be adjusted away casually. Inspect missing severity labels and whether the same classification rules operated in both groups. Small or empty cells can make within-stratum rates unstable. Missingness audit extends to the variables that define strata.
Publish both views with their denominators
Show counts, rates and uncertainty for every stratum, each group’s crude mix, the chosen reference mix, and the standardized result. Explain why the crude and standardized summaries differ rather than hiding one. If policy makers want to know expected volume for each center’s actual queue, crude rates may remain operationally important even while standardized rates support a fairer process comparison. The project tests whether all target strata are observed before accepting either rank.
Implementation
def weighted_rate(routine_rate, complex_rate, complex_share):
if not all(0 <= value <= 1 for value in
(routine_rate, complex_rate, complex_share)):
raise ValueError("rates and mix must lie between zero and one")
return routine_rate * (1 - complex_share) + complex_rate * complex_share
candidate_crude = weighted_rate(0.02, 0.12, 0.60)
incumbent_crude = weighted_rate(0.03, 0.14, 0.20)
assert round(candidate_crude, 6) == 0.08
assert round(incumbent_crude, 6) == 0.052
assert weighted_rate(0.02, 0.12, 0.50) < weighted_rate(0.03, 0.14, 0.50)
Performance and operating cost
Each two-stratum rate is O(1) time and space; s strata take O(s). Data quality dominates computation: inconsistent severity coding or unsupported cells can invalidate a polished aggregate. The cheap crude rate remains useful for actual workload, but not as a like-for-like process comparison.
Common Mistakes
- Calling a reversal a calculation bug without checking group composition.
- Presenting only the standardized result while hiding actual workload rates.
- Adjusting for a stratum created after the workflow change without a causal plan.
- Ignoring small-cell uncertainty inside an apparently stable weighted average.
Read next
- Standardized rates: compare groups under one declared population mix
- Project: compare claims centers under a shared severity mix
- Population, estimand and sampling frame: name the quantity before calculating
- Rare proportions: keep interval uncertainty visible at zero and one
- Missing outcomes: count absence before choosing an estimator
