Risk difference, risk ratio and odds ratio summarize different aspects of the same two-group table and are not interchangeable.
Binary outcomes: report absolute risk, risk ratio, and odds on their own scales
Start with event counts and denominators
A duplicate-payment control flags 3 of 240 reviewed claims under one rule and 9 of 240 under another. Event risks are 1.25 percent and 3.75 percent. The absolute risk difference is minus 2.5 percentage points if the first rule is labeled candidate. The candidate risk ratio is one third. An odds ratio is calculated from event odds, not event probabilities, and is similar to the risk ratio here only because both event rates are low. Sparse-table testing addresses a null model, not which effect scale serves the decision.
Choose the scale before seeing the sign
A risk difference tells an operations team how many events might change per hundred comparable claims. A risk ratio expresses proportional change. An odds ratio is natural in logistic modeling and certain sampling designs, but should not be described to readers as a direct relative probability. A baseline risk near zero makes a ratio unstable or undefined even when an absolute difference remains meaningful. The logistic lesson translates fitted odds carefully.
Show uncertainty and design limitations
Report intervals appropriate to each effect scale and sparse counts, plus the actual table and exposure period. A naive Wald interval can behave badly near zero. If assignment was not randomized, changes in case mix or detection effort may explain the contrast. If claims are clustered by account or reviewer, standard errors based on independent claims can be too small. The cluster lesson starts from the independent unit.
Interpret zero events honestly
When the candidate group has no observed events, the point risk is zero, but population risk is not proved zero. If the baseline group also has no events, a risk ratio is undefined. The code returns no ratio in that case rather than inventing infinity or adding hidden pseudo-events. A reviewer should see both event counts and an uncertainty bound before a rule is declared safe. The project links effect scale to the cost per missed duplicate.
Implementation
def binary_effects(candidate_events, candidate_total,
baseline_events, baseline_total):
if min(candidate_total, baseline_total) <= 0 or not 0 <= candidate_events <= candidate_total or not 0 <= baseline_events <= baseline_total:
raise ValueError("valid events and positive denominators required")
candidate_risk = candidate_events / candidate_total
baseline_risk = baseline_events / baseline_total
return {"risk_difference": candidate_risk - baseline_risk,
"risk_ratio": candidate_risk / baseline_risk if baseline_risk else None}
comparison = binary_effects(3, 240, 9, 240)
assert round(comparison["risk_difference"], 4) == -0.025
assert round(comparison["risk_ratio"], 6) == round(1 / 3, 6)
assert binary_effects(0, 47, 0, 47)["risk_ratio"] is None
Performance and operating cost
The point-effect calculation is O(1) time and space. Intervals, cluster adjustment and design review cost more than arithmetic. A ratio with a tiny baseline can vary sharply as one event is reclassified, so report counts and absolute changes with it.
Common Mistakes
- Calling odds ratios risk ratios when outcomes are common.
- Suppressing denominators or using different observation windows.
- Calling zero observed events proof of zero population risk.
- Selecting only the scale that makes a change look largest.
Read next
- Sparse contingency tables: exact conditional inference for a two-by-two contrast
- Project: audit a rare duplicate-payment rule before replacing it
- Rare proportions: keep interval uncertainty visible at zero and one
- Logistic regression: translate odds coefficients back to risk
- Diagnostic performance: separate sensitivity, specificity and predictive value
