A conditional exact test enumerates tables with the observed margins instead of relying on a large-cell approximation.
Sparse contingency tables: exact conditional inference for a two-by-two contrast
Specify what the four cells count
An audit flags duplicate claim payments under two review policies. The table records flagged and unflagged claims in each policy group, using one independent claim as the unit. Preserve the actual denominators, assignment process, and observation window; two flags on the same claim are not two independent trials. Sparse cells can make a large-sample chi-square approximation inaccurate. Rare-proportion intervals are also needed when a rate estimate is the main output.
Condition on the observed margins
For a fixed two-by-two table under the null odds ratio of one, the upper-left count follows a hypergeometric distribution when row and column totals are conditioned on. Enumerate feasible counts, calculate each table probability, and sum probability at or above the observed count for a prespecified one-sided alternative. The code below does exactly that. Two-sided exact p-values have more than one definition, so its one-sided tail must not be relabeled as a generic two-sided result. The testing lesson separates a tail probability from effect magnitude.
Look at design and scale
An exact conditional calculation is exact for its stated sampling model and tail rule; it cannot fix biased assignment, unequal follow-up, or dependent units. Report the four counts, absolute event rates, risk difference and a defensible uncertainty interval alongside any p-value. The odds ratio can be large while the absolute change is tiny when events are rare. The effect-scale lesson explains why product decisions often need both.
Do not repair zeros silently
A zero event cell does not license adding a hidden half-count and presenting the result as raw evidence. Some estimators use continuity corrections or penalized models, each with an interpretation cost. Choose the effect measure and interval method before inspecting the favorable direction, and disclose any correction. The project holds a launch claim if counted units, exposure windows, or policy assignment cannot be reconciled.
Implementation
from math import comb
def upper_tail_fixed_margins(policy_events, policy_total,
all_events, all_claims):
if not 0 <= policy_events <= policy_total <= all_claims or not 0 <= all_events <= all_claims or all_events - policy_events > all_claims - policy_total:
raise ValueError("inconsistent contingency margins")
denominator = comb(all_claims, policy_total)
upper = min(policy_total, all_events)
return sum(comb(all_events, event_count) *
comb(all_claims - all_events, policy_total - event_count)
for event_count in range(policy_events, upper + 1)
if 0 <= policy_total - event_count <= all_claims - all_events) / denominator
assert upper_tail_fixed_margins(2, 4, 2, 8) == 3 / 14
Performance and operating cost
For N claims and a feasible upper-left-cell range of width k, direct enumeration takes O(k) combinatorial evaluations and O(1) extra space. Large counts need numerically stable probability routines. The computation is trivial next to verifying claim independence and consistent policy exposure.
Common Mistakes
- Calling a one-sided tail a two-sided exact p-value.
- Treating a small p-value as a large operational benefit.
- Counting repeated claim flags as independent claims.
- Assuming exact arithmetic repairs selection or follow-up bias.
Read next
- Binary outcomes: report absolute risk, risk ratio, and odds on their own scales
- Project: audit a rare duplicate-payment rule before replacing it
- Rare proportions: keep interval uncertainty visible at zero and one
- Hypothesis tests: pair the decision rule with an effect size
- Experiment design: assign the right unit and guard against interference
