For a binary outcome, the full-population event rate lies between observed events divided by all eligible units and observed events plus missing outcomes divided by all eligible units.
Worst-case bounds for missing binary outcomes
Keep the full eligible denominator
A branch has 87 eligible tickets. The deadline result is known for 72; 19 of those breached. The other 15 outcomes are unavailable because one event stream failed. If every missing ticket avoided a breach, the overall rate is 19/87. If every missing ticket breached, it is 34/87. The observed-case rate 19/72 is neither endpoint and should not be labeled the full-branch rate. The assumption ledger names the binary support behind the bound.
Recognize what makes the bounds valid
The result needs a complete eligibility frame, correct observed labels and a binary outcome that must be zero or one for every eligible unit. If missing tickets could have been excluded under the frozen policy, eligibility itself is uncertain and the simple formula does not solve that problem. If outcome labels are wrong, add a separate measurement-error analysis. The label-error project covers that failure route.
Do not assume randomness of missingness
The bounds remain valid whether lost outcomes are mostly successes, mostly failures or selectively absent after a system outage, as long as the listed support and frame hold. They often become wide when many outcomes are missing. That width is the price of declining an unsupported missing-at-random assumption, not a flaw in the arithmetic. If all outcomes are observed, both endpoints coincide at the observed rate.
Report loss as a count and a fraction
Show 15/87 missing alongside the bounds. A one-point interval in a huge trial and a one-point interval in a tiny branch have different sampling uncertainty even if their identification widths match. The code works from counts; it does not produce a confidence interval for the endpoints. The group contrast combines two such intervals when assignment supports a comparison.
Use added information to tighten, not disguise
If operations recover outcomes for six of the missing tickets, update observed events and reduce the missing count. Each recovered binary outcome narrows the worst-case rate interval by 1/87, regardless of whether that ticket breached. The center may move in either direction. This points to a concrete data-collection plan instead of filling missing cells from an arbitrary average.
Implementation
def binary_rate_bounds(eligible, observed, observed_events):
if not (isinstance(eligible, int) and isinstance(observed, int)
and isinstance(observed_events, int)):
raise ValueError("counts must be integers")
if eligible <= 0 or not 0 <= observed_events <= observed <= eligible:
raise ValueError("inconsistent outcome counts")
missing = eligible - observed
return observed_events / eligible, (observed_events + missing) / eligible
lower, upper = binary_rate_bounds(eligible=87, observed=72,
observed_events=19)
assert abs(lower - 19 / 87) < 1e-12
assert abs(upper - 34 / 87) < 1e-12
assert abs((upper - lower) - 15 / 87) < 1e-12Performance and operating cost
The bound calculation is O(1) time and space after counts are available. Building and validating the eligibility frame is O(N) over N units; log recovery can dominate the entire analysis.
Common Mistakes
- Do not divide observed events by observed tickets when reporting a bound for all eligible tickets.
- Do not present worst-case endpoints as confidence limits.
- Do not treat an uncertain eligibility frame as a solved missing-outcome problem.
