Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Worst-case bounds for missing binary outcomes

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

For a binary outcome, the full-population event rate lies between observed events divided by all eligible units and observed events plus missing outcomes divided by all eligible units.

Keep the full eligible denominator

A branch has 87 eligible tickets. The deadline result is known for 72; 19 of those breached. The other 15 outcomes are unavailable because one event stream failed. If every missing ticket avoided a breach, the overall rate is 19/87. If every missing ticket breached, it is 34/87. The observed-case rate 19/72 is neither endpoint and should not be labeled the full-branch rate. The assumption ledger names the binary support behind the bound.

Recognize what makes the bounds valid

The result needs a complete eligibility frame, correct observed labels and a binary outcome that must be zero or one for every eligible unit. If missing tickets could have been excluded under the frozen policy, eligibility itself is uncertain and the simple formula does not solve that problem. If outcome labels are wrong, add a separate measurement-error analysis. The label-error project covers that failure route.

Do not assume randomness of missingness

The bounds remain valid whether lost outcomes are mostly successes, mostly failures or selectively absent after a system outage, as long as the listed support and frame hold. They often become wide when many outcomes are missing. That width is the price of declining an unsupported missing-at-random assumption, not a flaw in the arithmetic. If all outcomes are observed, both endpoints coincide at the observed rate.

Report loss as a count and a fraction

Show 15/87 missing alongside the bounds. A one-point interval in a huge trial and a one-point interval in a tiny branch have different sampling uncertainty even if their identification widths match. The code works from counts; it does not produce a confidence interval for the endpoints. The group contrast combines two such intervals when assignment supports a comparison.

Use added information to tighten, not disguise

If operations recover outcomes for six of the missing tickets, update observed events and reduce the missing count. Each recovered binary outcome narrows the worst-case rate interval by 1/87, regardless of whether that ticket breached. The center may move in either direction. This points to a concrete data-collection plan instead of filling missing cells from an arbitrary average.

Implementation

python
def binary_rate_bounds(eligible, observed, observed_events):
    if not (isinstance(eligible, int) and isinstance(observed, int)
            and isinstance(observed_events, int)):
        raise ValueError("counts must be integers")
    if eligible <= 0 or not 0 <= observed_events <= observed <= eligible:
        raise ValueError("inconsistent outcome counts")
    missing = eligible - observed
    return observed_events / eligible, (observed_events + missing) / eligible

lower, upper = binary_rate_bounds(eligible=87, observed=72,
                                  observed_events=19)
assert abs(lower - 19 / 87) < 1e-12
assert abs(upper - 34 / 87) < 1e-12
assert abs((upper - lower) - 15 / 87) < 1e-12

Performance and operating cost

The bound calculation is O(1) time and space after counts are available. Building and validating the eligibility frame is O(N) over N units; log recovery can dominate the entire analysis.

Common Mistakes

  • Do not divide observed events by observed tickets when reporting a bound for all eligible tickets.
  • Do not present worst-case endpoints as confidence limits.
  • Do not treat an uncertain eligibility frame as a solved missing-outcome problem.

Read next

ai-data
data-science
Storage details