Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Bounded missing-outcome risk scenarios

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A missing-outcome risk restriction narrows worst-case bounds by assigning a credible lower and upper event fraction to the missing units in each group.

State the restriction as a rate

Suppose 10 assigned-policy outcomes are missing. Rather than pretend none or all breached, an audit of similar outage tickets may support a missing-ticket breach fraction between 0.20 and 0.60. For five missing control outcomes, a separate range from 0.10 to 0.40 may be plausible. These are assumptions about missing units, not observed rates and not probabilities assigned to particular records. Keep the rationale and population match in the ledger. The ledger lesson distinguishes observed facts from restrictions.

Calculate arm bounds before the contrast

The treated group has 15 observed breaches among 50 assigned and 10 missing. Its rate lies between (15 + 10 × 0.20)/50 = 0.34 and (15 + 10 × 0.60)/50 = 0.42. The control group has 20 observed breaches, five missing, and risk range 0.10 to 0.40, yielding 0.41 to 0.44. The treated-minus-control contrast lies between -0.10 and +0.01. It still crosses zero. The worst-case contrast was wider.

Do not confuse plausible with identified

The observed data do not verify that missing tickets have those risk ranges. A validation sample drawn from a different outage or branch may not transfer. If the imposed ranges exclude a realistic failure mode, the narrow interval is false reassurance. Show the unrestricted result alongside the restricted result, with the rule that made each endpoint smaller. The code rejects risk ranges outside zero to one and inconsistent arm counts.

Preserve group-specific missingness

Applying one missing-risk range to both groups may be unjustified if policy launch changes logging or ticket mix. Conversely, inventing different ranges solely to obtain a preferred effect is equally weak. Use process logs, recovered ticket samples and domain constraints to motivate each arm. Missing-not-at-random sensitivity examines a related problem using model-based shifts.

Propagate further uncertainty separately

The scenario interval is conditional on exact counts and fixed risk restrictions. It is not a statistical confidence interval and does not account for small-sample variation in the observed breaches. If a random sample of missing outcomes is recovered, the plausible ranges can be re-estimated with uncertainty, but the selection mechanism of that recovery sample must be documented. The decision lesson checks whether the remaining range matters.

Implementation

python
def risk_restricted_rate(assigned, observed, breaches, missing_risk):
    low_risk, high_risk = missing_risk
    if assigned <= 0 or not 0 <= breaches <= observed <= assigned:
        raise ValueError("inconsistent arm counts")
    if not 0 <= low_risk <= high_risk <= 1:
        raise ValueError("invalid missing-outcome risk range")
    missing = assigned - observed
    return ((breaches + missing * low_risk) / assigned,
            (breaches + missing * high_risk) / assigned)

treated = risk_restricted_rate(50, 40, 15, (.20, .60))
control = risk_restricted_rate(50, 45, 20, (.10, .40))
effect = (treated[0] - control[1], treated[1] - control[0])
assert all(abs(actual - expected) < 1e-12
           for actual, expected in zip(treated, (.34, .42)))
assert all(abs(actual - expected) < 1e-12
           for actual, expected in zip(control, (.41, .44)))
assert all(abs(actual - expected) < 1e-12
           for actual, expected in zip(effect, (-.10, .01)))

Performance and operating cost

Two-arm bound evaluation is O(1) time and space. A grid over K plausible risk-range pairs costs O(K). Establishing credible restrictions through audits and preserving a sampling design is more expensive than the interval arithmetic.

Common Mistakes

  • Do not label an assumption-dependent range a confidence interval.
  • Do not choose missing-risk values after inspecting which ones favor the rollout.
  • Do not silently use the observed-case risk as the missing-case risk.

Read next

ai-data
data-science
Storage details