Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Nonresponse bounds: show what missing binary outcomes could change

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

When selected units have missing binary outcomes, the full selected-sample rate lies between treating all missing as zero and all missing as one.

Keep selected and answered units separate

A compliance team draws 91 branches from its register; 73 return a yes-or-no audit result and 18 do not. The observed respondent rate is successes divided by 73, but the selected-sample rate uses 91 branches. If missing outcomes may differ from observed outcomes, a response-weighted estimate needs assumptions or useful auxiliary information. The crude bounds require no such outcome assumption: divide known successes by selected count for the low end, and add missing count for the high end. The missing-outcome ledger records why each branch is absent.

Read bounds as a sensitivity range

Suppose 52 of 73 responding branches pass. Among the 91 selected, the pass rate is at least 52 divided by 91 and at most 70 divided by 91. These are logical bounds for the selected set, not a confidence interval for all 347 branches; sampling uncertainty around the population target is another layer. The code enforces the arithmetic and reports response fraction. A narrow range arises when few selected outcomes are missing, not when respondents happen to look similar on an unrelated field.

Investigate selection before weighting

Compare response by region, branch size, visit type and prior compliance if those fields exist for every selected branch. A high response fraction alone does not rule out bias when the few missing branches are systematically different; a low fraction does not prove a large bias. Weighting on observed fields can reduce some imbalance if response is explainable by those fields, but it cannot identify missing outcomes without assumptions. Response weighting and calibration need those assumptions made explicit.

Use a decision threshold honestly

If even the lower bound exceeds a required pass threshold, missing outcomes cannot reverse the selected-sample decision. If the interval straddles the threshold, follow-up measurement has direct value. Do not apply a finite-population correction to erase nonresponse; it addresses random selection variance only. Report selected count, answered count, reasons, outcome bounds, response-by-group checks and the target population. The project combines these checks without confusing logical bounds with uncertainty intervals.

Implementation

python
def selected_binary_bounds(selected_count, answered_count, known_successes):
    if not 0 < selected_count or not 0 <= known_successes <= answered_count <= selected_count:
        raise ValueError("counts must describe one selected sample")
    missing_count = selected_count - answered_count
    return (known_successes / selected_count,
            (known_successes + missing_count) / selected_count,
            answered_count / selected_count)

low, high, response_fraction = selected_binary_bounds(91, 73, 52)
assert round(low, 3) == 0.571 and round(high, 3) == 0.769
assert round(response_fraction, 3) == 0.802

Performance and operating cost

The bounds use O(1) time and space from reconciled counts. Building a per-branch response ledger costs O(n) storage and enables group audits. Extra modeling may narrow a decision range, but only by adding assumptions and independent information that the simple bounds intentionally avoid.

Common Mistakes

  • Calling the logical bounds a confidence interval for the population rate.
  • Using the respondent denominator for the selected-sample bound.
  • Assuming a high response rate guarantees no nonresponse bias.
  • Claiming the finite-population correction fixes unobserved outcomes.

Read next

ai-data
applied-statistics
Storage details