Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Nonresponse bias: diagnose the missing outcomes before adjusting

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Nonresponse occurs when a selected unit lacks the outcome needed for analysis; its bias depends on how nonrespondents differ from respondents.

Separate absence types

A case never selected is not a nonrespondent. A selected case whose reviewer could not retrieve its attachment is. Record selection, contact or retrieval attempt, response state and reason independently. Compare response by channel, week and severity using fields known for the full selected sample. Missing-data policy also distinguishes unknown evidence from a measured zero.

Quantify what is known

For binary outcomes, a simple bound within the selected sample treats every missing case as non-defective for the lower bound and defective for the upper bound. If 31 of 47 selected cases were reviewed and 6 defects were found, the selected-sample defect fraction lies from 6/47 to 22/47 without further assumptions. These are deliberately wide bounds; they are not population confidence intervals and do not correct the sampling design.

Use follow-up to learn

Choose a random subset of nonrespondents for a more expensive retrieval attempt. Record the second-stage selection probability and outcome. A successful follow-up can reveal whether ordinary respondents were unusually easy or hard cases. Adjustment weights based on observed covariates only reduce bias under assumptions about response within those groups. If the missing outcome itself drives response, weighting cannot make the problem disappear.

Publish sensitivity, not certainty

Report the original response rate, response by stratum, follow-up yield and estimates under several plausible assumptions. Separate the observed rate among respondents from a weighted target-population estimate. When results change sign under modest assumptions about missing cases, the decision is not stable. More sampling may not help until the retrieval process improves.

Implementation

python
def binary_outcome_bounds(positive_count, reviewed_count, selected_count):
    if not 0 <= positive_count <= reviewed_count <= selected_count or selected_count == 0:
        raise ValueError("invalid selected-sample counts")
    missing_count = selected_count - reviewed_count
    return positive_count / selected_count, (positive_count + missing_count) / selected_count

assert binary_outcome_bounds(6, 31, 47) == (6 / 47, 22 / 47)

Performance and operating cost

The bound is O(1) time and space. Stratified follow-up requires storing selection and response events for each selected unit; inference cost depends on that design, not on this arithmetic.

Common Mistakes

  • Do not replace a missing binary outcome with zero without stating an assumption.
  • Do not call response weighting a guaranteed correction.
  • Do not present selected-sample bounds as population confidence intervals.

Read next

Continue the workflow: Proxy measurements and label error in operational datasets.

Continue the workflow: Cognitive pilots and survey branch-logic audits.

Continue the workflow: Nonresponse adjustment cells and the positivity check.

ai-data
data-science
Storage details