Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Missing outcomes: count absence before choosing an estimator

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An outcome analysis starts with the eligible frame, observed outcomes, and reasons that observations are absent.

Separate eligibility from observation

A returns team tracks whether each approved refund was completed within its service window. The denominator is every eligible refund entering a mature cohort, including cases whose final status cannot be reconstructed. An export of only completed cases changes the population and can make processing look faster. Define entry date, eligibility, outcome window, extraction date, and missing-status reason before computing a percentage. The estimand and frame determine what the result can describe.

Audit patterns in recorded variables

Count observed and missing outcomes by branch, intake week, priority, and handling channel. A difference in observation rates across these recorded fields warns that a complete-case average may be selected. It cannot prove that absence is unrelated to the unobserved outcome; that assumption cannot be read directly from the missing cells. The example counts a small segment ledger rather than filling gaps. The observation-process lesson distinguishes reasons for absence from statistical labels.

Choose the method for the claim

If observation can be explained by recorded covariates and every relevant group has a nonzero chance of observation, a weighted or model-based estimate may recover the target under those assumptions. If difficult refunds disappear because they are difficult, this assumption can fail even after conditioning. Record a plausible range for the unobserved outcome and carry it into the decision. The weighting lesson makes the observed-data assumption visible; bounds require less belief but may be wide.

Report the missingness itself

Deliver the eligible count, observed count, missing count, reasons, segment rates, and comparison of observed cases with all eligible cases. Then report the estimate and uncertainty under each stated assumption. A single filled value for each missing outcome hides model uncertainty and should not be presented as measured truth. The audit should also flag whether source-system changes altered which outcomes are visible. The project treats an apparent service improvement as provisional until those counts reconcile.

Implementation

python
from collections import defaultdict

def observation_by_branch(refund_rows):
    counts = defaultdict(lambda: [0, 0])
    for branch_id, outcome_observed in refund_rows:
        if not branch_id:
            raise ValueError("branch ID required")
        counts[branch_id][0] += 1
        counts[branch_id][1] += bool(outcome_observed)
    return {branch: {"eligible": total, "observed": observed}
            for branch, (total, observed) in counts.items()}

ledger = observation_by_branch([("north", True), ("north", False),
                                ("south", True), ("south", True)])
assert ledger["north"] == {"eligible": 2, "observed": 1}

Performance and operating cost

The ledger scan is O(n) time and O(g) space for n refunds and g branches. Reconstructing outcome history can cost much more than counting rows, but dropping absent outcomes cheaply changes the population. A more elaborate estimator does not rescue an incomplete eligibility ledger.

Common Mistakes

  • Dividing by observed outcomes instead of the eligible frame without saying so.
  • Calling missingness random because observed segments look similar.
  • Filling absent outcomes with one value and treating them as measured.
  • Comparing branches when their reporting systems have different capture rules.

Read next

Continue the workflow: Multiple imputation: combine estimates and both sources of uncertainty.

Continue the workflow: Nonresponse bounds: show what missing binary outcomes could change.

ai-data
applied-statistics
Storage details