Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Complete-case selection: know whose outcome remains

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Deleting incomplete rows can change the population represented by a rate or model, even when the remaining table looks clean.

Name the target quantity

Suppose the service question is the late-delivery rate for all 47 dispatched parcels. Calculating late share only among parcels with a final scan estimates a rate in the scanned subset unless additional assumptions connect that subset to all dispatched parcels. A clean denominator is not the same as the intended denominator. State the eligible population, the outcome window and the point at which a parcel enters it.

Show selection numerically

Consider 20 partner parcels, of which eight have a missing final scan, and 27 ordinary parcels, of which one is missing. Dropping incomplete rows retains 12 partner and 26 ordinary parcels. The partner share falls from 20/47, about 42.6%, to 12/38, about 31.6%. Even before estimating lateness, the selected table represents a different mix. Whether the late rate is biased depends on how scan absence relates to lateness within these groups.

Use a selection diagram in words

A carrier affects both scan availability and delay risk. Conditioning on scan availability by dropping absent rows can change the carrier mix; a further relationship between delay and scanning may remain inside each carrier. Compare retained and excluded rows on variables known for everyone. The observation process tells you which variables to check. A reweighting calculation needs estimated observation probabilities and positivity, not just a wish to restore the row count.

Preserve honest bounds

If the binary late status of nine parcels is truly unknown, and seven of the other 38 were late, the overall late rate is between 7/47 and 16/47 without further assumptions: about 14.9% to 34.0%. These wide bounds may still answer whether a service threshold is definitely met. If they straddle the threshold, identify what follow-up evidence would narrow them. Sensitivity analysis can explore narrower, stated assumptions.

Report what was removed

Publish original eligible count, excluded count, reasons for exclusion, retained group mix and the analysis denominator. Keep excluded records in an audit table. Complete-case analysis can be appropriate for some estimands under some observation processes; it is not automatically valid because the missing fraction is small, nor automatically invalid because the fraction is large.

Implementation

python
def binary_rate_bounds(observed_late, observed_total, missing_total):
    if min(observed_late, observed_total, missing_total) < 0:
        raise ValueError("counts must be nonnegative")
    if observed_late > observed_total:
        raise ValueError("late count exceeds observed count")
    eligible = observed_total + missing_total
    if eligible == 0:
        raise ValueError("no eligible parcels")
    return observed_late / eligible, (observed_late + missing_total) / eligible

lower, upper = binary_rate_bounds(7, 38, 9)
assert round(lower, 3) == 0.149 and round(upper, 3) == 0.340

Performance and operating cost

The bound calculation is O(1) time and space after counts are known. Producing group-level counts from N rows is O(N). Bounds are intentionally broad because no unobserved late statuses are invented.

Common Mistakes

  • Do not present a complete-case rate as the full-population rate without assumptions.
  • Do not hide a changed group mix behind one overall missing percentage.
  • Do not impute a binary outcome as zero simply because its final scan is absent.

Read next

ai-data
data-science
Storage details