When missing outcomes may be systematically worse or better, vary that assumption and show when the decision reverses.
Missing-not-at-random sensitivity: test what unseen outcomes could change
Start with identified facts
For 47 eligible parcels, seven of 38 observed outcomes are late; nine outcomes are missing. The observed late share is 7/38, about 18.4%, but it is not necessarily the all-parcel late rate. The assumption-free range is 7/47 to 16/47. A sensitivity analysis narrows this range only by adding stated assumptions about the unseen parcels. The target population must stay fixed across scenarios.
Define transparent scenarios
Consider missing-parcel late shares of 0%, 25%, 50%, 75% and 100%. The implied full-population rate is (7 + 9 × assumed share)/47. These are scenario expectations and need not be integer late-parcel counts; do not portray them as observed outcomes. A carrier investigation could motivate a narrower set. The scenario grid is useful because readers can see precisely how much unobserved lateness the result requires.
Find the tipping point
Suppose the service target is a late rate below 25%. Solving (7 + 9 × q)/47 < 0.25 gives q < 19/36, about 52.8%. At or above that assumed late share among missing parcels, the target is no longer met under this model. This threshold is easier to review than a claim that the missing data are harmless. It does not estimate q.
Stress structure, not only one percentage
Missingness may concentrate in the partner carrier, where delay rates differ from ordinary parcels. A single q for every missing record hides that structure. Repeat the calculation by carrier and date with the original eligible denominators, then aggregate with the actual population mix. Also check whether scans simply arrive after the analysis cutoff; an ingestion lag is a measurement issue that can be investigated directly.
Connect the result to action
If plausible scenarios fall on both sides of the 25% target, report the target as unresolved. Request delayed scan recovery, a manual sample of missing parcels or carrier logs before making a stronger claim. If a decision must be made now, state which conservative rule it follows. Mechanism evidence informs the scenario range, while the sensitivity calculation shows its consequences.
Implementation
def implied_late_rate(observed_late, observed_total, missing_total,
assumed_missing_late_share):
if not 0 <= assumed_missing_late_share <= 1:
raise ValueError("missing late share must be between zero and one")
eligible = observed_total + missing_total
if missing_total < 0 or eligible <= 0 or not 0 <= observed_late <= observed_total:
raise ValueError("invalid observed counts")
return (observed_late + missing_total * assumed_missing_late_share) / eligible
assert round(implied_late_rate(7, 38, 9, 0.5), 3) == 0.245
assert implied_late_rate(7, 38, 9, 0.75) > 0.25Performance and operating cost
Each scenario rate costs O(1) time and space; G scenarios cost O(G) time. Stratified scenarios require count aggregation first. No computation can identify the unobserved share from observed records alone without additional assumptions or data.
Common Mistakes
- Do not call a scenario assumption an estimated fact.
- Do not report a single imputed rate without showing a range or tipping point.
- Do not combine groups with different missingness patterns before checking them separately.
Read next
- Missingness mechanisms: model why a value is absent
- Complete-case selection: know whose outcome remains
- Train-only imputation and missingness indicators
- Project: audit missing delivery scans before reporting service rates
Continue the workflow: Partial identification and an assumption ledger.
Continue the workflow: Observation weighting: state positivity and test missing-outcome shifts.
