Diagnostics compare plausible completed values with observed data and expose decisions that depend on assumptions about unseen outcomes.
Imputation diagnostics: inspect draws and stress missing-not-at-random shifts
Inspect by the variables that matter
A combined estimate can look stable while the completed resolution times for urgent cases violate known service limits. Compare observed and filled distributions within branch, priority, calendar period, and other predictors used by the decision. Review trace plots across imputation iterations if the algorithm samples them, impossible values, and changes in the substantive coefficient across completed datasets. A visual match alone does not prove the missingness assumption. Combination rules describe uncertainty conditional on the model that generated those draws.
Hold out observed cells as a diagnostic
Mask a declared subset of known outcomes, generate fills without access to them, and compare filled distributions and prediction errors with held-out truths. This checks whether the model can reconstruct cases resembling the observed ones. It cannot directly validate cases absent because unusually slow repairs never received a closing timestamp. Preserve the selection logic and random seed so an apparent improvement cannot be created by testing only easy cases.
Move the unseen outcomes deliberately
For sensitivity analysis, shift imputed missing-case resolution times upward by a plausible extra delay and refit or recompute the decision for each shift. A positive shift represents systematically slower unseen cases on the working scale. Choose the range from operational knowledge, not from the threshold that makes the result favorable. The code computes a simple completed mean from supplied observed and filled values to show how a shift changes the conclusion; a full analysis must re-estimate its model and uncertainty across imputations. The ledger provides the count exposed to this shift.
Report a tipping range
Present the smallest shift that changes the operational call, the missing fraction in every important group, and the assumption needed to exclude such a shift. If the sign changes under a modest plausible delay, mark the decision as fragile. A large number of imputations cannot make an implausible model safe. The project requires a predeclared service threshold and an explicit hold when the conclusion hinges on unseen urgent cases.
Implementation
def shifted_completed_mean(observed_minutes, filled_minutes, extra_missing_delay):
if not observed_minutes or not filled_minutes:
raise ValueError("observed and filled groups required")
if extra_missing_delay < 0 or any(value < 0 for value in
observed_minutes + filled_minutes):
raise ValueError("durations and shift must be nonnegative")
total = sum(observed_minutes) + sum(filled_minutes)
return (total + len(filled_minutes) * extra_missing_delay) / (len(observed_minutes) + len(filled_minutes))
assert shifted_completed_mean([18, 27, 33], [24], 0) == 25.5
assert shifted_completed_mean([18, 27, 33], [24], 12) == 28.5
Performance and operating cost
The mean calculation is O(n) time and O(1) extra space. Repeating the full imputation and analysis over s shifts and m imputations costs roughly s times m model fits. A quick shift grid is useful for screening, but final uncertainty must reflect the analysis actually used.
Common Mistakes
- Reading matching observed and filled histograms as proof of missing-at-random.
- Using post-outcome predictors unavailable at the intended decision time.
- Selecting the sensitivity range after seeing the desired headline.
- Shifting the point estimate while leaving a final uncertainty interval unchanged.
Read next
- Multiple imputation: combine estimates and both sources of uncertainty
- Project: decide whether a service recovery target survives missing outcomes
- Missing outcomes: count absence before choosing an estimator
- Observation weighting: state positivity and test missing-outcome shifts
- Population, estimand and sampling frame: name the quantity before calculating
