Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Imputation diagnostics: inspect draws and stress missing-not-at-random shifts

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Diagnostics compare plausible completed values with observed data and expose decisions that depend on assumptions about unseen outcomes.

Inspect by the variables that matter

A combined estimate can look stable while the completed resolution times for urgent cases violate known service limits. Compare observed and filled distributions within branch, priority, calendar period, and other predictors used by the decision. Review trace plots across imputation iterations if the algorithm samples them, impossible values, and changes in the substantive coefficient across completed datasets. A visual match alone does not prove the missingness assumption. Combination rules describe uncertainty conditional on the model that generated those draws.

Hold out observed cells as a diagnostic

Mask a declared subset of known outcomes, generate fills without access to them, and compare filled distributions and prediction errors with held-out truths. This checks whether the model can reconstruct cases resembling the observed ones. It cannot directly validate cases absent because unusually slow repairs never received a closing timestamp. Preserve the selection logic and random seed so an apparent improvement cannot be created by testing only easy cases.

Move the unseen outcomes deliberately

For sensitivity analysis, shift imputed missing-case resolution times upward by a plausible extra delay and refit or recompute the decision for each shift. A positive shift represents systematically slower unseen cases on the working scale. Choose the range from operational knowledge, not from the threshold that makes the result favorable. The code computes a simple completed mean from supplied observed and filled values to show how a shift changes the conclusion; a full analysis must re-estimate its model and uncertainty across imputations. The ledger provides the count exposed to this shift.

Report a tipping range

Present the smallest shift that changes the operational call, the missing fraction in every important group, and the assumption needed to exclude such a shift. If the sign changes under a modest plausible delay, mark the decision as fragile. A large number of imputations cannot make an implausible model safe. The project requires a predeclared service threshold and an explicit hold when the conclusion hinges on unseen urgent cases.

Implementation

python
def shifted_completed_mean(observed_minutes, filled_minutes, extra_missing_delay):
    if not observed_minutes or not filled_minutes:
        raise ValueError("observed and filled groups required")
    if extra_missing_delay < 0 or any(value < 0 for value in
                                      observed_minutes + filled_minutes):
        raise ValueError("durations and shift must be nonnegative")
    total = sum(observed_minutes) + sum(filled_minutes)
    return (total + len(filled_minutes) * extra_missing_delay) /            (len(observed_minutes) + len(filled_minutes))

assert shifted_completed_mean([18, 27, 33], [24], 0) == 25.5
assert shifted_completed_mean([18, 27, 33], [24], 12) == 28.5

Performance and operating cost

The mean calculation is O(n) time and O(1) extra space. Repeating the full imputation and analysis over s shifts and m imputations costs roughly s times m model fits. A quick shift grid is useful for screening, but final uncertainty must reflect the analysis actually used.

Common Mistakes

  • Reading matching observed and filled histograms as proof of missing-at-random.
  • Using post-outcome predictors unavailable at the intended decision time.
  • Selecting the sensitivity range after seeing the desired headline.
  • Shifting the point estimate while leaving a final uncertainty interval unchanged.

Read next

ai-data
applied-statistics
Storage details