Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multiple imputation: combine estimates and both sources of uncertainty

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Multiple imputation creates several plausible completed datasets, fits the same analysis to each, and combines the resulting estimates and variances.

Begin with the missingness design

A repair network measures resolution time for eligible cases, but some completion timestamps are absent. Define eligibility, maturity, outcome, and observed predictors before considering filled values. The imputation model should preserve variables that matter to the analysis, including branch and priority interactions when the final model uses them. Draw several plausible values for each missing field rather than reusing one predicted mean. The missing-outcome ledger still counts every absent value; imputation does not erase it.

Fit first, combine second

Run the same substantive analysis on each completed dataset and collect one coefficient or contrast with its estimated sampling variance. Average the estimates. Average their variances to obtain within-imputation uncertainty; compute the sample variance of estimates to obtain between-imputation uncertainty. With m completed datasets, total variance is within plus (1 + 1/m) times between. The small code block combines supplied outputs only. It neither generates credible imputations nor supplies a final confidence interval, whose degrees of freedom need separate care.

Understand what changes the answer

If all completed datasets produce almost identical estimates, between-imputation variance is small; a large spread says missing values materially affect the result. Extra imputations reduce Monte Carlo noise but cannot repair a wrong observation model. A predictor measured after resolution may leak the outcome into the imputation process. If missingness depends on the unrecorded delay after conditioning on observed predictors, ordinary missing-at-random imputation can remain biased. Observation weighting shares that identification limit.

Publish the analysis boundary

Report eligible and missing counts, imputation model variables, number of completed datasets, diagnostics, pooled estimate, within and between components, and sensitivity to a worse missing-case scenario. Keep transformed coefficients on the analysis scale used for combining, then convert for interpretation if justified. A single imputed dataset or a pooled row-level average understates uncertainty. The diagnostics lesson tests whether plausible draws and sensitivity ranges agree with the operational decision.

Implementation

python
from math import sqrt

def combine_imputation_estimates(estimates, sampling_variances):
    if len(estimates) < 2 or len(estimates) != len(sampling_variances):
        raise ValueError("at least two paired analyses are required")
    if any(variance < 0 for variance in sampling_variances):
        raise ValueError("sampling variances must be nonnegative")
    count = len(estimates)
    pooled = sum(estimates) / count
    within = sum(sampling_variances) / count
    between = sum((estimate - pooled) ** 2 for estimate in estimates) / (count - 1)
    return pooled, sqrt(within + (1 + 1 / count) * between)

estimate, standard_error = combine_imputation_estimates([4.2, 5.1, 4.8],
                                                         [0.36, 0.49, 0.25])
assert round(estimate, 3) == 4.7
assert round(standard_error ** 2, 6) == round((0.36 + 0.49 + 0.25) / 3 +
                                             (4 / 3) * 0.21, 6)

Performance and operating cost

Combining m scalar estimates is O(m) time and O(1) extra space after the results exist. Generating m completed datasets and refitting the substantive model is the dominant cost. More runs reduce simulation noise; they do not buy identification when missingness depends on unrecorded outcomes.

Common Mistakes

  • Using one filled dataset and its usual standard error as the final analysis.
  • Averaging completed rows before fitting, which discards between-imputation spread.
  • Leaving an interaction out of the imputation model when it defines the target contrast.
  • Claiming missing-at-random behavior was proved by comparing observed groups.

Read next

ai-data
applied-statistics
Storage details