Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Observation weighting: state positivity and test missing-outcome shifts

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Inverse observation weights can target an eligible population under measured-data assumptions, while a sensitivity shift tests departures.

Estimate a chance of being seen

Suppose completed refund status is recovered for some but not all mature cases. Fit or specify an observation probability using variables recorded for every eligible refund, such as branch, intake week, and priority. Each observed case contributes inverse-probability weight, so under-observed groups receive more influence. The normalized weighted mean below uses observed outcomes and supplied probabilities; it does not estimate those probabilities. The eligible ledger is required first.

Check positivity and unstable weights

If a branch and priority combination has no observed outcomes, weighting cannot identify its mean from that stratum. Very small observation probabilities create huge weights and unstable estimates. Report their distribution, effective weight concentration, and comparisons before and after any predeclared cap. A cap trades variance for possible bias and cannot be chosen solely because it produces a preferred result. The survey-weight lesson describes related population weighting but a distinct sampling mechanism.

Ask what measured variables miss

The estimate needs an assumption that, conditional on included recorded variables, observation does not depend on the unobserved outcome. That can fail if difficult refunds are less likely to have a final status even within every recorded group. For a binary on-time outcome, examine a range in which missing cases are less successful than comparable observed cases; a sensitivity shift or bound shows how much the decision depends on that belief. The missing-not-at-random lesson makes such scenarios explicit.

Present assumption and consequence together

Show the complete-case value, weighted value, missing fraction, probability-model variables, smallest probabilities, and sensitivity range on one decision page. Uncertainty must account for estimating observation probabilities, not just for averaging known weights. A precise point estimate with a fragile observation assumption is not a precise operating conclusion. The project refuses a branch comparison when an entire high-priority stratum is unseen.

Implementation

python
def normalized_observation_mean(observed_outcomes, observation_probs):
    if not observed_outcomes or len(observed_outcomes) != len(observation_probs):
        raise ValueError("aligned observed outcomes required")
    if any(not 0 < probability <= 1 for probability in observation_probs):
        raise ValueError("each observation probability must be positive")
    weights = [1 / probability for probability in observation_probs]
    return sum(weight * outcome for weight, outcome in
               zip(weights, observed_outcomes)) / sum(weights)

assert normalized_observation_mean([6, 12], [0.5, 0.25]) == 10

Performance and operating cost

The weighted mean is O(n) time and O(n) space as written; a streaming accumulator can use O(1) extra space. Fitting the observation model and its uncertainty costs more. A zero-probability stratum is an identification failure, not a performance issue solved by faster code.

Common Mistakes

  • Estimating observation probabilities only among observed cases.
  • Allowing zero or near-zero probabilities without a support audit.
  • Calling weighted results immune to outcome-dependent missingness.
  • Ignoring uncertainty in fitted weights or choosing a cap after seeing the result.

Read next

Continue the workflow: Imputation diagnostics: inspect draws and stress missing-not-at-random shifts.

Continue the workflow: Nonresponse bounds: show what missing binary outcomes could change.

ai-data
applied-statistics
Storage details