Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Overlap and heterogeneity diagnostics for uplift models

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An uplift estimate is locally credible only where both actions could have been observed for comparable eligible decisions.

Locate unsupported learners

If a policy always sends reminders to learners with an overdue assignment, the observational log contains no untreated comparison for that state. A model may still output a number, but it extrapolates beyond evidence. Inspect treatment propensity distributions for both arms and count decisions in predeclared slices. Overlap and weighting diagnostics explain why extreme inverse weights destabilize averages.

Use effective sample size, not raw count

For weights w, effective sample size is (sum w)^2 divided by sum(w squared). It can be far below the logged row count when a few rare assignments dominate. Report it separately for treated and control observations within each policy-relevant segment. A segment with 300 rows and an effective control sample of 18 should not receive a confident fine-grained ranking.

Separate heterogeneity from noise

Split learners by a feature defined before treatment, such as whether they have previously completed a course. Compare treated and control outcomes within each segment with uncertainty intervals, then ask whether segment effects differ. A positive point estimate in one tiny slice and a negative estimate in another is not evidence of a stable interaction. Reserve a randomized holdout for final evaluation.

Respect the unit of uncertainty

One learner may generate several eligible decisions. Resample learners, not rows, when estimating uncertainty for an average effect or policy value. Calendar weeks may also be correlated through campaigns and exams; block by week when this is material. Sample-ratio checks can reveal randomization defects before any uplift plot is interpreted.

Declare a fallback policy

Where overlap fails, do not smooth away the unsupported region and quietly target it. Keep a uniform randomized holdout, use a simple default action, or collect exploration data under a governed policy. Action logging and support connect this decision to the bandit setting; the current uplift analysis stays with one-shot treatment assignment.

Implementation

python
logged_decisions = [
    {"learner": "acct-47", "action": 1, "treatment_probability": 0.28},
    {"learner": "acct-62", "action": 0, "treatment_probability": 0.47},
    {"learner": "acct-83", "action": 1, "treatment_probability": 0.61},
    {"learner": "acct-94", "action": 0, "treatment_probability": 0.74},
]

def assignment_weight(record):
    probability = record["treatment_probability"]
    if not 0 < probability < 1:
        raise ValueError("missing action support")
    observed_probability = probability if record["action"] else 1 - probability
    return 1 / observed_probability

weights = [assignment_weight(record) for record in logged_decisions]
effective_count = sum(weights) ** 2 / sum(weight ** 2 for weight in weights)
assert round(max(weights), 3) == 3.846
assert effective_count < len(logged_decisions)

Performance and operating cost

A pass over N logged decisions costs O(N) time and O(N) memory if weights are retained, or O(1) extra memory for streaming sums. Segment and clustered-uncertainty checks add repeated passes or resampling fits. Trimming changes the target population, so record its rule and excluded share.

Common Mistakes

  • Do not treat an out-of-support prediction as an observed causal effect.
  • Do not report only raw segment counts when weights are concentrated.
  • Do not bootstrap repeated decisions as independent learner rows.

Read next

ai-data
machine-learning
Storage details