An augmented uplift score combines predictions under each action with an inverse-assignment correction for the action actually observed.
Cross-fitted augmented weighting scores for uplift
Build the score from two nuisance estimates
Let m1 and m0 be predicted completion probabilities under reminder and control, and let e be the probability of reminder assignment. For a treated row, add (outcome minus m1) divided by e to m1 minus m0. For a control row, subtract (outcome minus m0) divided by (1 minus e). The resulting pseudo-outcome is noisy; it is not a literal personal counterfactual.
Check the arithmetic
Suppose m1 is 0.42, m0 is 0.31 and e is 0.40. A treated learner who completes gets 0.42 minus 0.31 plus (1 minus 0.42)/0.40 = 1.56. A score above one is legitimate: inverse weighting magnifies a rare observed outcome. It should not be displayed as a probability. Averaging scores over supported groups can target an average effect under appropriate assumptions.
Cross-fit before learning heterogeneity
Partition learners into folds. Fit outcome and propensity models on every fold except one, predict the held-out fold, construct scores there, then train a final effect model from those out-of-fold scores. Keep every episode of one learner in the same fold. This reduces own-row overfitting; it does not create identification when unmeasured factors affect both reminders and completion.
Use logged randomization when available
If assignment was randomized with logged e, use that known probability rather than estimating it. For observational assignment, propensity estimation needs only pre-action confounders, a credible causal graph, and stable action definitions. Covariate timing is the first defense. A wrong outcome model and wrong propensity model can both fail; two-model correction never means immunity to missing confounders.
Inspect unstable corrections
When e approaches zero or one, residual terms explode. Report effective sample size, weight tails and treatment counts by score band; predeclare any trimming policy. Overlap diagnostics decide where estimates are supportable. Keep model tuning out of the final randomized test, where policy value is measured.
Implementation
def doubly_robust_uplift_score(action, outcome, treated_prediction,
control_prediction, treatment_probability):
if action not in (0, 1) or outcome not in (0, 1):
raise ValueError("binary action and outcome required")
if not 0 < treatment_probability < 1:
raise ValueError("both actions need positive probability")
baseline = treated_prediction - control_prediction
if action == 1:
return baseline + (outcome - treated_prediction) / treatment_probability
return baseline - (outcome - control_prediction) / (1 - treatment_probability)
treated_score = doubly_robust_uplift_score(1, 1, 0.42, 0.31, 0.40)
assert abs(treated_score - 1.56) < 1e-12Performance and operating cost
One score costs O(1) time and space after nuisance predictions exist; N scores cost O(N). Cross-fitting K folds trains nuisance models K times and stores out-of-fold predictions. Inverse weights can dominate variance when one action is rare, so compute cost is seldom the binding limit; data support is.
Common Mistakes
- Do not call a pseudo-outcome a personal treatment probability.
- Do not use in-sample nuisance predictions when claiming cross-fitted evaluation.
- Do not infer causal protection from two-model correction if confounders are missing.
