An S-learner fits one outcome model with action as an input; a T-learner fits separate outcome models by action and subtracts their predictions.
S-learners and T-learners for treatment-effect baselines
Use a small baseline first
For reminder targeting, estimate completion under action and control within prespecified learner segments before fitting a flexible model. If eight eligible learners in a segment were randomized, four assigned reminders with three completions and four assigned control with one completion, the observed segment contrast is 3/4 minus 1/4, or 0.50. It is noisy. It is still a useful arithmetic test for the event pipeline.
Understand what each learner fits
An S-learner uses one function m(x,a) and computes m(x,1) minus m(x,0). Shared parameters help small arms, but heavy regularization may suppress the action feature and flatten real heterogeneity. A T-learner fits m1(x) only on treated rows and m0(x) only on controls. It can capture distinct response patterns, yet each arm has fewer examples and extrapolates where its own arm is sparse.
Keep the split outside both fits
Split by learner and decision time before tuning features, hyperparameters or segment boundaries. Fit preprocessing only on the training period. A categorical encoding learned from future cohorts can leak changes in the catalog or learner mix. Temporal validation protects the forecast setting; uplift evaluation checks whether model differences actually improve decisions.
Distinguish outcome fit from effect fit
Low log loss for both arm models does not establish accurate differences. If both models make the same 11-point upward error for one learner, their difference might survive; if the errors oppose each other, uplift can be badly wrong despite decent overall accuracy. Compare effect ranking, policy value, arm-specific calibration and segment support on held-out randomized assignments.
Choose complexity for the evidence
If treatment is rare or segments are thin, shrink toward a simpler common model and report uncertainty rather than claiming a precise per-learner effect. In observational data, separate arm regressions alone do not remove confounding. An X-learner can use imputed effects when arm sizes differ, while augmented weighting scores also use assignment probabilities.
Implementation
from statistics import mean
randomized_segment = [
{"action": 1, "completed": 1}, {"action": 1, "completed": 1},
{"action": 1, "completed": 0}, {"action": 1, "completed": 1},
{"action": 0, "completed": 1}, {"action": 0, "completed": 0},
{"action": 0, "completed": 0}, {"action": 0, "completed": 0},
]
def segment_uplift(assignments):
treated = [row["completed"] for row in assignments if row["action"] == 1]
control = [row["completed"] for row in assignments if row["action"] == 0]
if not treated or not control:
raise ValueError("both randomized arms are required")
return mean(treated) - mean(control)
assert segment_uplift(randomized_segment) == 0.5Performance and operating cost
The arithmetic segment baseline is O(N) time and O(N) temporary space. For a T-learner, fitting cost is the sum of two outcome-model fits; an S-learner fits once but predicts twice per learner. Cross-validation multiplies those costs. Arm-specific sample size, rather than total row count, limits credible heterogeneity.
Common Mistakes
- Do not treat an outcome-model accuracy score as an uplift score.
- Do not create segments after looking at test outcomes.
- Do not fit a separate arm model where that arm has no local support.
