Skip to content
AITroveRead. Build. Understand.
Make this comfortable

S-learners and T-learners for treatment-effect baselines

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An S-learner fits one outcome model with action as an input; a T-learner fits separate outcome models by action and subtracts their predictions.

Use a small baseline first

For reminder targeting, estimate completion under action and control within prespecified learner segments before fitting a flexible model. If eight eligible learners in a segment were randomized, four assigned reminders with three completions and four assigned control with one completion, the observed segment contrast is 3/4 minus 1/4, or 0.50. It is noisy. It is still a useful arithmetic test for the event pipeline.

Understand what each learner fits

An S-learner uses one function m(x,a) and computes m(x,1) minus m(x,0). Shared parameters help small arms, but heavy regularization may suppress the action feature and flatten real heterogeneity. A T-learner fits m1(x) only on treated rows and m0(x) only on controls. It can capture distinct response patterns, yet each arm has fewer examples and extrapolates where its own arm is sparse.

Keep the split outside both fits

Split by learner and decision time before tuning features, hyperparameters or segment boundaries. Fit preprocessing only on the training period. A categorical encoding learned from future cohorts can leak changes in the catalog or learner mix. Temporal validation protects the forecast setting; uplift evaluation checks whether model differences actually improve decisions.

Distinguish outcome fit from effect fit

Low log loss for both arm models does not establish accurate differences. If both models make the same 11-point upward error for one learner, their difference might survive; if the errors oppose each other, uplift can be badly wrong despite decent overall accuracy. Compare effect ranking, policy value, arm-specific calibration and segment support on held-out randomized assignments.

Choose complexity for the evidence

If treatment is rare or segments are thin, shrink toward a simpler common model and report uncertainty rather than claiming a precise per-learner effect. In observational data, separate arm regressions alone do not remove confounding. An X-learner can use imputed effects when arm sizes differ, while augmented weighting scores also use assignment probabilities.

Implementation

python
from statistics import mean

randomized_segment = [
    {"action": 1, "completed": 1}, {"action": 1, "completed": 1},
    {"action": 1, "completed": 0}, {"action": 1, "completed": 1},
    {"action": 0, "completed": 1}, {"action": 0, "completed": 0},
    {"action": 0, "completed": 0}, {"action": 0, "completed": 0},
]

def segment_uplift(assignments):
    treated = [row["completed"] for row in assignments if row["action"] == 1]
    control = [row["completed"] for row in assignments if row["action"] == 0]
    if not treated or not control:
        raise ValueError("both randomized arms are required")
    return mean(treated) - mean(control)

assert segment_uplift(randomized_segment) == 0.5

Performance and operating cost

The arithmetic segment baseline is O(N) time and O(N) temporary space. For a T-learner, fitting cost is the sum of two outcome-model fits; an S-learner fits once but predicts twice per learner. Cross-validation multiplies those costs. Arm-specific sample size, rather than total row count, limits credible heterogeneity.

Common Mistakes

  • Do not treat an outcome-model accuracy score as an uplift score.
  • Do not create segments after looking at test outcomes.
  • Do not fit a separate arm model where that arm has no local support.

Read next

ai-data
machine-learning
Storage details