Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Uplift ranking and offline policy-value evaluation

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A useful uplift model sends the action to people whose outcomes change because of it; outcome accuracy alone cannot establish that ranking.

Freeze scores before the test

Train an effect model on earlier decisions, then score a later randomized test cohort using only pre-assignment features. Sort eligible learners by predicted uplift and predeclare capacity bands, such as the top 15% and 30%. For each band, estimate treated-minus-control completion with assignment probabilities accounted for. Plotting an uplift curve after changing bands to flatter the model turns the test into training.

Evaluate the policy that will run

Define a policy pi(x) returning reminder or control. Under a randomized logging policy with known probability for each action, inverse-propensity policy value averages observed outcome divided by the probability of the logged action only when pi(x) matches that action. This estimates expected outcome under the new policy if consistency, support, and stable outcome measurement hold. Inverse-propensity policy value develops the estimator.

Compare incremental value, not just value

A policy value of 0.41 means little without a comparable baseline. Report value minus the no-reminder policy on the same test population, plus reminder cost and operational capacity. For 1,000 eligible learners, estimated uplift of 0.04 among 120 targeted learners implies about 4.8 additional completions in expectation, not 40 across the full cohort. A small benefit may disappear under uncertainty or message fatigue.

Understand metric failure modes

AUC or accuracy on completion may reward finding learners who would finish anyway. A naive uplift curve can be distorted by uneven treatment rates across score bands; use randomized arms or appropriate weighting. Ranking each positive example against a tiny sampled set of negatives also changes the question. Check assignment balance and held-out arm outcome rates before presenting any curve.

Make uncertainty operational

Use learner-level resampling or a suitable influence-function approach for confidence intervals; keep test choices frozen. A high-variance inverse-weighted estimate may require a larger randomized test or an outcome-model-assisted estimate. Augmented weighting can improve stability when its nuisance models behave, but it cannot replace randomized support.

Implementation

python
randomized_holdout = [
    {"segment": "high", "action": 1, "outcome": 1, "logged_probability": 0.5},
    {"segment": "high", "action": 0, "outcome": 0, "logged_probability": 0.5},
    {"segment": "low", "action": 1, "outcome": 0, "logged_probability": 0.5},
    {"segment": "low", "action": 0, "outcome": 1, "logged_probability": 0.5},
]

def inverse_propensity_policy_value(records, policy):
    if not records:
        raise ValueError("empty evaluation cohort")
    weighted_outcome = 0.0
    for record in records:
        if not 0 < record["logged_probability"] <= 1:
            raise ValueError("invalid probability of logged action")
        if policy(record) == record["action"]:
            weighted_outcome += record["outcome"] / record["logged_probability"]
    return weighted_outcome / len(records)

target_high = lambda record: int(record["segment"] == "high")
assert inverse_propensity_policy_value(randomized_holdout, target_high) == 1.0

Performance and operating cost

Policy scoring is O(N) for N test decisions; sorting all scores is O(N log N), or O(N log K) with a heap for a top-K capacity. Inverse weighting can have high variance when the new policy selects rarely logged actions. Keep the test cohort large enough for the intended decision and its uncertainty.

Common Mistakes

  • Do not use completion AUC as the sole uplift metric.
  • Do not select the best-looking capacity band on the final test.
  • Do not divide by the count of matching actions; the inverse-weighted denominator is the full eligible cohort.

Read next

ai-data
machine-learning
Storage details