An augmented propensity estimate combines a reward model for candidate actions with a propensity-weighted correction for observed reward errors.
Augmented propensity policy value
Build two separate components
For each context, predict reward under each eligible action using a reward model fitted without its evaluation outcome. Average those predictions under the candidate policy. Then add a correction: candidate probability divided by logged selection probability, multiplied by observed reward minus predicted reward for the logged action. The code handles a deterministic rapid-lane candidate and supplied reward estimates. IPS provides the correction term and its support requirement.
Read the name precisely
Under the usual contextual-bandit identification assumptions, the estimate can remain consistent if either the reward model or the logging propensity model is correctly specified. That does not mean it is immune to missing actions, incorrect reward definitions, changed outcomes, extreme weights or leakage. If logging propensities are exact by design, the reward model can reduce variance, but a small example cannot verify that advantage. Action logging is still foundational.
Keep nuisance fitting away from evaluation outcomes
Fit reward estimates on an earlier or cross-fitted partition; do not evaluate each row with a reward model trained on its own observed reward. Record the reward-model version and the features available at the original decision. A model that uses a post-routing signal introduces a different form of leakage even if its prediction error looks low. Pipeline isolation and feature timing both apply.
Diagnose rather than trust one number
Report the direct reward-model average, raw IPS, augmented estimate, largest weights and support by relevant context. Large disagreement among estimators deserves investigation. It is not a rule to pick the most favorable number. Repeated depot events and an adaptive logging policy complicate intervals; use an uncertainty method that respects those dependencies. Paired evaluation offers a starting point, not a blanket solution.
Protect the decision boundary
An offline estimate is a screening tool. Promotion requires operational eligibility, latency, review capacity, mature reward coverage and a bounded prospective plan when policy actions themselves affect future outcomes. The release project combines these conditions.
Implementation
# Logged action, logging probability, observed reward, predicted reward
# for standard and rapid actions from a separately fitted reward model.
records = [
("rapid", 0.50, 1.0, 0.2, 0.6),
("standard", 0.75, 0.0, 0.3, 0.7),
("rapid", 0.25, 0.0, 0.4, 0.5),
("standard", 0.50, 1.0, 0.5, 0.8),
]
def rapid_policy_dr(rows):
terms = []
for action, logged_probability, reward, standard_estimate, rapid_estimate in rows:
if action not in {"standard", "rapid"} or not 0 < logged_probability <= 1:
raise ValueError("invalid logged action")
chosen_estimate = rapid_estimate if action == "rapid" else standard_estimate
correction = ((reward - chosen_estimate) / logged_probability
if action == "rapid" else 0.0)
terms.append(rapid_estimate + correction)
return sum(terms) / len(terms)
estimate = rapid_policy_dr(records)
assert round(estimate, 3) == 0.35Performance and operating cost
Scoring N rows costs O(N) time and O(1) running memory when streamed, beyond fitting the reward model. That fit must be repeated across suitable folds or historical partitions to avoid evaluation leakage. Extreme inverse probabilities can still create high variance even when the reward model is useful.
Common Mistakes
- Do not fit and evaluate the reward model on the same row without a valid split.
- Do not read the two-part construction as protection against missing support or bad reward measurement.
- Do not select the largest estimate among methods after seeing results.
