An offline policy estimate needs known logging probabilities, action support and enough effective observations to be useful.
Bandit offline evaluation: support, variance and release limits
State the estimand
Historical pickup decisions reveal reward only for the notification timing actually sent. To estimate a candidate policy from those logs, weight each observed reward by candidate probability for the logged action divided by the logging probability for that action. Average across eligible mature decisions. This inverse-propensity estimate targets the candidate policy under the logged context distribution when probabilities are correct and support exists; it is not a controlled live experiment. The decision ledger supplies the denominator.
Reject absent support
If the candidate places weight on late notifications in a context where the logger always chose early, there is no observed late reward for that context. The estimate cannot rescue it by dividing by zero, dropping those contexts or guessing their reward. Mark the candidate unsupported and gather controlled exploration or narrow the candidate to supported actions. Even a tiny nonzero logging probability creates very large weights and a noisy estimate. Report the weight distribution and effective sample size, not just the estimated reward.
Keep reward and cohort definitions fixed
Use mature rewards, identical eligibility filters and a pinned reward revision for incumbent and candidate. A delivery change that speeds pickups but raises cancellations may score well on the wrong reward. Evaluate safety cohorts and notification burden separately. The result can be distorted if logged action probabilities were wrong, if users interfere with one another or if new production contexts differ from the logged frame. Interference controls and observation coverage show why an offline score has limits.
Promote through a bounded experiment
Require valid logs, target support, a minimum effective sample size, acceptable cohort bounds and a fixed rollback rule before a canary. The offline estimate is a filter for weak candidates, not a promise of live lift. A small online randomized allocation can then measure actual outcomes under the new policy, with mature rewards and guardrails. The project rejects a high estimated reward driven by one rare, heavily weighted action.
Implementation
def ips_estimate(decisions):
if not decisions:
return {"state": "hold:no-decisions"}
weights = []
weighted_rewards = []
for decision in decisions:
for action, target_probability in decision["target_probs"].items():
if target_probability > 0 and decision["behavior_probs"].get(action, 0) <= 0:
return {"state": "hold:unsupported-action"}
chosen = decision["chosen_action"]
weight = decision["target_probs"].get(chosen, 0) / decision["behavior_probs"][chosen]
weights.append(weight)
weighted_rewards.append(weight * decision["reward"])
weight_square_sum = sum(weight * weight for weight in weights)
effective_n = sum(weights) ** 2 / weight_square_sum if weight_square_sum else 0
return {"state": "measured", "value": sum(weighted_rewards) / len(decisions),
"effective_n": effective_n}
logs = [{"chosen_action": "early", "behavior_probs": {"early": 0.5,
"late": 0.5}, "target_probs": {"early": 0.5, "late": 0.5},
"reward": 1},
{"chosen_action": "late", "behavior_probs": {"early": 0.5,
"late": 0.5}, "target_probs": {"early": 0.5, "late": 0.5},
"reward": 0}]
assert ips_estimate(logs)["value"] == 0.5
assert ips_estimate(logs)["effective_n"] == 2
Performance and operating cost
The estimate is O(n × a) time for n decisions and a candidate actions per context, with O(n) temporary weight storage. Heavy weights reduce effective sample size even when raw n is large. A stricter support gate may delay rollout, but replacing missing evidence with a guessed reward would make the estimate look precise without data.
Common Mistakes
- Using a chosen-action probability of zero or an invented default.
- Reporting raw sample count without weight concentration.
- Treating an offline estimate as evidence of live causal lift.
- Discarding unsupported contexts until the remaining estimate looks favorable.
Read next
- Bandit decision logs: preserve action probabilities and delayed rewards
- Project: evaluate a pickup-notification bandit before rollout
- Model experiment guardrails: stop harm without misreading the sample
- Selective labels: measure what the model never lets reviewers see
- Feedback policy shift: compare models when labels depend on routing
