Inverse-propensity scoring estimates the reward of a candidate action policy by weighting observed rewards according to candidate and logging action probabilities.
Inverse-propensity policy value
Evaluate a declared target policy
A proposed routing policy gives each eligible action a probability in each recorded context. For each logged decision, divide its candidate probability for the action actually taken by the probability that the logging policy assigned that action. Multiply the observed reward by this ratio, then average over all decisions. A deterministic candidate gives zero weight to rows where the logging action differs. The action log supplies the denominator.
Understand the assumptions
The logging probability must describe the actual randomized selection after eligibility constraints, and the observed reward must be attached to the chosen action. Candidate actions need positive logging support. Context and action versions must be consistent. If operators secretly override assigned actions in a way the log omits, the recorded probability is wrong. The estimator then need not identify the proposed policy value. Guardrails must be visible in the log.
Report weight concentration
Rarely selected actions receive large weights. The code reports raw IPS, a self-normalized ratio and effective sample size. Self-normalization can reduce instability but generally changes the finite-sample bias; it is a diagnostic companion, not a silent substitute for raw IPS. Display the largest weights and their contributing cases. A handful of rewarded shipments must not be presented as a certain operational gain. The augmented propensity estimate is another comparison, with its own assumptions.
Keep policy selection separate from evaluation
If many candidate policies are tried and the one with the largest IPS estimate is reported on those same cases, selection optimism returns. Use historical development logs to design the candidate, then a later untouched logged cohort for evaluation. Repeated decisions for the same depot or customer may require clustered uncertainty. The time split preserves that boundary.
Check reward definition before arithmetic
A rapid-lane reward might combine timely handoff with staffing cost. Declare units, observation window and missing-reward rule before comparing policies. Off-policy estimators cannot fix a reward that omits the downstream cost the business actually cares about. Reward maturity covers unresolved outcomes; the release review checks the full policy.
Implementation
# Selected action, logging probability, candidate probability, mature reward.
logged_decisions = [
("rapid", 0.25, 1.0, 1.0),
("standard", 0.75, 0.0, 0.0),
("rapid", 0.50, 1.0, 0.0),
("standard", 0.50, 0.0, 1.0),
]
def policy_value(rows):
if not rows:
raise ValueError("no mature logged decisions")
weighted_rewards = 0.0
weights = []
for action, logging_probability, candidate_probability, reward in rows:
if action not in {"rapid", "standard"}:
raise ValueError("unknown action")
if not 0 < logging_probability <= 1 or not 0 <= candidate_probability <= 1:
raise ValueError("invalid policy probability")
weight = candidate_probability / logging_probability
weighted_rewards += weight * reward
weights.append(weight)
raw_ips = weighted_rewards / len(rows)
normalized = weighted_rewards / sum(weights) if sum(weights) else None
effective_size = sum(weights) ** 2 / sum(weight ** 2 for weight in weights)
return raw_ips, normalized, effective_size
ips, normalized, effective_size = policy_value(logged_decisions)
assert ips == 1.0
assert normalized == 2 / 3
assert effective_size == 1.8Performance and operating cost
For N mature decisions, scoring takes O(N) time and O(1) running memory; the teaching implementation keeps O(N) weights for inspection. Effective sample size can be far smaller than N when probabilities are small. Logging, reward maturity and credible uncertainty estimates usually dominate compute.
Common Mistakes
- Do not divide by a policy score that was not the actual selection probability.
- Do not compare a candidate action with zero logging support.
- Do not report a self-normalized estimate without naming its changed bias behavior.
Read next
- Contextual bandit action logging and support
- Augmented propensity policy value
- Exploration budget and action guardrails
- Contextual bandit policy release review project
Continue the workflow: Uplift ranking and offline policy-value evaluation.
