A bandit decision log needs the chosen action, its selection probability, context revision and eventual reward to support evaluation.
Bandit decision logs: preserve action probabilities and delayed rewards
Log the decision before its result
A delivery app chooses one of three pickup-notification timings for each eligible order. The policy sees a context, chooses an action and later observes completion only for that action. Store decision ID, eligible action set, chosen action, probability assigned by the logging policy, policy revision, context feature revision and event time before sending the notification. A model score is not the probability that the logging policy chose the action. Experiment guardrails still protect user cohorts, but a bandit needs the extra probability trail.
Join delayed rewards without inventing counterfactuals
The reward might arrive after a pickup window closes. Join by decision ID, set a maturity cutoff and retain missing or corrected outcomes as separate states. Do not treat an unobserved completion for an unchosen timing as a failure; that timing was never sent. Preserve the reward definition and revision, such as completed pickup within a fixed window, because changing it alters every policy comparison. Outcome maturity gives the same discipline to supervised decisions.
Check support and integrity
For any action a candidate policy might choose in a recorded context, the logging policy needs a nonzero probability of that action if historical logs are to evaluate it with inverse-propensity weighting. A deterministic logger with no exploration leaves other actions unsupported. Validate that logged probability is finite, positive and no greater than one, that the chosen action was eligible and that a decision ID appears once. Quarantine broken records instead of silently replacing missing probability with one. Offline policy gates reject unsupported comparisons.
Keep rollout evidence separable
During a policy change, log the logging policy revision on every decision rather than pooling all events as though one policy served them. Track action mix, reward maturity, cancellation outcomes and notification burden by cohort. The policy may optimize completion but annoy users with excess alerts, so the release needs service guardrails as well as reward. The project catches a candidate that shifts heavily toward an action the historical logger almost never tried.
Implementation
from math import isfinite
def validate_bandit_decision(decision):
probability = decision["logged_probability"]
if not isfinite(probability) or not 0 < probability <= 1:
return "reject:probability"
if decision["chosen_action"] not in decision["eligible_actions"]:
return "reject:ineligible-action"
if not decision["decision_id"] or not decision["policy_revision"]:
return "reject:identity"
return "admit"
record = {"decision_id": "pickup-47", "policy_revision": "timing-r8",
"eligible_actions": {"early", "middle", "late"},
"chosen_action": "middle", "logged_probability": 0.27}
assert validate_bandit_decision(record) == "admit"
assert validate_bandit_decision({**record, "logged_probability": 0}) == "reject:probability"
Performance and operating cost
Validation is O(1) expected time and space for a small eligible-action set. Storing decision probabilities and reward joins grows linearly with decisions. Exploration can reduce short-term reward, so the policy budget and user-impact guardrails must be explicit before logs are collected.
Common Mistakes
- Recording a model score instead of the probability of the chosen action.
- Filling an unchosen action with an invented negative reward.
- Mixing decisions from multiple logging policies without their revisions.
- Assuming a deterministic policy generated support for an untried action.
Read next
- Bandit offline evaluation: support, variance and release limits
- Project: evaluate a pickup-notification bandit before rollout
- Model experiment guardrails: stop harm without misreading the sample
- Prediction-outcome joins: evaluate only mature, matched decisions
- Feedback policy shift: compare models when labels depend on routing
