Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Bandit decision logs: preserve action probabilities and delayed rewards

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A bandit decision log needs the chosen action, its selection probability, context revision and eventual reward to support evaluation.

Log the decision before its result

A delivery app chooses one of three pickup-notification timings for each eligible order. The policy sees a context, chooses an action and later observes completion only for that action. Store decision ID, eligible action set, chosen action, probability assigned by the logging policy, policy revision, context feature revision and event time before sending the notification. A model score is not the probability that the logging policy chose the action. Experiment guardrails still protect user cohorts, but a bandit needs the extra probability trail.

Join delayed rewards without inventing counterfactuals

The reward might arrive after a pickup window closes. Join by decision ID, set a maturity cutoff and retain missing or corrected outcomes as separate states. Do not treat an unobserved completion for an unchosen timing as a failure; that timing was never sent. Preserve the reward definition and revision, such as completed pickup within a fixed window, because changing it alters every policy comparison. Outcome maturity gives the same discipline to supervised decisions.

Check support and integrity

For any action a candidate policy might choose in a recorded context, the logging policy needs a nonzero probability of that action if historical logs are to evaluate it with inverse-propensity weighting. A deterministic logger with no exploration leaves other actions unsupported. Validate that logged probability is finite, positive and no greater than one, that the chosen action was eligible and that a decision ID appears once. Quarantine broken records instead of silently replacing missing probability with one. Offline policy gates reject unsupported comparisons.

Keep rollout evidence separable

During a policy change, log the logging policy revision on every decision rather than pooling all events as though one policy served them. Track action mix, reward maturity, cancellation outcomes and notification burden by cohort. The policy may optimize completion but annoy users with excess alerts, so the release needs service guardrails as well as reward. The project catches a candidate that shifts heavily toward an action the historical logger almost never tried.

Implementation

python
from math import isfinite

def validate_bandit_decision(decision):
    probability = decision["logged_probability"]
    if not isfinite(probability) or not 0 < probability <= 1:
        return "reject:probability"
    if decision["chosen_action"] not in decision["eligible_actions"]:
        return "reject:ineligible-action"
    if not decision["decision_id"] or not decision["policy_revision"]:
        return "reject:identity"
    return "admit"

record = {"decision_id": "pickup-47", "policy_revision": "timing-r8",
          "eligible_actions": {"early", "middle", "late"},
          "chosen_action": "middle", "logged_probability": 0.27}
assert validate_bandit_decision(record) == "admit"
assert validate_bandit_decision({**record, "logged_probability": 0})        == "reject:probability"

Performance and operating cost

Validation is O(1) expected time and space for a small eligible-action set. Storing decision probabilities and reward joins grows linearly with decisions. Exploration can reduce short-term reward, so the policy budget and user-impact guardrails must be explicit before logs are collected.

Common Mistakes

  • Recording a model score instead of the probability of the chosen action.
  • Filling an unchosen action with an invented negative reward.
  • Mixing decisions from multiple logging policies without their revisions.
  • Assuming a deterministic policy generated support for an untried action.

Read next

ai-data
mlops
Storage details