Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Contextual bandit action logging and support

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A contextual bandit records the context, available actions, selected action, selection probability and eventual reward for each decision.

Define one decision at a time

A warehouse can either send a shipment to a rapid review lane or keep the standard lane. At intake, the routing system sees backlog and service tier, chooses one action, and later observes the outcome under that action. The outcome for the action not taken is missing. This is a different data shape from a fully labeled classifier: the logged reward belongs only to the selected action. Decision costing explains why the action, rather than a score alone, is the object of evaluation.

Record the probability that produced the action

A logging policy can choose rapid review with probability 0.25 in one context and 0.60 in another. Save the probability assigned to the action actually taken at decision time. Recomputing it from a later policy version loses the information needed for inverse-propensity evaluation. Log the candidate action set, policy version, feature snapshot, action time and reward-ready time with an immutable decision ID. Feature timing prevents post-decision values from entering context.

Check support before counterfactual analysis

A proposed policy must only choose actions that the logging policy had a nonzero chance of taking in matching contexts. If rapid review was forbidden for premium shipments, no logged record can identify its effect there. A small nonzero chance may technically provide support but yield a high-variance estimate. The code checks probabilities and reports the minimum observed selection probability; it cannot prove support for every unobserved context. The IPS lesson makes the consequence explicit.

Keep eligibility separate from preference

Some shipments are legally or operationally ineligible for rapid review. The candidate policy must obey the same action constraints; an exploration scheme must never override them. Record why an action was unavailable. Otherwise a missing action might be mistaken for a random draw and a candidate could be evaluated on impossible decisions. Exploration controls make the action set explicit.

Treat delayed rewards as data, not zeros

The handoff outcome may arrive after dispatch or be corrected later. Join it by decision ID and apply a declared maturity window. An incomplete reward is not a failed shipment. Record all decisions, including those still pending, so evaluation can expose selective observation. Reward maturity handles that clock.

Implementation

python
# Decision ID, context, eligible actions, selected action, selection probability.
decision_log = [
    ("R801", "high-backlog", ("standard", "rapid"), "rapid", 0.35),
    ("R802", "low-backlog", ("standard", "rapid"), "standard", 0.80),
    ("R803", "premium", ("standard",), "standard", 1.00),
]

def validate_action_log(rows):
    seen_ids = set()
    lowest_probability = 1.0
    for decision_id, context, eligible, action, action_probability in rows:
        if decision_id in seen_ids or not context or action not in eligible:
            raise ValueError("invalid or duplicate decision")
        if not 0 < action_probability <= 1:
            raise ValueError("invalid logged action probability")
        seen_ids.add(decision_id)
        lowest_probability = min(lowest_probability, action_probability)
    return {"decisions": len(seen_ids), "min_logged_probability": lowest_probability}

summary = validate_action_log(decision_log)
assert summary == {"decisions": 3, "min_logged_probability": 0.35}

Performance and operating cost

Validating N decision records costs O(N) time and O(N) IDs for duplicate detection. Logging one probability is cheap; storing immutable feature snapshots, policy versions, eligibility rules and later rewards is the substantial systems work. No estimator can reconstruct a missing selection probability reliably after the policy changes.

Common Mistakes

  • Do not attach the reward of the chosen action to every available action.
  • Do not infer a historical selection probability from the current policy.
  • Do not treat an ineligible action as merely an unchosen eligible action.

Read next

Continue the workflow: Inverse-propensity policy value.

Continue the workflow: Bellman value iteration for stock control.

Continue the workflow: Click position bias and support for ranking evaluation.

ai-data
machine-learning
Storage details