An exploration policy assigns nonzero probability to eligible alternatives while respecting hard action constraints and a declared intervention budget.
Exploration budget and action guardrails
Start with the eligible set
At a depot, some shipments cannot use the rapid lane because capacity or handling rules forbid it. Filter those actions before randomization. The code uses an epsilon-greedy choice over the resulting action set and returns the exact probability of the chosen action. No exploration probability can make an ineligible action safe. The log schema must retain this action set and the reason it changed.
Know the actual probability
For K eligible actions, epsilon-greedy selects the currently preferred action with probability one minus epsilon plus epsilon divided by K, and each other action with epsilon divided by K. Record that probability after all guardrails, throttles and fallbacks have executed. If operators override a draw or a budget cap reroutes it, the pre-override probability is not the probability of the action delivered. IPS fails when that denominator is wrong.
Budget learning separately from service risk
A five-percent exploration budget is not permission to risk five percent of critical shipments. Define which contexts permit experimentation, maximum rapid-lane volume, per-site load and stop conditions. Compare reward and harm, not just average clicks or short-term speed. For high-stakes decisions, a randomized policy may be inappropriate without a formal review. Group error and review capacity are operating constraints.
Distinguish learning from evaluation
The same adaptive log can support learning, but candidate selection and evaluation should be separated by time or other valid design. If exploration probabilities shrink to nearly zero for alternatives, counterfactual estimates become unstable. Reserve a controlled, policy-approved evaluation cohort with enough support rather than trying to rescue a deterministic history by arithmetic. Augmented propensity evaluation cannot create data for absent actions.
Account for delayed outcomes
A policy update based only on quickly observed wins may prefer actions with short reward latency. Freeze a maturity rule, monitor pending outcomes and keep a rollback path when changing exploration. The reward guide separates pending cases from failures.
Implementation
import random
def choose_lane(eligible_actions, preferred_action, epsilon, rng):
if not eligible_actions or preferred_action not in eligible_actions:
raise ValueError("preferred action must be eligible")
if not 0 <= epsilon <= 1:
raise ValueError("epsilon outside probability range")
action_count = len(eligible_actions)
chosen = (rng.choice(eligible_actions) if rng.random() < epsilon
else preferred_action)
propensity = (1 - epsilon if chosen == preferred_action else 0)
propensity += epsilon / action_count
return chosen, propensity
local_rng = random.Random(47)
eligible = ("standard", "rapid")
decision = choose_lane(eligible, "standard", 0.20, local_rng)
assert decision[0] in eligible
assert decision[1] in {0.10, 0.90}
forced_standard = choose_lane(("standard",), "standard", 0.20, local_rng)
assert forced_standard == ("standard", 1.0)Performance and operating cost
Selecting among K eligible actions costs O(K) if constructing and sampling a list, with O(K) action-set storage; the simple code receives that list. Policy logs and safeguards must scale with decisions. More exploration can improve support but consumes real operational capacity and may impose measurable harm.
Common Mistakes
- Do not explore actions forbidden by operational rules.
- Do not log the pre-override probability as if it generated the final action.
- Do not treat a nominal exploration percentage as an acceptable harm budget.
Read next
- Contextual bandit action logging and support
- Inverse-propensity policy value
- Delayed bandit rewards and outcome maturity
- Contextual bandit policy release review project
Continue the workflow: Q-learning control and exploration.
