Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Exploration budget and action guardrails

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An exploration policy assigns nonzero probability to eligible alternatives while respecting hard action constraints and a declared intervention budget.

Start with the eligible set

At a depot, some shipments cannot use the rapid lane because capacity or handling rules forbid it. Filter those actions before randomization. The code uses an epsilon-greedy choice over the resulting action set and returns the exact probability of the chosen action. No exploration probability can make an ineligible action safe. The log schema must retain this action set and the reason it changed.

Know the actual probability

For K eligible actions, epsilon-greedy selects the currently preferred action with probability one minus epsilon plus epsilon divided by K, and each other action with epsilon divided by K. Record that probability after all guardrails, throttles and fallbacks have executed. If operators override a draw or a budget cap reroutes it, the pre-override probability is not the probability of the action delivered. IPS fails when that denominator is wrong.

Budget learning separately from service risk

A five-percent exploration budget is not permission to risk five percent of critical shipments. Define which contexts permit experimentation, maximum rapid-lane volume, per-site load and stop conditions. Compare reward and harm, not just average clicks or short-term speed. For high-stakes decisions, a randomized policy may be inappropriate without a formal review. Group error and review capacity are operating constraints.

Distinguish learning from evaluation

The same adaptive log can support learning, but candidate selection and evaluation should be separated by time or other valid design. If exploration probabilities shrink to nearly zero for alternatives, counterfactual estimates become unstable. Reserve a controlled, policy-approved evaluation cohort with enough support rather than trying to rescue a deterministic history by arithmetic. Augmented propensity evaluation cannot create data for absent actions.

Account for delayed outcomes

A policy update based only on quickly observed wins may prefer actions with short reward latency. Freeze a maturity rule, monitor pending outcomes and keep a rollback path when changing exploration. The reward guide separates pending cases from failures.

Implementation

python
import random

def choose_lane(eligible_actions, preferred_action, epsilon, rng):
    if not eligible_actions or preferred_action not in eligible_actions:
        raise ValueError("preferred action must be eligible")
    if not 0 <= epsilon <= 1:
        raise ValueError("epsilon outside probability range")
    action_count = len(eligible_actions)
    chosen = (rng.choice(eligible_actions) if rng.random() < epsilon
              else preferred_action)
    propensity = (1 - epsilon if chosen == preferred_action else 0)
    propensity += epsilon / action_count
    return chosen, propensity

local_rng = random.Random(47)
eligible = ("standard", "rapid")
decision = choose_lane(eligible, "standard", 0.20, local_rng)
assert decision[0] in eligible
assert decision[1] in {0.10, 0.90}
forced_standard = choose_lane(("standard",), "standard", 0.20, local_rng)
assert forced_standard == ("standard", 1.0)

Performance and operating cost

Selecting among K eligible actions costs O(K) if constructing and sampling a list, with O(K) action-set storage; the simple code receives that list. Policy logs and safeguards must scale with decisions. More exploration can improve support but consumes real operational capacity and may impose measurable harm.

Common Mistakes

  • Do not explore actions forbidden by operational rules.
  • Do not log the pre-override probability as if it generated the final action.
  • Do not treat a nominal exploration percentage as an acceptable harm budget.

Read next

Continue the workflow: Q-learning control and exploration.

ai-data
machine-learning
Storage details