Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Q-learning control and exploration

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Tabular Q-learning updates a state-action value toward observed reward plus the best estimated next-action value, while its behavior policy must still collect eligible alternatives.

Separate behavior from update target

A warehouse learner may sometimes choose a less-preferred eligible action to obtain evidence. The behavior policy generates the transition. The Q-learning target uses the largest estimated next-action value, not necessarily the next action the behavior policy would take. That distinction makes this an off-policy control update. Exploration guardrails still govern what actions can actually be tried.

Handle terminal and eligible actions

The code applies one update to a nonterminal low-stock transition and another to a terminal shortage transition. It looks up the maximum only among legal next actions and uses zero continuation at a true terminal. A value for a forbidden action must never be selected by a maximization routine. The action map defines this legal set.

Do not mistake a toy update for a deployment policy

Convergence statements for tabular Q-learning require conditions on exploration and learning rates in a stationary finite process. An adaptive warehouse with delayed rewards, shared constraints and incomplete state does not satisfy them by declaration. Audit coverage and compare a fixed safe policy on independent later episodes. Offline support is a separate release concern.

Inspect learning dynamics

Report visits by state and action, update size, realized long-run return, shortage rate and capacity incidents. A high estimated Q value for rarely tried replenishment may reflect noise or an optimistic initialization. A small maximum update is not enough to prove the system objective is correct. Return and TD clarify the bootstrapped target.

Keep exploration within operations

An epsilon-greedy or other behavior rule needs a logged selection probability after eligibility and overrides. Never explore a prohibited action merely to populate the Q table. If safe exploration is not possible, rely on a reviewed simulator or conservative policy instead of claiming new action effects from absent data. The project asks for a bounded pilot and rollback.

Implementation

python
eligible = {"low-stock": ("replenish", "wait"),
            "ready-stock": ("dispatch", "hold"), "closed": ()}
q_values = {("low-stock", "replenish"): 1.0,
            ("low-stock", "wait"): -2.0,
            ("ready-stock", "dispatch"): 4.0,
            ("ready-stock", "hold"): 0.0}

def q_update(table, state, action, reward, next_state, terminated,
             legal_actions, rate=0.5, gamma=0.8):
    if action not in legal_actions[state]:
        raise ValueError("ineligible action")
    continuation = (0.0 if terminated else
                    max(table[(next_state, next_action)]
                        for next_action in legal_actions[next_state]))
    previous = table[(state, action)]
    table[(state, action)] = previous + rate * (
        reward + gamma * continuation - previous)
    return table[(state, action)]

first = q_update(q_values, "low-stock", "replenish", -2.0,
                 "ready-stock", False, eligible)
terminal = q_update(q_values, "low-stock", "wait", -7.0,
                    "closed", True, eligible)
assert abs(first - 1.1) < 1e-9
assert terminal == -4.5

Performance and operating cost

One tabular update scans at most A eligible next actions, costing O(A) time and O(SA) stored values for S states and A actions. Collecting enough safe transitions is harder than table arithmetic. Function approximators add training cost and can amplify errors outside the visited state-action distribution.

Common Mistakes

  • Do not bootstrap from a terminal state.
  • Do not maximize over forbidden or unobserved actions without checking support.
  • Do not infer a good operational policy from a few stable-looking table updates.

Read next

ai-data
machine-learning
Storage details