Tabular Q-learning updates a state-action value toward observed reward plus the best estimated next-action value, while its behavior policy must still collect eligible alternatives.
Q-learning control and exploration
Separate behavior from update target
A warehouse learner may sometimes choose a less-preferred eligible action to obtain evidence. The behavior policy generates the transition. The Q-learning target uses the largest estimated next-action value, not necessarily the next action the behavior policy would take. That distinction makes this an off-policy control update. Exploration guardrails still govern what actions can actually be tried.
Handle terminal and eligible actions
The code applies one update to a nonterminal low-stock transition and another to a terminal shortage transition. It looks up the maximum only among legal next actions and uses zero continuation at a true terminal. A value for a forbidden action must never be selected by a maximization routine. The action map defines this legal set.
Do not mistake a toy update for a deployment policy
Convergence statements for tabular Q-learning require conditions on exploration and learning rates in a stationary finite process. An adaptive warehouse with delayed rewards, shared constraints and incomplete state does not satisfy them by declaration. Audit coverage and compare a fixed safe policy on independent later episodes. Offline support is a separate release concern.
Inspect learning dynamics
Report visits by state and action, update size, realized long-run return, shortage rate and capacity incidents. A high estimated Q value for rarely tried replenishment may reflect noise or an optimistic initialization. A small maximum update is not enough to prove the system objective is correct. Return and TD clarify the bootstrapped target.
Keep exploration within operations
An epsilon-greedy or other behavior rule needs a logged selection probability after eligibility and overrides. Never explore a prohibited action merely to populate the Q table. If safe exploration is not possible, rely on a reviewed simulator or conservative policy instead of claiming new action effects from absent data. The project asks for a bounded pilot and rollback.
Implementation
eligible = {"low-stock": ("replenish", "wait"),
"ready-stock": ("dispatch", "hold"), "closed": ()}
q_values = {("low-stock", "replenish"): 1.0,
("low-stock", "wait"): -2.0,
("ready-stock", "dispatch"): 4.0,
("ready-stock", "hold"): 0.0}
def q_update(table, state, action, reward, next_state, terminated,
legal_actions, rate=0.5, gamma=0.8):
if action not in legal_actions[state]:
raise ValueError("ineligible action")
continuation = (0.0 if terminated else
max(table[(next_state, next_action)]
for next_action in legal_actions[next_state]))
previous = table[(state, action)]
table[(state, action)] = previous + rate * (
reward + gamma * continuation - previous)
return table[(state, action)]
first = q_update(q_values, "low-stock", "replenish", -2.0,
"ready-stock", False, eligible)
terminal = q_update(q_values, "low-stock", "wait", -7.0,
"closed", True, eligible)
assert abs(first - 1.1) < 1e-9
assert terminal == -4.5Performance and operating cost
One tabular update scans at most A eligible next actions, costing O(A) time and O(SA) stored values for S states and A actions. Collecting enough safe transitions is harder than table arithmetic. Function approximators add training cost and can amplify errors outside the visited state-action distribution.
Common Mistakes
- Do not bootstrap from a terminal state.
- Do not maximize over forbidden or unobserved actions without checking support.
- Do not infer a good operational policy from a few stable-looking table updates.
