Skip to content
AITroveRead. Build. Understand.
Make this comfortable

DQN transitions, replay sampling and reward contracts

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A value network cannot repair a faulty decision log: state timing, valid actions, terminal flags and reward accounting determine what replay teaches.

Define one decision boundary

A warehouse replenishment agent might choose whether to move a spare carton before the next dispatch. Store the observation available immediately before that choice, the selected action, the reward earned over the following interval, the next observation and whether the episode ended. A later stock count cannot appear in the earlier state. Attach event time, policy version and action-eligibility rules to the transition, then test whether the same record can be reconstructed from the event ledger. Causal feature construction prevents a learned shortcut from reading the future.

Make reward units explicit

A positive reward for completed orders and a negative reward for late orders sounds simple until a single late order gets counted on every subsequent time step. Define when each outcome is booked and what a terminal penalty means. Keep physical units alongside the scalar reward so an engineer can audit one episode by hand. If the organization changes the cost of an expedited move, version the reward function and do not mix labels from two versions in one unexplained training run.

Separate termination from truncation

A terminal state means there is no next decision in the modeled episode; a log cut because the shift ended may merely truncate observation. These cases can require different bootstrap treatment. The code models a terminal flag and refuses sampling before the replay buffer contains enough distinct transitions. Retain a stable episode ID to identify adjacent records. The target lesson uses the flag to suppress bootstrap on true terminal transitions.

Sample without hiding coverage

Uniform replay breaks the tight correlation of neighboring observations and reuses costly experience, but it does not invent actions the logger never tried. Record counts by warehouse, action, stock band and failure mode. A buffer dominated by quiet hours can underrepresent rare stockouts; stratified analysis is still needed even if mini-batches are random. Split evaluation by episode or time period before filling training replay, since shared adjacent transitions across train and test inflate performance.

Keep exploration out of unsafe decisions

Epsilon-greedy exploration can choose a random action only from the permitted set. An action mask belongs to the state contract and must be applied both while collecting decisions and while estimating future values. In an operational setting, begin with simulation or a tightly bounded shadow policy and define a rollback to the existing dispatcher. The project checks action legality, order completion and stockout cost together.

Implementation

python
from collections import deque
from dataclasses import dataclass
from random import Random

@dataclass(frozen=True)
class DispatchTransition:
    stock_before: tuple[int, int]
    action: int
    reward: float
    stock_after: tuple[int, int]
    terminal: bool
    episode_id: str

replay = deque(maxlen=47)
replay.extend([
    DispatchTransition((8, 3), 1, 2.0, (7, 4), False, "shift-47"),
    DispatchTransition((7, 4), 0, -1.0, (6, 4), True, "shift-47"),
    DispatchTransition((5, 2), 1, 3.0, (4, 3), True, "shift-48"),
])
randomizer = Random(47)
batch = randomizer.sample(list(replay), k=2)
assert len(batch) == 2 and len(set(batch)) == 2
assert all(record.episode_id and record.action in (0, 1) for record in batch)

Performance and operating cost

Appending to a bounded deque is O(1) amortized time and O(C) memory for capacity C. Converting the entire deque to a list for each sample is O(C); a large production replay store needs indexed sampling or batched storage. One sampled mini-batch is O(B) after indexing for batch size B. Environment interaction may dominate training cost, while uniform random replay still leaves distribution mismatch between historical behavior and a proposed policy.

Common Mistakes

  • Do not place a post-decision count in the pre-decision state.
  • Do not treat a log cutoff as a true terminal event without a declared rule.
  • Do not claim replay covers actions that the logging policy never selected.

Read next

Continue the workflow: PPO rollouts, advantage estimates and bootstrap boundaries.

ai-data
deep-learning
Storage details