An actor-critic update begins with fresh policy data; its advantage calculation must distinguish true endings from collection cutoffs.
PPO rollouts, advantage estimates and bootstrap boundaries
Collect one policy version
A warehouse simulator should record the observation available before dispatch, the chosen action, its old log probability, reward, value estimate, next observation and action mask. Keep the policy and reward versions with each rollout. PPO may optimize several times against this fixed batch, but newly collected data belongs to a new policy version. Mixing the two makes probability ratios meaningless. The transition lesson covers observation timing shared with DQN.
Separate endings from cutoffs
A true terminal shift cannot bootstrap a future value. A collector may stop after a fixed number of steps while the shift continues; that cutoff can use the critic value of the next observation. Both cases stop advantage recursion before an unrelated trajectory. The code stores terminal and cutoff separately. Treating every cutoff as terminal biases returns downward; carrying a recurrence into the next shift mixes unrelated outcomes.
Work backward through rewards
At each step compute reward plus discounted next value when nonterminal, minus the current value. Propagate later errors backward with discount and trace decay until an episode boundary. This produces an advantage estimate while allowing control over variance and bias. Save the critic predictions made during collection; do not silently recompute old values after a critic update. Check several tiny episodes by hand before training at scale.
Respect legal actions
Save the eligibility mask used when the action was sampled. A later mask change alters its log probability and can corrupt the policy ratio. For a transfer action, stock availability must be determined from the pre-decision state, not a post-decision inventory count. The action-mask lesson explains why a neural score never overrides legality. Invalid logged actions should fail data validation rather than enter the optimizer.
Evaluate independent shifts
Normalize advantages only within a declared training batch and keep reward values in their original units for evaluation. Report order completion, late orders, stockouts and illegal actions by independent held-out shift. A lower surrogate loss can coexist with worse decisions. The applied project compares a policy against a fixed dispatch baseline under matched demand draws.
Implementation
def advantages(rewards, values, following, terminal, cutoff,
discount=0.93, decay=0.87):
count = len(rewards)
if not count or any(len(part) != count for part in
(values, following, terminal, cutoff)):
raise ValueError("rollout fields must align")
result = [0.0] * count
carry = 0.0
for step in range(count - 1, -1, -1):
next_value = 0.0 if terminal[step] else discount * following[step]
error = rewards[step] + next_value - values[step]
carry = error + (0.0 if cutoff[step] else discount * decay * carry)
result[step] = carry
return result
shift_advantages = advantages([2., -1., 3.], [1., .5, 1.2],
[.5, 0., .7], [False, True, False],
[False, True, True])
assert len(shift_advantages) == 3
assert abs(shift_advantages[1] + 1.5) < 1e-9
assert abs(shift_advantages[2] - 2.451) < 1e-9Performance and operating cost
For T rollout steps, this backward pass uses O(T) time and O(T) output memory. Observations and action masks usually dominate collection storage. Several optimizer epochs reuse the same environment experience but cost repeated network passes and can move the policy too far from its recorded behavior. Simulator steps may dominate wall time even when this recurrence is cheap.
Common Mistakes
- Do not bootstrap from a truly terminal state.
- Do not propagate advantages across unrelated shifts.
- Do not recompute old action probabilities with updated actor weights.
