Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit a warehouse-shift PPO dispatcher

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Train an actor-critic policy in a versioned simulator and gate release on episode service, legality and response time.

Specify the environment

At each order event, the actor sees stock, known inbound units, backlog and remaining shift time. Actions are wait or transfer one spare unit when available. Book service and transfer costs exactly once, then log observation, mask, decision, reward and next state. Freeze demand seeds and supplier-delay assumptions for each comparison. The rollout lesson separates terminal shifts from collection cutoffs.

Establish fixed baselines

Run the existing threshold dispatch rule and a no-transfer policy under the same demand draws. Record completed orders, late orders, stockouts, expedited moves and invalid attempts in natural units before compressing them into reward. Choose reward weights only with development simulations. A learned actor that raises weighted reward while increasing stockouts fails the service gate. Hold back later demand periods and different delay regimes.

Train with measured updates

Collect fresh rollouts with old action log probabilities, compute advantages and returns, then apply a limited number of clipped actor-critic updates. Track value error, ratio clipping, approximate KL, entropy and action frequency. The snippet makes one synthetic update to verify the tensor path; it is not a trained warehouse controller. The update lesson explains the early-stop guard.

Stress independent shifts

Evaluate multiple held-out demand seeds, burst periods, supplier delays and missing stock observations. Report mean and lower-tail return, stockouts, late orders, illegal actions and disagreements against the baseline. Replay each decision with only information available at that moment. Inspect any shift where return improves but a service metric worsens. Save failure traces with environment and policy versions for reproduction.

Release behind a hard filter

Shadow-score operational events before allowing any action. Keep a deterministic legality filter, timeout fallback and the old dispatcher. Require no increase in stockouts, a defined lateness budget and bounded p95 decision time. Package actor, critic, feature schema, reward version and rollback pointer. Simulator gains alone do not establish safe unsupervised dispatch.

Implementation

python
import torch
from torch import nn
from torch.distributions import Categorical
from torch.nn import functional as functional

torch.manual_seed(47)
states = torch.tensor([[8., 3., 2.], [5., 0., 1.], [7., 2., 3.]])
legal = torch.tensor([[True, True], [True, False], [True, True]])
actions = torch.tensor([1, 0, 0])
returns = torch.tensor([2.4, -1.1, 3.2])
actor, critic = nn.Linear(3, 2), nn.Linear(3, 1)
optimizer = torch.optim.AdamW(list(actor.parameters()) + list(critic.parameters()), lr=.0007)
new_log_prob = Categorical(logits=actor(states).masked_fill(~legal, -1e9)).log_prob(actions)
old_log_prob = new_log_prob.detach() - torch.tensor([.1, -.08, .04])
values = critic(states).squeeze(1)
advantage = returns - values.detach()
ratio = (new_log_prob - old_log_prob).exp()
actor_loss = -torch.minimum(ratio * advantage, ratio.clamp(.83, 1.17) * advantage).mean()
loss = actor_loss + .4 * functional.smooth_l1_loss(values, returns)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
assert torch.isfinite(loss) and all(legal[row, action] for row, action in enumerate(actions))

Performance and operating cost

A training update evaluates actor and critic on B samples; repeated epochs multiply that cost. Rollout generation takes E environment steps and may dominate elapsed time, while stored masks can require O(EA) memory for A actions. The synthetic batch checks only finite gradients and legal recorded choices. A release needs independent episode measurements, simulator validation and a hard serving guard.

Common Mistakes

  • Do not mistake a synthetic update for measured dispatch performance.
  • Do not reuse stale rollout probabilities after a policy revision.
  • Do not remove the hard action filter when the actor scores well.

Read next

ai-data
deep-learning
Storage details