Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit a neural inventory dispatch policy before release

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Train a masked value model in a reproducible simulator, then reject any policy that improves reward while violating stock or service gates.

Specify the decision process

Use a small warehouse with two stock locations and one spare transfer choice at each order event. The state includes only known stock, inbound arrivals already confirmed, backlog and time within a shift. An action either transfers one spare unit or waits; transfers are illegal when the source has no usable stock. Record demand seed, lead-time assumptions, reward version and episode termination. The transition contract gives each event an auditable before-and-after record.

Build baselines before fitting

Run the existing reorder rule and a fixed safe action under the same demand draws. Count completed orders, late orders, stockouts and expedited moves in natural units before compressing them into a reward. A candidate that raises a weighted reward by causing more stockouts has failed even if the score looks better. Keep a development set for reward and model selection and a final set of unseen demand periods.

Wire the model without pretending it is trained

The code performs one masked value update on a tiny synthetic batch. It verifies finite loss, legal next-action selection and a detached target branch. It is a wiring check, not a trained dispatch policy; the full project requires generating many independent episodes and evaluating checkpoints on held-out seeds. The value lesson explains the two estimators and why masks belong to both paths.

Compare under distribution changes

Test slower supplier lead times, bursts of demand and inventory-record delays. Report mean and lower-tail episode return, stockout count, late-order count and illegal-action attempts; group uncertainty by independent episode. Inspect actions near the limits of the replay distribution. A model trained only on quiet periods can overvalue transfer actions in a shortage, so no broad release follows from one simulator setting.

Gate and observe the release

Require no legal-action violations, no increase in stockouts against the baseline, an acceptable late-order rate and bounded inference latency. Start with shadow decisions and log disagreements against the current dispatch rule. Permit a narrow rollout only after operational review, with a reversible switch back to the baseline and a fixed data-collection window. Publish the environment assumptions, replay support table, seeds and failure episodes with the model artifact.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
dispatch_values = nn.Sequential(nn.Linear(3, 16), nn.ReLU(), nn.Linear(16, 2))
target_values = nn.Sequential(nn.Linear(3, 16), nn.ReLU(), nn.Linear(16, 2))
target_values.load_state_dict(dispatch_values.state_dict())
observations = torch.tensor([[8., 3., 1.], [5., 0., 2.]])
next_observations = torch.tensor([[7., 4., 1.], [4., 0., 2.]])
selected_actions = torch.tensor([[1], [0]])
rewards = torch.tensor([2., -3.])
terminal = torch.tensor([False, True])
legal_next = torch.tensor([[True, True], [True, False]])
optimizer = torch.optim.AdamW(dispatch_values.parameters(), lr=0.0007)
with torch.no_grad():
    next_choice = dispatch_values(next_observations).masked_fill(~legal_next, -1e9).argmax(1, keepdim=True)
    future_value = target_values(next_observations).gather(1, next_choice).squeeze(1)
    target = rewards + 0.93 * (~terminal).float() * future_value
prediction = dispatch_values(observations).gather(1, selected_actions).squeeze(1)
loss = functional.smooth_l1_loss(prediction, target)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
assert torch.isfinite(loss) and target_values[0].weight.grad is None

Performance and operating cost

Each update processes B transitions and A action scores; the dominant arithmetic is the neural network forward and backward passes, approximately proportional to B times its parameter operations. Replay requires O(CD) storage for C transitions of dimension D, plus metadata such as episode and policy IDs. Simulator cost scales with episodes and steps per episode. The toy batch demonstrates the update path only; rollout gates require independent episode measurements and an operational fallback.

Common Mistakes

  • Do not let a weighted reward conceal worsening stockouts.
  • Do not release from a simulator calibrated on one quiet period.
  • Do not execute an out-of-support action without a hard eligibility check.

Read next

Continue the workflow: Project: audit a warehouse-shift PPO dispatcher.

ai-data
deep-learning
Storage details