Skip to content
AITroveRead. Build. Understand.
Make this comfortable

PPO clipped policy ratios, critic loss and KL guards

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Clipping limits the update incentive on sampled actions, while critic error and policy movement still need separate checks.

Use the recorded denominator

Subtract the old log probability captured during rollout from the current actor log probability for the same state, action and legal-action mask. Exponentiating gives the probability ratio. A ratio close to one means little movement on that recorded action, not necessarily every action. The rollout contract freezes the old denominator and makes accidental reuse of stale data visible.

Clip the objective

Multiply the ratio by the estimated advantage and compare that contribution with the ratio clamped around one. Choose the smaller contribution for maximization, or minimize its negative. Positive and negative advantages respond differently; test both signs on a small batch. The code checks finite gradients in this objective. The clamp limits a particular incentive, but repeated optimization or other losses can still move the full policy substantially.

Keep critic diagnostics separate

The critic predicts discounted return; it is not a second action chooser. Fit it against returns built from rollout estimates and advantages. Report value error apart from actor loss, especially if both share a feature encoder. Record entropy and its coefficient if used for exploration. Optimizer settings and clipping thresholds belong in the checkpoint, not in undocumented notebook cells.

Watch policy movement

Track approximate KL on collected states, ratio spread, clipping fraction and entropy by optimizer epoch. Define a KL stop rule before the experiment. The estimate is diagnostic, not a guarantee about unvisited states. If most ratios are clipped, more epochs may waste compute or worsen decisions. Compare across independent simulator seeds instead of selecting the smoothest training curve.

Gate operating outcomes

Evaluate order completion, stockouts, late orders, transfer cost and action-mask violations on held-out demand sequences. A healthy surrogate objective cannot compensate for a reward booked twice or an incorrect simulator. The project keeps episode-level service gates and a deterministic rollback policy alongside training metrics.

Implementation

python
import torch
from torch import nn
from torch.distributions import Categorical

torch.manual_seed(47)
stock_states = torch.tensor([[8., 3.], [5., 0.], [7., 2.]])
legal_actions = torch.tensor([[True, True], [True, False], [True, True]])
recorded_actions = torch.tensor([1, 0, 0])
estimated_advantages = torch.tensor([1.3, -0.7, 0.4])
actor = nn.Linear(2, 2)
new_logits = actor(stock_states).masked_fill(~legal_actions, -1e9)
new_log_prob = Categorical(logits=new_logits).log_prob(recorded_actions)
old_log_prob = new_log_prob.detach() - torch.tensor([0.12, -0.09, 0.05])
ratio = (new_log_prob - old_log_prob).exp()
unclipped = ratio * estimated_advantages
clipped = ratio.clamp(0.83, 1.17) * estimated_advantages
actor_loss = -torch.minimum(unclipped, clipped).mean()
actor_loss.backward()
assert torch.isfinite(actor_loss) and actor.weight.grad is not None

Performance and operating cost

For B samples and A actions, the probability calculation and mask add O(BA) operations around the actor forward. Extra optimizer epochs multiply network compute and activation storage. Old log probabilities, actions, masks and advantages occupy O(B) to O(BA) rollout memory. Clipping has no hard guarantee on complete policy divergence; measure KL and held-out episode behavior.

Common Mistakes

  • Do not recompute the denominator with the new actor.
  • Do not call objective clipping a hard safety constraint.
  • Do not hide critic collapse inside a combined loss value.

Read next

ai-data
deep-learning
Storage details