Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Adversarial training and the clean–stress tradeoff

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Training on perturbed inputs can reduce failures under a chosen test, but it can also change clean accuracy and does not cover threats outside the declared search.

Keep a clean baseline

Train and freeze a reference receipt classifier under the existing split, preprocessing and update budget. Measure normal validation and final-test decisions by defect and capture device. The adversarial-training candidate must use the same labeled training identities and a declared extra compute budget. If it sees more real examples or gets more optimizer updates, any difference cannot be assigned to the perturbation policy alone. The evaluation loop should preserve group boundaries and identical class maps.

Generate perturbations against the current candidate

For each training batch, compute a gradient with respect to the input while holding reviewed labels fixed. Apply a bounded change in the declared pixel space, clamp it and detach the perturbed input before the parameter-update forward pass. The code combines clean and perturbed losses in one optimizer update. It is a one-step stress objective, not an exhaustive inner maximization. If model weights or preprocessing change, cached perturbations become stale. The threat-model lesson defines the numeric bound.

Balance losses and inspect slices

A weighted clean-plus-stress loss exposes a choice: too little stress weight may leave the attack effective; too much may lower performance on ordinary images. Tune the weight and radius on development data, then lock them before final test. Monitor clipped-edge recall, false accepts, confidence and manual-review volume under clean, numeric and natural-corruption conditions. A smaller stress error with a larger clean false-accept rate is not automatically an improvement for the product. The confidence audit makes this operational cost visible.

Avoid gradient masking claims

A model may appear resistant to one gradient attack because of saturated preprocessing or a non-smooth operation, not because the decision is stable. Test more than one search strength, multiple starting points where appropriate and physically plausible corruptions. Check input gradients for finite, nontrivial values and compare against a clean model under the same attack. Do not state a certified guarantee without a method that actually proves one. Keep attack code and normalization versioned so a future audit can reproduce the reported failures.

Decide on measured constraints

Set release gates before inspecting final test: minimum clean rare-defect recall, maximum false accepts, acceptable review capacity and target stress performance. If no candidate satisfies all gates, keep the reference model and revise data or review policy. The project runs this comparison across numeric and natural perturbations. The training objective is one engineering control, not a substitute for capture quality checks or human review.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
quality_model = nn.Sequential(nn.Flatten(), nn.Linear(16, 3))
optimizer = torch.optim.AdamW(quality_model.parameters(), lr=0.0007)
clean_pixels = torch.rand(5, 1, 4, 4)
reviewed_labels = torch.tensor([0, 1, 2, 1, 0])
attack_source = clean_pixels.detach().clone().requires_grad_(True)
attack_loss = functional.cross_entropy(quality_model(attack_source), reviewed_labels)
input_gradient = torch.autograd.grad(attack_loss, attack_source)[0]
perturbed_pixels = (attack_source.detach() + 0.04 * input_gradient.sign()).clamp(0, 1)

optimizer.zero_grad(set_to_none=True)
clean_loss = functional.cross_entropy(quality_model(clean_pixels), reviewed_labels)
stress_loss = functional.cross_entropy(quality_model(perturbed_pixels), reviewed_labels)
combined_loss = 0.65 * clean_loss + 0.35 * stress_loss
combined_loss.backward()
optimizer.step()
assert torch.isfinite(combined_loss)
assert not perturbed_pixels.requires_grad

Performance and operating cost

This one-step training example adds an attack forward and backward pass, then evaluates the clean and perturbed batches for the parameter update. It can therefore cost several times a clean-only update, depending on model and batching. Stronger inner searches increase cost with search steps. Input perturbation storage is O(BCHW) and ordinary model gradients still require parameter memory. A numerical stress gain is useful only when clean performance, natural-corruption behavior and deployment latency remain inside the declared gates.

Common Mistakes

  • Do not report robustness only at the radius used during training.
  • Do not train on final-test receipt identities or tune the clean–stress weight on final test.
  • Do not mistake weak input gradients for proven resistance to input changes.

Read next

ai-data
deep-learning
Storage details