Skip to content
AITroveRead. Build. Understand.
Make this comfortable

FGSM input scale and perturbation threat models

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A gradient-sign perturbation tests one precisely bounded input change; its numeric radius means nothing until image units and attacker access are defined.

State what can change

Choose whether the test perturbs raw image intensities, normalized tensors, a document crop or a full camera frame. Specify the norm, maximum change, valid pixel range and whether the perturbation may cover text, background or borders. A radius of 0.06 in zero-to-one pixel space is not the same as 0.06 after standardization by a small channel deviation. Real capture problems such as motion blur and paper folds are separate stressors. Augmentation policy may already cover some natural changes, but it does not define an adversarial test.

Differentiate the input, not parameters

For a fixed model and reviewed target label, create a clean input tensor with gradients enabled and compute loss. Obtain the gradient with respect to that input, then add radius times its sign for an untargeted one-step perturbation. Clamp to the valid pixel range and, if preprocessing uses normalization, return to the intended space before applying the bound. Detach the resulting image for evaluation. The code uses a tiny classifier and checks the infinity-norm bound; it says nothing about how a trained receipt model will behave.

Keep the threat model honest

A one-step gradient-sign search is a diagnostic under white-box access to model gradients, not a proof that all allowed perturbations have been explored. Results depend on clean correctness, label quality and input clipping. Evaluate attack success on examples the model originally classified correctly as well as on the whole population, and report both denominators. Repeated or stronger searches can find failures a single step misses. Do not describe one failed attack as certification of robustness or as a guarantee against camera and print changes.

Inspect perturbation visibility

Save clean and perturbed pairs at actual display scaling, with the changed-pixel range and regions affected. Receipt text and cut-off edges are semantically sensitive; an imperceptible numeric change may still alter OCR or downstream human interpretation in a low-contrast area. Conversely, a large numeric change may be unrealistic for a capture pipeline. Measure class-specific failures and confidence shifts, especially confident false accepts. Selective risk shows whether a review policy catches those cases.

Build a repeatable stress suite

Version the model, preprocessing, label map, input units, radius and attack implementation. Test multiple radii and independent natural corruption groups without combining them into a single ambiguous score. Keep physical receipt identities out of training if they are in the held-out stress set. The applied project pairs numeric attacks with physically plausible changes and a release decision.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
quality_model = nn.Sequential(nn.Flatten(), nn.Linear(16, 3))
clean_pixels = torch.rand(4, 1, 4, 4).detach().requires_grad_(True)
reviewed_labels = torch.tensor([0, 2, 1, 0])
clean_logits = quality_model(clean_pixels)
clean_loss = functional.cross_entropy(clean_logits, reviewed_labels)
input_gradient = torch.autograd.grad(clean_loss, clean_pixels)[0]
radius = 0.06
perturbed_pixels = (clean_pixels.detach() + radius * input_gradient.sign()).clamp(0, 1)
maximum_change = (perturbed_pixels - clean_pixels.detach()).abs().amax()
assert maximum_change <= radius + 1e-7
assert bool(((perturbed_pixels >= 0) & (perturbed_pixels <= 1)).all())

Performance and operating cost

One gradient-sign step adds a model forward and backward pass per stress batch, followed by O(BCHW) input arithmetic. Evaluating several radii repeats model work unless the same input gradient is intentionally reused under a stated protocol. More intensive searches can cost many forwards and backwards per example. The cost is acceptable for an offline audit, but applying such a search to every live request would change latency and the threat model. Report hardware, batch size and search budget with robustness results.

Common Mistakes

  • Do not apply a raw-pixel radius directly to standardized tensors without conversion.
  • Do not call one unsuccessful attack a robustness certificate.
  • Do not count initially misclassified examples as new attack successes without separating denominators.

Read next

ai-data
deep-learning
Storage details