A gradient-sign perturbation tests one precisely bounded input change; its numeric radius means nothing until image units and attacker access are defined.
FGSM input scale and perturbation threat models
State what can change
Choose whether the test perturbs raw image intensities, normalized tensors, a document crop or a full camera frame. Specify the norm, maximum change, valid pixel range and whether the perturbation may cover text, background or borders. A radius of 0.06 in zero-to-one pixel space is not the same as 0.06 after standardization by a small channel deviation. Real capture problems such as motion blur and paper folds are separate stressors. Augmentation policy may already cover some natural changes, but it does not define an adversarial test.
Differentiate the input, not parameters
For a fixed model and reviewed target label, create a clean input tensor with gradients enabled and compute loss. Obtain the gradient with respect to that input, then add radius times its sign for an untargeted one-step perturbation. Clamp to the valid pixel range and, if preprocessing uses normalization, return to the intended space before applying the bound. Detach the resulting image for evaluation. The code uses a tiny classifier and checks the infinity-norm bound; it says nothing about how a trained receipt model will behave.
Keep the threat model honest
A one-step gradient-sign search is a diagnostic under white-box access to model gradients, not a proof that all allowed perturbations have been explored. Results depend on clean correctness, label quality and input clipping. Evaluate attack success on examples the model originally classified correctly as well as on the whole population, and report both denominators. Repeated or stronger searches can find failures a single step misses. Do not describe one failed attack as certification of robustness or as a guarantee against camera and print changes.
Inspect perturbation visibility
Save clean and perturbed pairs at actual display scaling, with the changed-pixel range and regions affected. Receipt text and cut-off edges are semantically sensitive; an imperceptible numeric change may still alter OCR or downstream human interpretation in a low-contrast area. Conversely, a large numeric change may be unrealistic for a capture pipeline. Measure class-specific failures and confidence shifts, especially confident false accepts. Selective risk shows whether a review policy catches those cases.
Build a repeatable stress suite
Version the model, preprocessing, label map, input units, radius and attack implementation. Test multiple radii and independent natural corruption groups without combining them into a single ambiguous score. Keep physical receipt identities out of training if they are in the held-out stress set. The applied project pairs numeric attacks with physically plausible changes and a release decision.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
quality_model = nn.Sequential(nn.Flatten(), nn.Linear(16, 3))
clean_pixels = torch.rand(4, 1, 4, 4).detach().requires_grad_(True)
reviewed_labels = torch.tensor([0, 2, 1, 0])
clean_logits = quality_model(clean_pixels)
clean_loss = functional.cross_entropy(clean_logits, reviewed_labels)
input_gradient = torch.autograd.grad(clean_loss, clean_pixels)[0]
radius = 0.06
perturbed_pixels = (clean_pixels.detach() + radius * input_gradient.sign()).clamp(0, 1)
maximum_change = (perturbed_pixels - clean_pixels.detach()).abs().amax()
assert maximum_change <= radius + 1e-7
assert bool(((perturbed_pixels >= 0) & (perturbed_pixels <= 1)).all())Performance and operating cost
One gradient-sign step adds a model forward and backward pass per stress batch, followed by O(BCHW) input arithmetic. Evaluating several radii repeats model work unless the same input gradient is intentionally reused under a stated protocol. More intensive searches can cost many forwards and backwards per example. The cost is acceptable for an offline audit, but applying such a search to every live request would change latency and the threat model. Report hardware, batch size and search budget with robustness results.
Common Mistakes
- Do not apply a raw-pixel radius directly to standardized tensors without conversion.
- Do not call one unsuccessful attack a robustness certificate.
- Do not count initially misclassified examples as new attack successes without separating denominators.
Read next
- Image augmentation: split originals first and preserve the label
- Adversarial training and the clean–stress tradeoff
- Project: stress-test receipt decisions under bounded image changes
- Project: audit receipt confidence and review handoff
- Inference contracts: preserve preprocessing and measure tail latency
