A vision transformer receives an ordered patch sequence; image padding, patch size and class-token position determine which visual evidence reaches attention.
Vision transformer patch grids and token contracts
Compute the grid before the encoder
A 24-by-32 receipt with 8-by-8 nonoverlapping patches produces a 3-by-4 grid, or twelve image tokens. A separate class token raises the sequence length to thirteen. Record channel count, color normalization, resize policy, patch dimensions, grid order and any padded border. A crop whose size is not divisible by the patch size needs a declared pad or crop rule. The code computes exactly this miniature grid. Convolution geometry predicts the projection output.
Treat position as part of the input contract
Self-attention alone cannot tell whether the same token content came from the top-left or bottom-right patch. Add a position representation after fixing row-major token order. If a class token is prepended, it needs its own position slot; reshaping it into the two-dimensional image grid is wrong. A token permutation with unchanged positions changes the image meaning. Tests should assert that flattening and unflattening preserve row and column identity.
Inspect patch-size information loss
A tiny clipped edge can occupy less than one patch. Coarse patches reduce token count and attention cost but make it easier to lose boundary detail in the projection. Smaller patches increase the number of tokens approximately with inverse patch area and can raise attention-map memory quadratically in token count. Compare patch choices on the same physical-receipt holdout, including small defects and blur. Pixel masks require an additional dense-output design beyond one class token.
Keep padding from becoming a cue
If only one defect class tends to use padded images, a classifier may learn border geometry instead of the defect. Record valid image area and use identical resizing rules across train and evaluation. For variable aspect ratios, test whether the model supports variable grids and whether padded patch tokens must be masked in attention. A standard fixed-grid encoder may simply expect one image size. Do not infer variable-size support from a successful tensor reshape.
Compare with a cheaper reference
Run a convolutional classifier under the same split, preprocessing and compute budget. Report small-defect recall, false alerts, p95 preprocessing-plus-model latency and peak memory, not only aggregate accuracy. The patch sequence may be appropriate when broad context matters, but an image model earns its complexity through measured improvement. The project tests that decision and preserves the token contract for serving.
Implementation
import torch
from torch import nn
receipt_images = torch.randn(2, 1, 24, 32)
patch_size = 8
patch_projection = nn.Conv2d(1, 16, kernel_size=patch_size,
stride=patch_size)
patch_grid = patch_projection(receipt_images)
assert patch_grid.shape == (2, 16, 3, 4)
image_tokens = patch_grid.flatten(2).transpose(1, 2)
class_token = nn.Parameter(torch.zeros(1, 1, 16))
position_vectors = nn.Parameter(torch.zeros(1, 13, 16))
tokens = torch.cat((class_token.expand(2, -1, -1), image_tokens), dim=1)
tokens = tokens + position_vectors
assert tokens.shape == (2, 13, 16)
assert image_tokens.size(1) == (24 // patch_size) * (32 // patch_size)Performance and operating cost
Patch projection is proportional to the image area and embedding width. For B images, N tokens and embedding width D, dense self-attention has approximately O(BN²D) arithmetic and O(BN²) attention-map storage per layer before implementation-specific optimizations. Halving both patch dimensions roughly quadruples N and can increase the attention map by about sixteenfold. The small example checks dimensions only; it does not train an attention model or establish accuracy. Measure memory at the intended resolution.
Common Mistakes
- Do not forget the extra class-token position when checking sequence length.
- Do not assume a patch boundary preserves a one-pixel defect.
- Do not resize or pad different receipt classes with different rules.
