Train a small patch-token classifier for receipt-quality decisions, then test whether its added attention cost improves small-defect and device-shift outcomes.
Project: compare receipt vision tokens with a convolutional baseline
Fix the data and decision
Use a physical-receipt split with labels for usable image, clipped edge and severe fold. Keep crop, grayscale normalization and review policy identical for the token model and convolutional baseline. Define whether a receipt can carry more than one defect; the miniature code uses one mutually exclusive label per image. Save class IDs and the exact resized image dimensions. The patch lesson determines token count.
Run a token wiring check
The code converts 24-by-32 images into twelve patches, prepends a class token, adds position vectors, applies one encoder layer and updates a three-class head. It checks that the token count and loss are finite. It is an architecture smoke test on random images, not evidence of field accuracy. Train the real candidate from a declared initialization with deterministic validation and compare accepted-example and compute budgets. Logit loss should receive raw class scores.
Test resolution and transfer as separate changes
If using transferred weights at a different image size, resize only the spatial position grid and document head replacement. First compare architectures at one resolution. Then compare old and new resolutions for the selected architecture, holding splits and annotation rules fixed. This prevents a resolution improvement from being misreported as an attention improvement. The position lesson details this step.
Look where the classifier fails
Report clipped-edge recall by edge width, false alerts on clean receipts, device-specific error, calibration and p95 end-to-end latency. Inspect border patches and occluded text. If a class token scores the image correctly but a safety-critical clipped edge is missed, a pixel-level model or targeted crop may be a better workflow. Keep examples of wrong predictions with image and label provenance rather than relying on a single aggregate score. Segmentation is a related alternative.
Package a reversible release
Predeclare minimum small-defect recall, maximum false alerts, memory ceiling and p95 latency. Export patch size, image dimensions, normalization, position table, class map and model revision. Reload and compare logits on fixed images in a clean process. Keep the prior convolutional path if the token model misses any gate. A release report should state the measured advantage and its serving cost, not merely that a transformer was trained.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
receipt_images = torch.rand(3, 1, 24, 32)
quality_labels = torch.tensor([0, 2, 1], dtype=torch.long)
patch_projection = nn.Conv2d(1, 24, kernel_size=8, stride=8)
class_token = nn.Parameter(torch.zeros(1, 1, 24))
positions = nn.Parameter(torch.zeros(1, 13, 24))
encoder = nn.TransformerEncoderLayer(d_model=24, nhead=3,
dim_feedforward=48, batch_first=True)
quality_head = nn.Linear(24, 3)
parameters = list(patch_projection.parameters()) + [class_token, positions]
parameters += list(encoder.parameters()) + list(quality_head.parameters())
optimizer = torch.optim.AdamW(parameters, lr=0.0007)
optimizer.zero_grad(set_to_none=True)
patch_tokens = patch_projection(receipt_images).flatten(2).transpose(1, 2)
all_tokens = torch.cat((class_token.expand(3, -1, -1), patch_tokens), dim=1)
assert all_tokens.shape == (3, 13, 24)
quality_logits = quality_head(encoder(all_tokens + positions)[:, 0])
loss = functional.cross_entropy(quality_logits, quality_labels)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)Performance and operating cost
For N tokens, D embedding width and B images, a standard dense attention layer incurs approximately O(BN²D) arithmetic and O(BN²) attention-map storage, plus projection and feed-forward work. At production resolution N may be far larger than thirteen. The example allocates one tiny encoder layer and performs one update, so it tests wiring only. Compare wall time and peak memory including preprocessing against a convolutional model on the intended hardware and batch policy.
Common Mistakes
- Do not label random-input smoke-test loss as model performance.
- Do not credit architecture for a gain caused by a different image resolution.
- Do not release without small-defect recall and full-pipeline latency checks.
