Build a field-box workflow with reviewed coordinate transforms, class-aware duplicate control and whole-receipt release gates.
Project: detect total, date and merchant fields on receipts
Create a field-level annotation contract
Label total, date and merchant boxes on full receipt images. Record original dimensions, edge-coordinate convention and annotation uncertainty; group every crop from one physical receipt in the same split. Define whether an unreadable field is absent, present but ignored, or needs manual review. Audit overlays after resize and crop, especially at narrow page edges. The box lesson fixes the geometry and one-to-one match policy.
Prove the output wiring
The code performs one update for a deliberately restricted one-box-per-image toy head. It predicts a class and a valid normalized box, then combines class loss with box regression loss. It does not represent a full multi-object detector and must not be used to claim recall for receipts with several fields. Replace it with a proper multi-box architecture and assignment rule for the actual task. Keep a tiny deterministic fixture to test class order, box validity and gradient flow.
Train and postprocess under one policy
Use paired image–box transforms and a per-class matching rule. Choose score cutoff and class-aware suppression on development receipts, then freeze them. Preserve all candidate boxes for diagnostic review, but expose only postprocessed detections to the application. Test crowded layouts where date and merchant text overlap and cases with two totals. The suppression lesson describes how a real field can disappear.
Evaluate whole-receipt usefulness
Report per-class precision and recall at a declared IoU threshold, duplicate boxes per receipt, exact presence of all required fields and time saved in review. Break results out by store format, device, blur and cropped edge. A detector can improve average precision while failing the receipts that need the most human attention. Include omitted or uncertain annotations in an explicit ignored-region count rather than silently counting them as background.
Release with a reversible manifest
Package image normalization, class map, box convention, score threshold, suppression IoU, model revision and expected output schema. Replay fixed receipts in a clean process and check coordinate restoration to original pixels. Gate on field recall, duplicate load, p95 end-to-end latency and a manual fallback for missing fields. Keep the prior review flow if any gate fails. The serving contract should include postprocessing time.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
receipt_images = torch.rand(3, 1, 32, 32)
reviewed_classes = torch.tensor([0, 2, 1])
reviewed_boxes = torch.tensor([[0.12, 0.48, 0.66, 0.71],
[0.09, 0.08, 0.57, 0.23],
[0.18, 0.27, 0.81, 0.46]])
field_head = nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(),
nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(8, 7))
optimizer = torch.optim.AdamW(field_head.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
raw = field_head(receipt_images)
center = raw[:, 3:5].sigmoid()
extent = raw[:, 5:7].sigmoid() * 0.5
predicted_boxes = torch.cat(((center - extent / 2).clamp(0, 1),
(center + extent / 2).clamp(0, 1)), dim=1)
assert torch.all(predicted_boxes[:, :2] <= predicted_boxes[:, 2:])
loss = (functional.cross_entropy(raw[:, :3], reviewed_classes) +
functional.smooth_l1_loss(predicted_boxes, reviewed_boxes))
loss.backward()
optimizer.step()
assert torch.isfinite(loss)Performance and operating cost
The toy head makes one class and box prediction per image, so its cost is dominated by the small convolution and cannot stand in for a multi-box detector. Production detectors add candidate generation, assignment and often O(N²) worst-case suppression for N candidates. Box annotations and whole-receipt review are material operating costs. Report preprocessing, model and postprocessing latency separately, plus candidate counts and memory at the real image size.
Common Mistakes
- Do not score the one-box toy as though it detected three fields in one receipt.
- Do not evaluate boxes before restoring original-image coordinates.
- Do not ship a detector without class-aware duplicate and missing-field audits.
Read next
- Bounding-box transforms and one-to-one detection matching
- Class-aware suppression and detection score thresholds
- Project: segment clipped edges and folds on receipts
- Vision decision metrics: separate localization, class errors and abstention
- Inference contracts: preserve preprocessing and measure tail latency
