Build a pixel-level defect detector that keeps coordinate alignment across convolution, downsampling and a residual block, then audit failures at crop edges.
Project: localize receipt defects with a residual CNN
Write a mask contract
For each receipt image, store an image identity, native height and width, a binary mask for unreadable regions and a crop transform. Split by receipt identity before extracting overlapping crops. A mask cannot be resized with the same smooth interpolation as a photograph: use nearest-neighbor for discrete labels, and verify that every positive pixel still maps to the right location. Split-before-augmentation prevents neighboring crops from contaminating validation.
Design the spatial path
Use a stride-one feature stem and one stride-two residual block. Upsample to the original crop size for a one-channel logit mask; size-based interpolation removes ambiguity for odd heights and widths. The network below is a minimal shape-checked reference, not a claim that one block is enough for every receipt. Compare its receptive field against the width of known defects. Geometry accounting explains the half-resolution feature grid.
Keep the loss and metric honest
Train on logits with a binary cross-entropy loss; threshold probabilities only for an evaluation decision. Report pixel recall, intersection over union and receipt-level miss rate, with edge defects separated from center defects. A model can earn a good global pixel score by predicting background everywhere when defects are rare. Fix the threshold using validation data; leave a held-out test split untouched. Logits and loss must keep the same target scale.
Challenge the padding choice
Create a test set with marks touching the crop edge and one with the same marks translated into the center. Compare false-negative rates. Test zero padding against an explicit alternative under the same split, optimizer and threshold policy. A change in crop size can also change the feature-map grid, so assert output mask dimensions for 47 by 83, 48 by 84 and a narrow minimum crop. Do not infer edge reliability from center-only visual examples.
Package the serving boundary
Save a full state dictionary, class definition version, normalization values, resize policy, mask threshold and output-coordinate transform. Reload the artifact and check pixel-aligned output for the same fixed image. Log latency and peak memory at the largest accepted crop. The project passes only when an edge-defect counterexample, a blank receipt and a low-contrast receipt all have documented outcomes, including what the product does when confidence is low.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
class ReceiptDefectNet(nn.Module):
def __init__(self):
super().__init__()
self.stem = nn.Sequential(nn.Conv2d(1, 12, 3, padding=1, bias=False),
nn.BatchNorm2d(12), nn.ReLU())
self.main = nn.Sequential(nn.Conv2d(12, 24, 3, stride=2, padding=1,
bias=False), nn.BatchNorm2d(24), nn.ReLU(),
nn.Conv2d(24, 24, 3, padding=1, bias=False),
nn.BatchNorm2d(24))
self.shortcut = nn.Conv2d(12, 24, 1, stride=2)
self.head = nn.Conv2d(24, 1, 1)
def forward(self, receipt_images: torch.Tensor) -> torch.Tensor:
high_resolution = self.stem(receipt_images)
low_resolution = functional.relu(self.main(high_resolution)
+ self.shortcut(high_resolution))
low_resolution_logits = self.head(low_resolution)
return functional.interpolate(low_resolution_logits,
size=receipt_images.shape[-2:],
mode="bilinear", align_corners=False)
model = ReceiptDefectNet()
receipt_images = torch.rand(2, 1, 47, 83)
defect_masks = torch.zeros_like(receipt_images)
defect_masks[:, :, 11:16, 33:39] = 1
mask_logits = model(receipt_images)
loss = functional.binary_cross_entropy_with_logits(mask_logits, defect_masks)
loss.backward()
assert mask_logits.shape == defect_masks.shape
assert torch.isfinite(loss)Performance and operating cost
Convolution work grows with image area, kernel area and channel products; the retained high-resolution stem activation can dominate memory. Bilinear output resizing costs O(batch × image pixels) and does not recover detail already erased by a coarse feature grid. Track receipt-level miss rate alongside pixel metrics, and profile largest accepted crops rather than only the 47 by 83 reference.
Common Mistakes
- Do not split overlapping crops from one receipt across training and validation.
- Do not use smooth interpolation on a binary ground-truth mask.
- Do not treat a high background-dominated pixel accuracy as defect detection.
Read next
- Convolution output geometry and receptive fields
- Residual blocks and normalization state
- Image augmentation: split originals first and preserve the label
- Logits, cross-entropy and gradients: align the training calculation
- Inference contracts: preserve preprocessing and measure tail latency
Continue the workflow: Occlusion attribution and explanation sanity checks.
Continue the workflow: Diffusion reverse steps and conditional guidance.
Continue the workflow: Contrastive view identity and augmentation contracts.
Continue the workflow: Project: segment clipped edges and folds on receipts.
Continue the workflow: Project: compare receipt vision tokens with a convolutional baseline.
Continue the workflow: Project: detect total, date and merchant fields on receipts.
