Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: localize receipt defects with a residual CNN

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a pixel-level defect detector that keeps coordinate alignment across convolution, downsampling and a residual block, then audit failures at crop edges.

Write a mask contract

For each receipt image, store an image identity, native height and width, a binary mask for unreadable regions and a crop transform. Split by receipt identity before extracting overlapping crops. A mask cannot be resized with the same smooth interpolation as a photograph: use nearest-neighbor for discrete labels, and verify that every positive pixel still maps to the right location. Split-before-augmentation prevents neighboring crops from contaminating validation.

Design the spatial path

Use a stride-one feature stem and one stride-two residual block. Upsample to the original crop size for a one-channel logit mask; size-based interpolation removes ambiguity for odd heights and widths. The network below is a minimal shape-checked reference, not a claim that one block is enough for every receipt. Compare its receptive field against the width of known defects. Geometry accounting explains the half-resolution feature grid.

Keep the loss and metric honest

Train on logits with a binary cross-entropy loss; threshold probabilities only for an evaluation decision. Report pixel recall, intersection over union and receipt-level miss rate, with edge defects separated from center defects. A model can earn a good global pixel score by predicting background everywhere when defects are rare. Fix the threshold using validation data; leave a held-out test split untouched. Logits and loss must keep the same target scale.

Challenge the padding choice

Create a test set with marks touching the crop edge and one with the same marks translated into the center. Compare false-negative rates. Test zero padding against an explicit alternative under the same split, optimizer and threshold policy. A change in crop size can also change the feature-map grid, so assert output mask dimensions for 47 by 83, 48 by 84 and a narrow minimum crop. Do not infer edge reliability from center-only visual examples.

Package the serving boundary

Save a full state dictionary, class definition version, normalization values, resize policy, mask threshold and output-coordinate transform. Reload the artifact and check pixel-aligned output for the same fixed image. Log latency and peak memory at the largest accepted crop. The project passes only when an edge-defect counterexample, a blank receipt and a low-contrast receipt all have documented outcomes, including what the product does when confidence is low.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

class ReceiptDefectNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.stem = nn.Sequential(nn.Conv2d(1, 12, 3, padding=1, bias=False),
                                  nn.BatchNorm2d(12), nn.ReLU())
        self.main = nn.Sequential(nn.Conv2d(12, 24, 3, stride=2, padding=1,
                                            bias=False), nn.BatchNorm2d(24), nn.ReLU(),
                                  nn.Conv2d(24, 24, 3, padding=1, bias=False),
                                  nn.BatchNorm2d(24))
        self.shortcut = nn.Conv2d(12, 24, 1, stride=2)
        self.head = nn.Conv2d(24, 1, 1)

    def forward(self, receipt_images: torch.Tensor) -> torch.Tensor:
        high_resolution = self.stem(receipt_images)
        low_resolution = functional.relu(self.main(high_resolution)
                                         + self.shortcut(high_resolution))
        low_resolution_logits = self.head(low_resolution)
        return functional.interpolate(low_resolution_logits,
                                      size=receipt_images.shape[-2:],
                                      mode="bilinear", align_corners=False)

model = ReceiptDefectNet()
receipt_images = torch.rand(2, 1, 47, 83)
defect_masks = torch.zeros_like(receipt_images)
defect_masks[:, :, 11:16, 33:39] = 1
mask_logits = model(receipt_images)
loss = functional.binary_cross_entropy_with_logits(mask_logits, defect_masks)
loss.backward()
assert mask_logits.shape == defect_masks.shape
assert torch.isfinite(loss)

Performance and operating cost

Convolution work grows with image area, kernel area and channel products; the retained high-resolution stem activation can dominate memory. Bilinear output resizing costs O(batch × image pixels) and does not recover detail already erased by a coarse feature grid. Track receipt-level miss rate alongside pixel metrics, and profile largest accepted crops rather than only the 47 by 83 reference.

Common Mistakes

  • Do not split overlapping crops from one receipt across training and validation.
  • Do not use smooth interpolation on a binary ground-truth mask.
  • Do not treat a high background-dominated pixel accuracy as defect detection.

Read next

Continue the workflow: Occlusion attribution and explanation sanity checks.

Continue the workflow: Diffusion reverse steps and conditional guidance.

Continue the workflow: Contrastive view identity and augmentation contracts.

Continue the workflow: Project: segment clipped edges and folds on receipts.

Continue the workflow: Project: compare receipt vision tokens with a convolutional baseline.

Continue the workflow: Project: detect total, date and merchant fields on receipts.

ai-data
deep-learning
Storage details