Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: segment clipped edges and folds on receipts

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Train a pixel-level defect model with ignored annotation regions, then prove that mask overlap and image-level alerts survive held-out receipts and full-resolution inference.

Prepare paired annotations

Collect grayscale receipt images with reviewed masks for background, clipped edge and fold, plus an ignore code for uncertain borders. Split by physical receipt before generating crops and keep annotation versions in the manifest. Audit overlays from multiple capture devices, especially narrow clipped edges near the crop boundary. The tiny model below runs one training step to validate dimensions and ignored-pixel handling; it is not a production architecture. The alignment lesson defines the preprocessing contract.

Train with explicit pixel accounting

Use per-pixel multiclass logits and ignore uncertain pixels in the loss. Count accepted pixels and positive pixels per batch; reject empty-supervision batches. Sample images so rare defect masks are seen without repeating a few physical receipts so often that validation becomes misleading. Keep all geometric image transforms paired with nearest-neighbor mask transforms. Save a fixed image–mask–logit triplet after each code change. A model can lower loss by learning background while missing almost every clipped edge.

Validate overlap and detection

Choose a development-set checkpoint and any postprocessing threshold before final test. Report per-class Dice and IoU with an explicit both-empty policy, plus image-level clipped-edge recall and false alerts. Preserve counts by device and defect size. Compare to a simple image-level classifier: pixel masks cost annotation and runtime, so they should earn that cost through better localization or decisions. The metric lesson specifies counts and denominators.

Test full-resolution tiles

If inference uses overlapping tiles, account for padding, coordinate offsets and merging. A defect split across tile boundaries must not vanish or be double-counted. Compare full-image predictions from tiled and untiled runs where memory allows; inspect seam regions and near-edge masks. Record tile size, overlap, blending and output threshold with the model. Measure p95 end-to-end latency including image decode, tiling and mask assembly, not only the neural forward. The serving contract applies to masks too.

Release only with a reversible gate

Define minimum clipped-edge image recall, maximum false alerts, per-class overlap floor, review capacity and latency before final test. Export class map, preprocessing, tile policy and ignore-label treatment. Reload a fixed batch in a clean process and compare mask logits. Keep the previous classifier or manual review route if the segmentation candidate misses a gate. Deliver annotated examples, mask-overlap counts and a failure gallery, not a single aggregate score.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
receipt_images = torch.rand(4, 1, 8, 8)
reviewed_masks = torch.zeros(4, 8, 8, dtype=torch.long)
reviewed_masks[:, 2:4, 5:7] = 1
reviewed_masks[:, 5:7, 1:4] = 2
reviewed_masks[:, 0, 0] = 255
model = nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(),
                      nn.Conv2d(8, 3, 1))
optimizer = torch.optim.AdamW(model.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
logits = model(receipt_images)
assert logits.shape == (4, 3, 8, 8)
assert int(reviewed_masks.ne(255).sum()) == 252
loss = functional.cross_entropy(logits, reviewed_masks, ignore_index=255)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)

Performance and operating cost

Dense per-pixel logits use O(BCHW) storage and a full-resolution model may need substantially more activation memory than an image classifier. Tiled inference trades peak memory for repeated border compute and merging work; overlapping tiles can increase total pixels processed. Mask annotation and review are material project costs. The toy model has only two convolutions and proves tensor and loss wiring, not accuracy. Benchmark image decode, full-resolution assembly and serving latency on the target runtime.

Common Mistakes

  • Do not claim a high pixel score when clipped-edge image recall is poor.
  • Do not train on masks bilinearly resized into fractional class IDs.
  • Do not evaluate tiled images without checking seams and coordinate restoration.

Read next

Continue the workflow: Project: recognize receipt line items with CTC.

Continue the workflow: Project: compare receipt vision tokens with a convolutional baseline.

ai-data
deep-learning
Storage details