A pixel classifier is trained on image–mask geometry; a wrong interpolation rule or ignored-label denominator can turn accurate code into a different task.
Segmentation mask alignment and ignored pixels
Name the mask classes
For receipt defects, one mask might distinguish ordinary paper, clipped edge and obscuring fold. Record integer IDs, mutually exclusive or overlapping semantics, and an ignore value for pixels whose annotation cannot be trusted. A mutually exclusive target with one ID per pixel belongs with per-pixel multiclass logits and cross entropy. Independent overlapping defect layers require a different multi-label loss and threshold policy. Do not switch between these representations without changing the target and metrics. The logits rule applies at every pixel, not only once per image.
Keep geometric transforms paired
A crop, horizontal shift, perspective transform or resize applied to the image must apply the same geometric mapping to the mask. Image intensity can be interpolated smoothly; class IDs cannot. Use nearest-neighbor label interpolation and verify that resized masks contain only known IDs and the ignore value. An image resized independently of its mask may still have the expected width and height while every defect boundary has moved. Save several transformed image–mask overlays at fixed seeds. Split identities before generating augmented pairs.
Handle unknown borders explicitly
When a crop or warp introduces pixels outside the annotated source, decide whether they become background or ignored; the default should not silently invent negative evidence. In the example, 255 means ignored. A cross-entropy implementation excludes ignored targets from the loss denominator when class-index targets are used. Count accepted pixels per batch and reject a batch with no accepted pixels rather than dividing by zero or learning from an empty annotation. For rare clipped edges, report positive-pixel and positive-image counts separately.
Align model output resolution
A model may emit logits at one-quarter of input resolution. Either upsample logits to the mask grid with a documented continuous interpolation rule, or reduce the mask using a label-preserving policy that does not erase narrow defects. These two choices are not interchangeable for a one-pixel clipped edge. Record output stride, final-grid size and any tile padding. Convolution geometry predicts the feature-map shape, but visual overlays verify whether the physical boundary still lines up.
Verify the full target path
Run a tiny batch through decode, transform, model, loss and metric computation. Assert image–mask spatial dimensions, allowed mask values, finite loss and unchanged ignored-pixel count after transforms that should preserve area. Save a fixed batch to compare after library or preprocessing upgrades. A low training loss with most pixels marked background may say little about the defect class. The metric policy must therefore be fixed separately from the loss.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
receipt_image = torch.arange(16, dtype=torch.float32).view(1, 1, 4, 4) / 15
defect_mask = torch.tensor([[[0, 0, 1, 1], [0, 2, 2, 1],
[0, 2, 255, 0], [0, 0, 0, 0]]])
resized_image = functional.interpolate(receipt_image, size=(8, 8),
mode="bilinear", align_corners=False)
resized_mask = functional.interpolate(defect_mask.unsqueeze(1).float(),
size=(8, 8), mode="nearest-exact").squeeze(1).long()
segmentation_head = nn.Conv2d(1, 3, kernel_size=1)
pixel_logits = segmentation_head(resized_image)
valid_count = int(resized_mask.ne(255).sum())
assert valid_count > 0
pixel_loss = functional.cross_entropy(pixel_logits, resized_mask, ignore_index=255)
assert set(resized_mask.unique().tolist()) <= {0, 1, 2, 255}
assert pixel_logits.shape == (1, 3, 8, 8)
assert torch.isfinite(pixel_loss)Performance and operating cost
Image and mask transforms cost O(BHW) for B images of height H and width W, while the model usually dominates training compute. Per-pixel logits and gradients occupy O(BCHW) memory for C classes at output resolution. Upsampling logits to a large mask grid can raise activation memory sharply; downsampling masks can erase small defects. The example uses a tiny one-by-one head solely to test the label path. Measure accepted-pixel counts and peak memory at production image size.
Common Mistakes
- Do not bilinearly interpolate integer class IDs.
- Do not treat unannotated pixels as confirmed background without an annotation rule.
- Do not accept matching tensor shapes as proof that image and mask geometry align.
Read next
- Convolution output geometry and receptive fields
- Image augmentation: split originals first and preserve the label
- Dice, IoU, empty masks and threshold policy
- Project: segment clipped edges and folds on receipts
- Project: localize receipt defects with a residual CNN
Continue the workflow: Bounding-box transforms and one-to-one detection matching.
Continue the workflow: Point-set permutation invariance and pooled features.
