A convolutional stack changes spatial resolution and the portion of the input each output can inspect; both properties belong in the model contract.
Convolution output geometry and receptive fields
Calculate the grid before training
For one spatial axis, the output length is floor((input + 2 × padding − dilation × (kernel − 1) − 1) / stride + 1). Calculate height and width separately when kernels or strides differ. A 47 by 83 receipt crop sent through a three-wide convolution with padding one and stride two becomes 24 by 42, not 23 by 41. This floor operation matters at odd sizes. Assert the observed tensor shape against the calculation before connecting a classifier or decoder. The tensor contract also fixes channel order.
Separate resolution from context
A three-wide kernel at dilation one sees three neighboring positions in one layer. A second three-wide, stride-one layer expands the theoretical receptive field to five; a stride-two layer changes the spacing between positions used by later layers. Track two values: receptive-field width and jump, the distance in original pixels between adjacent outputs. Starting at one for both, a layer with kernel k, dilation d and stride s adds (k − 1) × d × previous jump to field width and multiplies jump by s. Large theoretical context is no guarantee that training uses distant pixels.
Choose padding for the task
Zero padding creates artificial borders. That may be acceptable for whole-image classification, but a defect next to a cropped edge can receive a different context from a defect in the center. Reflection padding can reduce the artificial step yet has its own constraint: the padding width must be smaller than the input extent. Keep crop policy, padding mode and any resize in the model artifact. Augmentation should model camera variation without moving evidence outside the label.
Inspect aliasing and detail loss
Stride two discards spatial samples. It can help cost and context, but tiny receipt marks may vanish after repeated reductions. Compare a shallow high-resolution branch against a deeper reduced branch on the small-defect slice, and inspect false negatives by mark width. Pooling is another reduction; it does not restore removed detail. When the target is a pixel mask, keep a clear map from output grid coordinates to input pixels and test odd dimensions as well as the usual square input.
Validate with a shape probe
A shape probe should include the smallest accepted crop, an odd-width crop and the maximum configured crop. Record each intermediate feature map and compare with the formula. An intentionally invalid kernel/input pair should fail before a training job begins. Put the probe in the same configuration path used by serving: a model that trains on resized 224-square crops and serves on native 47 by 83 crops has a different output geometry and potentially different edge behavior.
Implementation
from dataclasses import dataclass
import torch
from torch import nn
@dataclass(frozen=True)
class ConvPlan:
kernel: int
stride: int
padding: int
dilation: int = 1
def output_extent(input_extent: int, plan: ConvPlan) -> int:
numerator = input_extent + 2 * plan.padding - plan.dilation * (plan.kernel - 1) - 1
if numerator < 0:
raise ValueError("kernel does not fit the crop")
return numerator // plan.stride + 1
receipt_crop = torch.zeros(2, 1, 47, 83)
first_plan = ConvPlan(kernel=3, stride=2, padding=1)
edge_bank = nn.Conv2d(1, 12, kernel_size=first_plan.kernel,
stride=first_plan.stride, padding=first_plan.padding)
edge_maps = edge_bank(receipt_crop)
expected_grid = (output_extent(47, first_plan), output_extent(83, first_plan))
assert expected_grid == (24, 42)
assert edge_maps.shape == (2, 12, *expected_grid)Performance and operating cost
A standard convolution with input channels C, output channels F, kernel area K and output grid area G performs work proportional to G × C × F × K per example. Its weights occupy C × F × K parameters plus optional biases. Activations scale with batch × F × G and often dominate training memory at high resolution. A larger stride reduces G but may erase narrow evidence; measure the small-object slice before accepting that trade.
Common Mistakes
- Do not assume odd input dimensions divide evenly under a stride.
- Do not confuse output resolution with receptive-field width.
- Do not change resize or padding behavior between training and serving.
Read next
- Tensor contracts: shape, dtype, device and mask
- Image augmentation: split originals first and preserve the label
- Project: classify receipt image quality with a checked training contract
- Inference contracts: preserve preprocessing and measure tail latency
- Project: localize receipt defects with a residual CNN
Continue the workflow: Teacher feature adapters and layer alignment.
Continue the workflow: Segmentation mask alignment and ignored pixels.
Continue the workflow: CTC time lengths, blank labels and alignment feasibility.
Continue the workflow: Vision transformer patch grids and token contracts.
Continue the workflow: Video temporal convolutions and clip aggregation.
