Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Convolution output geometry and receptive fields

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A convolutional stack changes spatial resolution and the portion of the input each output can inspect; both properties belong in the model contract.

Calculate the grid before training

For one spatial axis, the output length is floor((input + 2 × padding − dilation × (kernel − 1) − 1) / stride + 1). Calculate height and width separately when kernels or strides differ. A 47 by 83 receipt crop sent through a three-wide convolution with padding one and stride two becomes 24 by 42, not 23 by 41. This floor operation matters at odd sizes. Assert the observed tensor shape against the calculation before connecting a classifier or decoder. The tensor contract also fixes channel order.

Separate resolution from context

A three-wide kernel at dilation one sees three neighboring positions in one layer. A second three-wide, stride-one layer expands the theoretical receptive field to five; a stride-two layer changes the spacing between positions used by later layers. Track two values: receptive-field width and jump, the distance in original pixels between adjacent outputs. Starting at one for both, a layer with kernel k, dilation d and stride s adds (k − 1) × d × previous jump to field width and multiplies jump by s. Large theoretical context is no guarantee that training uses distant pixels.

Choose padding for the task

Zero padding creates artificial borders. That may be acceptable for whole-image classification, but a defect next to a cropped edge can receive a different context from a defect in the center. Reflection padding can reduce the artificial step yet has its own constraint: the padding width must be smaller than the input extent. Keep crop policy, padding mode and any resize in the model artifact. Augmentation should model camera variation without moving evidence outside the label.

Inspect aliasing and detail loss

Stride two discards spatial samples. It can help cost and context, but tiny receipt marks may vanish after repeated reductions. Compare a shallow high-resolution branch against a deeper reduced branch on the small-defect slice, and inspect false negatives by mark width. Pooling is another reduction; it does not restore removed detail. When the target is a pixel mask, keep a clear map from output grid coordinates to input pixels and test odd dimensions as well as the usual square input.

Validate with a shape probe

A shape probe should include the smallest accepted crop, an odd-width crop and the maximum configured crop. Record each intermediate feature map and compare with the formula. An intentionally invalid kernel/input pair should fail before a training job begins. Put the probe in the same configuration path used by serving: a model that trains on resized 224-square crops and serves on native 47 by 83 crops has a different output geometry and potentially different edge behavior.

Implementation

python
from dataclasses import dataclass
import torch
from torch import nn

@dataclass(frozen=True)
class ConvPlan:
    kernel: int
    stride: int
    padding: int
    dilation: int = 1

def output_extent(input_extent: int, plan: ConvPlan) -> int:
    numerator = input_extent + 2 * plan.padding - plan.dilation * (plan.kernel - 1) - 1
    if numerator < 0:
        raise ValueError("kernel does not fit the crop")
    return numerator // plan.stride + 1

receipt_crop = torch.zeros(2, 1, 47, 83)
first_plan = ConvPlan(kernel=3, stride=2, padding=1)
edge_bank = nn.Conv2d(1, 12, kernel_size=first_plan.kernel,
                      stride=first_plan.stride, padding=first_plan.padding)
edge_maps = edge_bank(receipt_crop)
expected_grid = (output_extent(47, first_plan), output_extent(83, first_plan))
assert expected_grid == (24, 42)
assert edge_maps.shape == (2, 12, *expected_grid)

Performance and operating cost

A standard convolution with input channels C, output channels F, kernel area K and output grid area G performs work proportional to G × C × F × K per example. Its weights occupy C × F × K parameters plus optional biases. Activations scale with batch × F × G and often dominate training memory at high resolution. A larger stride reduces G but may erase narrow evidence; measure the small-object slice before accepting that trade.

Common Mistakes

  • Do not assume odd input dimensions divide evenly under a stride.
  • Do not confuse output resolution with receptive-field width.
  • Do not change resize or padding behavior between training and serving.

Read next

Continue the workflow: Teacher feature adapters and layer alignment.

Continue the workflow: Segmentation mask alignment and ignored pixels.

Continue the workflow: CTC time lengths, blank labels and alignment feasibility.

Continue the workflow: Vision transformer patch grids and token contracts.

Continue the workflow: Video temporal convolutions and clip aggregation.

ai-data
deep-learning
Storage details