Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Vision position grids, resolution changes and transfer

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Changing a pretrained image grid changes its patch-token count; spatial position vectors need a two-dimensional resize while the class-token vector stays separate.

Count old and new positions

A 3-by-4 patch grid has twelve image positions plus one class position. If image height doubles with width unchanged, the new grid is 6-by-4, giving twenty-four image positions plus one class position. Copying all thirteen old vectors into a twenty-five-token sequence fails by shape; blindly repeating vectors discards spatial meaning. Record the old and new grids, patch size, and position-embedding revision with the transferred model. The patch contract comes first.

Interpolate only the spatial part

Remove the class-token vector, reshape the remaining positions into their old two-dimensional grid, resize that grid, flatten in the original row-major order, then prepend the untouched class vector. The example performs this with a 3-by-4 to 6-by-4 resize and asserts both the result length and exact class-vector preservation. Its vectors are illustrative; a pretrained model also needs compatible patch projection and encoder weights. A one-dimensional interpolation over flattened positions can blend the end of one row with the start of the next.

Revisit the visual task after resizing

A higher input resolution changes more than positional vectors: an edge that occupied a fraction of one patch may now cross different tokens, and crop policy may expose another part of the receipt. Retrain or fine-tune under the new preprocessing and evaluate on a held-out physical-receipt set. Compare small defects, camera families and edge-of-frame cases. Staged unfreezing offers one controlled comparison; it is not a substitute for a new test.

Check class-head compatibility

A transferred encoder may have an output head for different labels. Reset or replace that head if the receipt classes change, and record which parameters were initialized versus copied. Verify class ID order, image normalization and color-channel adaptation. The presence of a loaded state dictionary does not prove that every weight was compatible or that the new head was trained. Save a fixed-image embedding and logits before and after the resize operation to catch accidental shifts.

Measure the new serving envelope

More patches raise attention cost and may invalidate a deployment latency or memory target. Compare the same batch, precision and device at the old and new image sizes, including decode and resize. If the high-resolution model misses the latency gate, test a smaller patch grid or a localized crop path before assuming an architectural win. The project makes resolution and rollback part of the release decision.

Implementation

python
import torch
from torch.nn import functional as functional

old_rows, old_columns = 3, 4
new_rows, new_columns = 6, 4
embedding_width = 8
old_positions = torch.arange((1 + old_rows * old_columns) * embedding_width,
                             dtype=torch.float32).view(1, 13, embedding_width)
class_position = old_positions[:, :1, :]
image_positions = old_positions[:, 1:, :].transpose(1, 2)
image_grid = image_positions.reshape(1, embedding_width, old_rows, old_columns)
larger_grid = functional.interpolate(image_grid, size=(new_rows, new_columns),
                                     mode="bicubic", align_corners=False)
larger_positions = torch.cat((class_position,
                              larger_grid.flatten(2).transpose(1, 2)), dim=1)
assert larger_positions.shape == (1, 25, embedding_width)
torch.testing.assert_close(larger_positions[:, :1, :], class_position)

Performance and operating cost

Resizing the position table costs O(ND) storage for N new patch positions and embedding width D, with interpolation work proportional to the new grid size. The dominant follow-on cost is the model: dense attention-map storage grows with the square of the new token count. A 3-by-4 to 6-by-4 grid changes image-token count from twelve to twenty-four, roughly quadrupling a dense image-token attention map before the class-token contribution. Validate exact shapes and run target-device memory and latency measurements.

Common Mistakes

  • Do not interpolate the class-token vector as though it were an image patch.
  • Do not flatten before spatial interpolation and blend unrelated row boundaries.
  • Do not compare resolutions under different cropping, class maps or test receipts.

Read next

ai-data
deep-learning
Storage details