Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pixel geometry and preprocessing: preserve the mapping back to the original image

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Image preprocessing changes dimensions, color values and coordinates; the model contract must record each transformation.

Declare the coordinate frame

A detector may output boxes on a resized 320-by-320 tensor while reviewers inspect a 1280-by-960 original. A direct overlay without inverse scaling puts boxes in the wrong place. Record original width and height, crop offsets, resize scale and any padding. Box annotations need one coordinate convention such as half-open x-min, y-min, x-max, y-max.

Handle aspect ratio

Stretching a long receipt into a square changes character shapes; padding preserves aspect ratio but changes the box origin. Compare the effect on the task and reproduce exactly the same policy at serving time. A crop may remove the merchant or total even when the remaining pixels look sharp. Keep an uncropped inspection view for difficult cases.

Make color explicit

A three-channel array may mean RGB or BGR depending on the decoder. Normalization can use byte values or scaled floats; silently applying it twice shifts the input distribution. Store channel order, value range and image orientation handling in the model manifest. Training-serving parity applies to pixels as well as tabular features.

Use a geometry fixture

Take an original image 1280 pixels wide and resize it to 320 without padding. A box whose original horizontal coordinates are 160 and 480 maps to 40 and 120. Reverse the transform and confirm the original coordinates return. Test a padded portrait image separately; one scale factor is insufficient without the padding offset.

Implementation

python
def scale_box_xyxy(box, source_size, target_size):
    source_width, source_height = source_size
    target_width, target_height = target_size
    if min(source_width, source_height, target_width, target_height) <= 0:
        raise ValueError("image dimensions must be positive")
    x_min, y_min, x_max, y_max = box
    if not 0 <= x_min < x_max <= source_width or not 0 <= y_min < y_max <= source_height:
        raise ValueError("box lies outside source image")
    return (x_min * target_width / source_width, y_min * target_height / source_height,
            x_max * target_width / source_width, y_max * target_height / source_height)

Performance and operating cost

Coordinate conversion is O(1) per box; resizing an image is O(P) in pixel count P, with a new tensor often requiring O(P) memory. Retain transform metadata so annotation overlays and error reviews do not require guessing how pixels moved.

Common Mistakes

  • Do not overlay resized coordinates on an unscaled original.
  • Do not assume RGB when the decoder returns another order.
  • Do not stretch receipts without checking task-relevant distortion.

Read next

Continue the workflow: Media preprocessing: image geometry, text spans and time alignment.

ai-data
computer-vision
Storage details