Image preprocessing changes dimensions, color values and coordinates; the model contract must record each transformation.
Pixel geometry and preprocessing: preserve the mapping back to the original image
Declare the coordinate frame
A detector may output boxes on a resized 320-by-320 tensor while reviewers inspect a 1280-by-960 original. A direct overlay without inverse scaling puts boxes in the wrong place. Record original width and height, crop offsets, resize scale and any padding. Box annotations need one coordinate convention such as half-open x-min, y-min, x-max, y-max.
Handle aspect ratio
Stretching a long receipt into a square changes character shapes; padding preserves aspect ratio but changes the box origin. Compare the effect on the task and reproduce exactly the same policy at serving time. A crop may remove the merchant or total even when the remaining pixels look sharp. Keep an uncropped inspection view for difficult cases.
Make color explicit
A three-channel array may mean RGB or BGR depending on the decoder. Normalization can use byte values or scaled floats; silently applying it twice shifts the input distribution. Store channel order, value range and image orientation handling in the model manifest. Training-serving parity applies to pixels as well as tabular features.
Use a geometry fixture
Take an original image 1280 pixels wide and resize it to 320 without padding. A box whose original horizontal coordinates are 160 and 480 maps to 40 and 120. Reverse the transform and confirm the original coordinates return. Test a padded portrait image separately; one scale factor is insufficient without the padding offset.
Implementation
def scale_box_xyxy(box, source_size, target_size):
source_width, source_height = source_size
target_width, target_height = target_size
if min(source_width, source_height, target_width, target_height) <= 0:
raise ValueError("image dimensions must be positive")
x_min, y_min, x_max, y_max = box
if not 0 <= x_min < x_max <= source_width or not 0 <= y_min < y_max <= source_height:
raise ValueError("box lies outside source image")
return (x_min * target_width / source_width, y_min * target_height / source_height,
x_max * target_width / source_width, y_max * target_height / source_height)Performance and operating cost
Coordinate conversion is O(1) per box; resizing an image is O(P) in pixel count P, with a new tensor often requiring O(P) memory. Retain transform metadata so annotation overlays and error reviews do not require guessing how pixels moved.
Common Mistakes
- Do not overlay resized coordinates on an unscaled original.
- Do not assume RGB when the decoder returns another order.
- Do not stretch receipts without checking task-relevant distortion.
Read next
- Vision annotations: validate boxes, masks and reviewer agreement
- Vision augmentation: preserve labels and match serving transforms
- Training-serving parity: compare feature values at one prediction clock
- Axis scales, baselines and labels: prevent a correct number from telling a false story
Continue the workflow: Media preprocessing: image geometry, text spans and time alignment.
