Augmentation changes training inputs to represent plausible capture variation, but it must not alter the task label by accident.
Vision augmentation: preserve labels and match serving transforms
Choose plausible variation
Mild rotation, exposure change or compression can resemble receipt capture conditions. A horizontal flip makes most receipt text unnatural and can invalidate text-location labels. Start with errors observed in the held-out capture channels, not an arbitrary list of transforms. Keep original and transformed images linked so failure analysis can identify harmful augmentation.
Transform targets together
If a crop moves a glare region, its box or mask must move with it. If a crop removes the labeled region entirely, the resulting label needs a defined rule. Reject boxes that become zero area. Geometry validation should run after each transform during development.
Keep validation clean
Do not add random training augmentation to validation or final test images. Serving must use deterministic decode, orientation, resize and normalization steps. A different channel order or resize interpolation at serving time can erase training gains. Parity checks should compare sampled tensors from both paths on identical raw files.
Test one transform
Rotate or crop a synthetic image with a known box, then check the transformed coordinates and inverse mapping. Use a receipt whose total is near the edge to catch a crop that silently removes the deciding text. Also test an unmodified image: a supposedly harmless augmentation pipeline should not change its label or size unexpectedly.
Implementation
def translate_box(box, x_offset, y_offset, crop_width, crop_height):
x_min, y_min, x_max, y_max = box
shifted = (max(0, x_min - x_offset), max(0, y_min - y_offset),
min(crop_width, x_max - x_offset), min(crop_height, y_max - y_offset))
if shifted[0] >= shifted[2] or shifted[1] >= shifted[3]:
return None
return shiftedPerformance and operating cost
A box transform is O(1); pixel augmentation is O(P) per image and can multiply training I/O if variants are materialized. Generate variants during training when practical, and record random seeds plus transform versions for reproducible review.
Common Mistakes
- Do not flip text images just because the library supports it.
- Do not transform an image without its boxes or masks.
- Do not apply random augmentation to the final test partition.
