A box marks a rectangular extent; a mask labels pixels, and each task needs a consistent annotation policy.
Vision annotations: validate boxes, masks and reviewer agreement
Choose the representation
For unreadable receipt detection, an image-level label may be enough. Localizing glare or a cropped total needs boxes or masks, which cost more to produce and review. Do not request pixel masks when the downstream decision only needs an image-level flag. State whether coordinates are absolute pixels or normalized fractions and whether right and bottom edges are included.
Validate geometry
A box must have positive area and lie inside the declared image bounds. A mask must match the image dimensions after any resize. If an image rotates due to orientation metadata, rotate its annotations too. Preprocessing metadata should make that transform reversible for review.
Measure label variation
Two reviewers can disagree on the border of a glare patch even if both agree the receipt needs recapture. Preserve both labels on a sample and adjudicate according to a versioned rubric. Report disagreement by capture channel, script and image quality. A large dataset with inconsistent borders can teach a detector contradictory geometry.
Build a failure fixture
Use a 640-by-480 image with a valid box from x 120 to 300 and y 40 to 190. Reject an inverted box, one that extends beyond the image and one with zero area. Then change image orientation and confirm that the displayed overlay still encloses the intended region.
Implementation
def validate_box(box, image_width, image_height):
x_min, y_min, x_max, y_max = box
if image_width <= 0 or image_height <= 0:
raise ValueError("invalid image size")
if not (0 <= x_min < x_max <= image_width and 0 <= y_min < y_max <= image_height):
raise ValueError("invalid bounding box")
return (x_max - x_min) * (y_max - y_min)Performance and operating cost
Validating B boxes costs O(B) time and O(1) extra space per image. Dense masks consume O(P) storage per image before compression, where P is pixel count; choose their fidelity only when the decision benefits from it.
Common Mistakes
- Do not mix normalized and pixel coordinates in one field.
- Do not rotate an image without rotating its labels.
- Do not interpret a reviewer disagreement as model error before adjudication.
Read next
- Pixel geometry and preprocessing: preserve the mapping back to the original image
- Vision evaluation splits: group captures and test acquisition shift
- Vision decision metrics: separate localization, class errors and abstention
- Population, estimand and sampling frame: name the quantity before calculating
Continue the workflow: Label contracts: decision unit, taxonomy and review rules.
Compare representations: Vector vs raster in machine learning: geometry, pixels and feature vectors.
Continue the workflow: Bounding-box transforms and one-to-one detection matching.
