A detection score depends on box coordinates, image transforms and a one-prediction-to-one-object matching rule before any average precision calculation.
Bounding-box transforms and one-to-one detection matching
Declare the coordinate system
Represent a box by left, top, right and bottom edges in one image coordinate space. This page uses continuous edge coordinates and area equal to width times height, with right and bottom edges excluded from the interior. Pixel-index-inclusive conventions require a different area rule. Record whether values are original-pixel, resized-pixel or normalized coordinates. Reject negative widths, out-of-bounds values and nonfinite coordinates before matching. Annotation geometry is part of the dataset contract.
Apply every image transform to boxes
A resize from 32-by-48 to 64-by-96 doubles each x and y edge. A crop subtracts its top-left offset, clips the remaining box to the crop and may remove the object entirely. Horizontal flips reverse x edges relative to the current image width. Do not transform an image and then compare predicted boxes in resized coordinates against annotations left in the original coordinate space. Save image overlays after each transform. Mask alignment follows a related but pixel-level rule.
Compute overlap on valid geometry
Intersection-over-union divides shared area by the union area. A pair of boxes with no overlap has zero IoU; a degenerate zero-area box should be rejected rather than awarded a score. The example uses independently chosen edges and checks the resulting overlap. A score near a matching threshold can flip with one annotation pixel, so report performance across agreed thresholds and by object size, especially tiny clipped corners.
Match once per reviewed object
Within each image and class, sort detections by confidence. A prediction can claim at most one unmatched reviewed box above the IoU threshold. Once a reviewed box is claimed, another overlapping prediction is a false positive, even if its IoU is high. An ignored region needs a separate rule; never count unlabeled areas as confirmed negatives by default. Matching across different images or classes can create fictitious true positives. Suppression reduces duplicates before this evaluation.
Keep metrics tied to the task
Average precision integrates precision over recall as a confidence threshold moves; it is not the accuracy of one fixed threshold. Report the chosen IoU matching threshold, class aggregation, ignored-region policy and score cutoff. For receipt fields, also report whether all required fields were found on a whole receipt and the count of duplicate boxes a reviewer must dismiss. The project gates those operational outcomes.
Implementation
from math import isclose
def box_iou(first, second):
for box in (first, second):
left, top, right, bottom = box
if not (left < right and top < bottom):
raise ValueError("box must have positive width and height")
overlap_width = max(0, min(first[2], second[2]) - max(first[0], second[0]))
overlap_height = max(0, min(first[3], second[3]) - max(first[1], second[1]))
intersection = overlap_width * overlap_height
first_area = (first[2] - first[0]) * (first[3] - first[1])
second_area = (second[2] - second[0]) * (second[3] - second[1])
return intersection / (first_area + second_area - intersection)
reviewed_total_box = (13, 17, 37, 41)
predicted_total_box = (19, 20, 43, 44)
assert isclose(box_iou(reviewed_total_box, predicted_total_box), 378 / 774)Performance and operating cost
A single box-pair IoU is O(1) time and space. Comparing P predictions with G reviewed boxes is O(PG) overlap work before matching, while sorting P detections by confidence is O(P log P). For large image batches, dense pairwise IoU tables occupy O(PG) memory; stream per image and class when practical. Matching policy can change reported precision without changing model weights, so version it with the evaluation output.
Common Mistakes
- Do not mix original and resized box coordinates.
- Do not match two predictions to the same reviewed object.
- Do not use one IoU area convention in training and another in evaluation.
