Vision evaluation depends on the task: image classification, object localization and segmentation count different errors.
Vision decision metrics: separate localization, class errors and abstention
Match metric to action
A receipt recapture gate needs false-negative and false-positive rates at a chosen threshold, because missed unreadable images delay review while unnecessary recaptures frustrate users. A detector that marks glare boxes also needs a localization match rule, such as intersection over union, before precision or recall is counted. Decision cost should set the operating point.
Define a box match
Intersection over union is overlap area divided by union area. State the coordinate convention and threshold; a prediction can have the right class but miss the affected region. Match predictions to ground truth one-to-one so duplicate boxes do not earn duplicate true positives. An empty prediction set or an image with no target boxes needs an explicit scoring policy.
Slice the errors
Report results by device, lighting, script and receipt format, not only one aggregate. Sparse slices need counts and intervals. A system that performs well on bright English receipts can fail on darker captures while the overall score barely moves. Acquisition shift is a deployment risk, not merely a test-set detail.
Use a geometry fixture
Two 20-by-20 boxes overlapping over a 10-by-20 region have intersection area 200 and union 600, so IoU is one third. Verify that disjoint boxes return zero. Then apply the recapture threshold to image-level probabilities separately; an IoU cutoff and a class-probability cutoff do different jobs.
Implementation
def box_iou(first, second):
left = max(first[0], second[0])
top = max(first[1], second[1])
right = min(first[2], second[2])
bottom = min(first[3], second[3])
intersection = max(0, right - left) * max(0, bottom - top)
first_area = (first[2] - first[0]) * (first[3] - first[1])
second_area = (second[2] - second[0]) * (second[3] - second[1])
if first_area <= 0 or second_area <= 0:
raise ValueError("boxes need positive area")
return intersection / (first_area + second_area - intersection)Performance and operating cost
One IoU calculation is O(1); comparing P predicted boxes with G ground-truth boxes naively costs O(PG) per image. One-to-one matching and per-slice metrics add work, but prevent a high aggregate score from concealing duplicate detections or capture-channel failures.
Common Mistakes
- Do not use image classification accuracy to claim good box localization.
- Do not count duplicate detections as multiple true positives.
- Do not choose the decision threshold on the final holdout.
Read next
- Vision annotations: validate boxes, masks and reviewer agreement
- Project: audit a receipt-image recapture gate
- Decision thresholds: choose an action from probabilities and error costs
- Uncertainty on charts: show denominators and intervals beside estimates
Continue the workflow: Bounding-box transforms and one-to-one detection matching.
