Duplicate boxes should be suppressed within the correct class after a declared score filter; a suppression rule can also erase nearby real objects.
Class-aware suppression and detection score thresholds
Filter before suppression
A detector may emit many candidate boxes with objectness and class scores. Define exactly which score is ranked, whether class scores are calibrated and which low-score candidates are removed before suppression. A threshold chosen on the final test set leaks evaluation into the policy. Retain a development split with whole-receipt identities and freeze the threshold there. The code uses already scored detections and does not claim to compute neural logits. Calibration is separate from box overlap.
Suppress within a class
Non-maximum suppression selects the highest-scored box, then removes lower-scored same-class boxes whose overlap exceeds a fixed IoU threshold. A total-field box and a date-field box may overlap without being duplicates, so class-agnostic suppression can discard a real field. The example keeps a same-location box of another class while dropping a near-duplicate total box. Verify that input boxes have valid geometry and that scores are finite before sorting.
Know when overlap is not duplication
Two adjacent line-item fields can share border pixels, especially after a coarse resize. A low suppression threshold may erase one. A high threshold may leave several copies of the same total. Review both duplicate count and missed-object count by class and object size. For crowded forms, consider a different suppression or assignment policy, but show the failure cases before adding complexity. One-to-one matching still governs evaluation after suppression.
Keep deterministic ordering
When two boxes have equal scores, a stable tie rule makes replay and debugging possible. The implementation sorts by score then original index and emits retained original indices. Production libraries may have backend-specific tie behavior, so test the exact deployed operator if audit reproducibility matters. Apply the same cutoff, coordinate clipping and suppression to validation and serving. A change in these policies is a model-release change even when weights are unchanged.
Measure the complete postprocessor
Record candidate count before and after score filtering, boxes after suppression, classwise false alerts and p95 postprocessing time. A tiny model can still have slow serving when thousands of dense candidates require pairwise comparisons. Compare detection quality at the frozen score and IoU settings on held-out receipts. The applied project packages both thresholds with the model and retains a rollback route.
Implementation
from math import isfinite
def overlap(first, second):
width = max(0, min(first[2], second[2]) - max(first[0], second[0]))
height = max(0, min(first[3], second[3]) - max(first[1], second[1]))
shared = width * height
first_area = (first[2] - first[0]) * (first[3] - first[1])
second_area = (second[2] - second[0]) * (second[3] - second[1])
return shared / (first_area + second_area - shared)
def class_aware_nms(detections, score_floor, overlap_limit):
for box, score, label in detections:
if not (box[0] < box[2] and box[1] < box[3] and isfinite(score)):
raise ValueError("invalid detection")
candidates = sorted((index for index, item in enumerate(detections)
if item[1] >= score_floor),
key=lambda index: (-detections[index][1], index))
retained = []
for index in candidates:
box, _, label = detections[index]
if all(detections[kept][2] != label or
overlap(box, detections[kept][0]) <= overlap_limit
for kept in retained):
retained.append(index)
return retained
receipt_fields = [((10, 10, 30, 30), 0.93, "total"),
((12, 12, 31, 31), 0.81, "total"),
((10, 10, 30, 30), 0.72, "date")]
assert class_aware_nms(receipt_fields, 0.4, 0.5) == [0, 2]Performance and operating cost
Sorting N candidates costs O(N log N); the simple retained-box scan can cost O(N²) overlap checks in the worst case and O(N) output storage. The code is an audit-sized reference, not a high-throughput GPU postprocessor. Score filtering can reduce N substantially, but an aggressive cutoff may remove rare small fields. Profile candidate counts and deployed operator latency as well as model forward time.
Common Mistakes
- Do not suppress boxes from different classes unless the task explicitly requires that.
- Do not tune score or IoU cutoffs on the final test set.
- Do not treat overlapping nearby objects as automatic duplicates.
