Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Dice, IoU, empty masks and threshold policy

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Pixel overlap scores need a declared treatment for ignored pixels, absent classes and probability thresholds before they can compare segmentation checkpoints.

Count intersection and union on valid pixels

For a chosen defect class, true positives are pixels marked defect in both prediction and reviewed mask; false positives are predicted defect on reviewed nondefect pixels; false negatives are missed reviewed defect pixels. IoU divides true positives by their union. Dice divides twice the intersection by the total predicted and reviewed positives. Exclude ignored pixels from every count, including predicted positives. Aggregate counts across images before computing a micro score if the intended unit is a pixel, and report per-image or per-class scores separately.

Declare the both-empty case

When neither prediction nor target contains a defect, the denominator of both IoU and Dice is zero. Returning one rewards easy empty images; returning zero penalizes correct absence; excluding the case changes the evaluated population. Choose the policy in advance and report how many images were both empty. The code returns no overlap score for both-empty images and keeps their count. For a defect detector, an absence false-positive rate is often more informative than assigning a perfect overlap to empty images. The project reports both.

Select thresholds away from final test

For a binary or multi-label mask, a probability threshold turns continuous output into a hard mask. Lowering it usually raises recall and false positives. Select one policy on a development group, then freeze it before the final physical-receipt holdout. A multiclass argmax has no independent per-class threshold unless the decision rule is changed. Count tiny connected defects separately: an average IoU can be high because large folds are easy while clipped-edge slivers disappear. Confidence calibration does not automatically select a pixel mask threshold.

Preserve the image-level question

A user may care whether any clipped edge is present, not its exact pixel boundary. Derive an image-level detection rule from the mask and report image-level recall and false alerts alongside overlap. Compare output after tile merging and any morphological cleanup, because those operations can remove small islands or create false regions. A model that gains Dice by expanding every defect may inflate false alerts on clean receipts. Keep the postprocessing revision and threshold with the model artifact.

Use uncertainty-aware reporting

Show raw true-positive, false-positive and false-negative counts by defect and capture device, plus the number of held-out physical receipts. Bootstrap by receipt identity rather than by pixel if an interval is needed, since pixels within one receipt are not independent examples. Do not choose a threshold on the same final test used to report its score. Mask alignment is a prerequisite; overlap numbers cannot rescue shifted annotations.

Implementation

python
def overlap_counts(target, prediction, ignored_value=255):
    if len(target) != len(prediction):
        raise ValueError("mask length mismatch")
    true_positive = false_positive = false_negative = 0
    for target_pixel, predicted_pixel in zip(target, prediction):
        if target_pixel == ignored_value:
            continue
        true_positive += target_pixel == 1 and predicted_pixel == 1
        false_positive += target_pixel != 1 and predicted_pixel == 1
        false_negative += target_pixel == 1 and predicted_pixel != 1
    denominator = true_positive + false_positive + false_negative
    if denominator == 0:
        return {"iou": None, "dice": None, "both_empty": True}
    return {"iou": true_positive / denominator,
            "dice": 2 * true_positive / (2 * true_positive + false_positive + false_negative),
            "both_empty": False}

reviewed_pixels = [0, 1, 1, 0, 255, 0, 0, 0]
predicted_pixels = [0, 1, 0, 0, 1, 0, 0, 0]
result = overlap_counts(reviewed_pixels, predicted_pixels)
assert result == {"iou": 0.5, "dice": 2 / 3, "both_empty": False}

Performance and operating cost

Counting overlap is O(HW) time and O(1) auxiliary memory for one image if pixels stream through the metric. Storing all probability maps for later threshold sweeps is O(NHW) and can be expensive; accumulate candidate-threshold counts when possible. Per-image IoU and micro-IoU answer different questions because large masks carry more pixels. Rare-class confidence intervals depend on independent receipt counts, not the much larger pixel count.

Common Mistakes

  • Do not silently assign an IoU of one to every both-empty mask.
  • Do not calculate overlap on ignored pixels or on masks with mismatched geometry.
  • Do not choose a threshold after viewing final-test performance.

Read next

ai-data
deep-learning
Storage details