Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-label metrics and denominators

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Multi-label evaluation needs both case-level and label-level measures because one case can have several correct and incorrect flags.

Report exact-set and per-label error separately

Exact-set accuracy credits a parcel only when every predicted flag matches every known true flag. Hamming loss counts incorrect label assignments divided by assessed case-label pairs. Exact-set accuracy becomes harder as the number of labels grows; Hamming loss can look small when most labels are negative. The code computes both on a fully annotated illustrative cohort. Unknown labels require masking and adjusted denominators in a real audit.

Distinguish micro from macro F1

Micro F1 pools true positives, false positives and false negatives across labels before combining precision and recall. Macro F1 computes F1 for each label and averages them equally. A frequent seal label can dominate the pooled result while a rare safety flag fails. Display each label’s support and confusion counts alongside both summaries. Rare-event precision explains why prevalence matters.

Choose metrics from the action

A parcel with one missed critical flag may be worse than three unnecessary low-cost flags. Exact-set and F1 numbers alone do not encode that distinction. Include per-label false-negative cost, unique parcels sent to review and deadline misses. Threshold selection should use the same action contract that evaluation reports.

Handle undefined results deliberately

A label with no positive cases has undefined recall; a label with no predicted positives has undefined precision. Do not silently turn these into a perfect score. Report support and the chosen zero-division convention, or withhold a metric until there is adequate evidence. Splitting by site and camera reduces support further. Group auditing keeps denominators visible.

Protect the evaluation population

Use parcel-grouped and time-appropriate held-out data. If only difficult parcels receive full inspection, the fully annotated subset is selected and its metrics may not describe routine traffic. Record inclusion rules and a random audit sample when feasible. The project asks for both model metrics and annotation coverage.

Implementation

python
label_order = ("seal", "water", "address")
actual = [(1, 1, 0), (1, 0, 0), (0, 1, 1), (0, 0, 1)]
predicted = [(1, 0, 0), (1, 0, 0), (0, 1, 0), (1, 0, 1)]

def multilabel_report(truth_rows, predicted_rows, labels):
    if len(truth_rows) != len(predicted_rows) or not truth_rows:
        raise ValueError("aligned evaluated cases required")
    exact = sum(truth == prediction for truth, prediction in
                zip(truth_rows, predicted_rows)) / len(truth_rows)
    mistakes = sum(left != right for truth, prediction in
                   zip(truth_rows, predicted_rows) for left, right in
                   zip(truth, prediction))
    counts = {}
    for position, label in enumerate(labels):
        tp = sum(t[position] == 1 and p[position] == 1 for t, p in zip(truth_rows, predicted_rows))
        fp = sum(t[position] == 0 and p[position] == 1 for t, p in zip(truth_rows, predicted_rows))
        fn = sum(t[position] == 1 and p[position] == 0 for t, p in zip(truth_rows, predicted_rows))
        counts[label] = {"tp": tp, "fp": fp, "fn": fn, "support": tp + fn}
    pooled = {name: sum(item[name] for item in counts.values()) for name in ("tp", "fp", "fn")}
    def f1(item):
        denominator = 2 * item["tp"] + item["fp"] + item["fn"]
        return 2 * item["tp"] / denominator if denominator else None
    return {"exact_set_accuracy": exact,
            "hamming_loss": mistakes / (len(truth_rows) * len(labels)),
            "micro_f1": f1(pooled),
            "macro_f1": sum(f1(item) for item in counts.values()) / len(counts),
            "by_label": counts}

report = multilabel_report(actual, predicted, label_order)
assert report["exact_set_accuracy"] == 0.25
assert report["hamming_loss"] == 0.25
assert report["by_label"]["water"]["fn"] == 1

Performance and operating cost

Computing counts over N cases and L labels costs O(NL) time and O(L) summary memory. A dense indicator matrix costs O(NL) storage; streaming counts avoids it when issued predictions and mature labels can be joined safely. Fully audited target cases are the limiting resource.

Common Mistakes

  • Do not describe exact-set accuracy as ordinary per-label accuracy.
  • Do not report macro F1 without per-label support and undefined-rate handling.
  • Do not include unknown annotations in a negative-label denominator.

Read next

ai-data
machine-learning
Storage details