Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-label calibration and label cardinality

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Each label can have a different probability calibration curve, while the number of true or predicted labels per case reveals a separate change in case mix.

Calibrate each condition on eligible outcomes

A score of 0.7 for water damage should be checked against observed water outcomes among comparable, fully inspected parcels. A good seal-breach reliability curve does not certify the water score. The code summarizes mean predicted risk and observed event rate by label over a fixed mature cohort. A single four-case mean is only a diagnostic; real calibration needs support across risk ranges. Probability calibration explains binning and independent calibration data.

Measure label cardinality separately

Cardinality is the average number of positive labels per case. It can change if parcels sustain more compound damage, if annotation rules expand or if model thresholds move. Compare actual cardinality on mature fully audited cases with predicted cardinality at the issued cutoffs. A changed prediction count alone is not proof of a changed real condition rate. The label schema must stay versioned.

Keep probability and action distinct

Calibrating each label’s score does not choose a review threshold or guarantee that combined case-level actions fit capacity. A model can be well calibrated yet send too many parcels to inspection when several labels trigger together. Per-label costing and review capacity turn risk into policy.

Watch selective inspection

A water label may be observed only when the parcel was already routed to review. Then observed water prevalence and apparent calibration are conditional on inspection. Keep a random audit stream or evaluate an identified fully inspected cohort; report what population the curve represents. Selection in labeling can invalidate a naive reliability chart.

Review changing co-occurrence

Even stable per-label rates can hide a rise in cases with both seal and water damage. Track common pairs with support and compare the cost of combined actions. Do not fit an elaborate dependence model from sparse pairs. Binary relevance is a sensible baseline; the release review checks the joint workload.

Implementation

python
label_order = ("seal", "water")
issued = [((0.8, 0.2), (1, 0)), ((0.6, 0.7), (1, 1)),
          ((0.3, 0.4), (0, 0)), ((0.2, 0.9), (0, 1))]
cutoffs = {"seal": 0.5, "water": 0.6}

def calibration_and_cardinality(rows, labels, thresholds):
    summary = {}
    for position, label in enumerate(labels):
        mean_risk = sum(scores[position] for scores, _ in rows) / len(rows)
        observed_rate = sum(truth[position] for _, truth in rows) / len(rows)
        summary[label] = {"mean_risk": mean_risk, "observed_rate": observed_rate}
    true_cardinality = sum(sum(truth) for _, truth in rows) / len(rows)
    predicted_cardinality = sum(
        sum(scores[position] >= thresholds[label]
            for position, label in enumerate(labels)) for scores, _ in rows
    ) / len(rows)
    return summary, true_cardinality, predicted_cardinality

calibration, actual_count, predicted_count = calibration_and_cardinality(
    issued, label_order, cutoffs)
assert actual_count == 1.0
assert predicted_count == 1.0
assert calibration["water"]["observed_rate"] == 0.5

Performance and operating cost

Scanning N mature cases and L labels costs O(NL) time and O(L) summary memory. Calibration by risk range and operating slice needs enough positive and negative outcomes in each cell, which often makes label collection more expensive than scoring.

Common Mistakes

  • Do not use a pooled calibration curve to certify every label.
  • Do not infer true cardinality from issued predictions without mature labels.
  • Do not evaluate only parcels selected for inspection and call it population calibration.

Read next

ai-data
machine-learning
Storage details