Each label can have a different probability calibration curve, while the number of true or predicted labels per case reveals a separate change in case mix.
Multi-label calibration and label cardinality
Calibrate each condition on eligible outcomes
A score of 0.7 for water damage should be checked against observed water outcomes among comparable, fully inspected parcels. A good seal-breach reliability curve does not certify the water score. The code summarizes mean predicted risk and observed event rate by label over a fixed mature cohort. A single four-case mean is only a diagnostic; real calibration needs support across risk ranges. Probability calibration explains binning and independent calibration data.
Measure label cardinality separately
Cardinality is the average number of positive labels per case. It can change if parcels sustain more compound damage, if annotation rules expand or if model thresholds move. Compare actual cardinality on mature fully audited cases with predicted cardinality at the issued cutoffs. A changed prediction count alone is not proof of a changed real condition rate. The label schema must stay versioned.
Keep probability and action distinct
Calibrating each label’s score does not choose a review threshold or guarantee that combined case-level actions fit capacity. A model can be well calibrated yet send too many parcels to inspection when several labels trigger together. Per-label costing and review capacity turn risk into policy.
Watch selective inspection
A water label may be observed only when the parcel was already routed to review. Then observed water prevalence and apparent calibration are conditional on inspection. Keep a random audit stream or evaluate an identified fully inspected cohort; report what population the curve represents. Selection in labeling can invalidate a naive reliability chart.
Review changing co-occurrence
Even stable per-label rates can hide a rise in cases with both seal and water damage. Track common pairs with support and compare the cost of combined actions. Do not fit an elaborate dependence model from sparse pairs. Binary relevance is a sensible baseline; the release review checks the joint workload.
Implementation
label_order = ("seal", "water")
issued = [((0.8, 0.2), (1, 0)), ((0.6, 0.7), (1, 1)),
((0.3, 0.4), (0, 0)), ((0.2, 0.9), (0, 1))]
cutoffs = {"seal": 0.5, "water": 0.6}
def calibration_and_cardinality(rows, labels, thresholds):
summary = {}
for position, label in enumerate(labels):
mean_risk = sum(scores[position] for scores, _ in rows) / len(rows)
observed_rate = sum(truth[position] for _, truth in rows) / len(rows)
summary[label] = {"mean_risk": mean_risk, "observed_rate": observed_rate}
true_cardinality = sum(sum(truth) for _, truth in rows) / len(rows)
predicted_cardinality = sum(
sum(scores[position] >= thresholds[label]
for position, label in enumerate(labels)) for scores, _ in rows
) / len(rows)
return summary, true_cardinality, predicted_cardinality
calibration, actual_count, predicted_count = calibration_and_cardinality(
issued, label_order, cutoffs)
assert actual_count == 1.0
assert predicted_count == 1.0
assert calibration["water"]["observed_rate"] == 0.5Performance and operating cost
Scanning N mature cases and L labels costs O(NL) time and O(L) summary memory. Calibration by risk range and operating slice needs enough positive and negative outcomes in each cell, which often makes label collection more expensive than scoring.
Common Mistakes
- Do not use a pooled calibration curve to certify every label.
- Do not infer true cardinality from issued predictions without mature labels.
- Do not evaluate only parcels selected for inspection and call it population calibration.
