A multi-label target can assign several independent flags to one case, and each flag needs a distinct positive, negative or unknown annotation state.
Multi-label schema and unknown targets
Choose the unit and vocabulary
A parcel incident can involve a broken seal and water damage at the same time. Define one parcel as the unit, then version the allowed label names and their meanings. A single multiclass field would force a false choice between coexisting conditions. Multiclass evaluation applies when exactly one class is assigned; the distinction changes both training data and metrics.
Keep unknown distinct from negative
A reviewer may have inspected the seal but not the interior. The missing interior annotation is unknown, not evidence that water damage is absent. The code validates explicit zero, one and unknown states, and counts observed supervision per label. If unknown is silently filled with zero, the model learns from invented negatives. Annotation guidelines must say which inspection establishes a negative.
Treat each label as an action contract
Some flags merely route a parcel for manual review; others stop dispatch. Record the action, deadline and cost of a missed condition for every flag. Two labels may require the same action, but a model output should preserve the distinction if investigation or audit depends on it. Per-label thresholds convert scores into actions without forcing one global cutoff.
Split at the parcel or incident group
Several photos, claims and scans from one parcel must remain in one training or evaluation partition. A related batch can share the same damage event. Splitting individual image rows at random makes memorization appear to be generalization. Group and time validation defines a credible held-out cohort.
Measure annotation coverage
Report how many cases were reviewed for each label and where unknown states concentrate. A rare label with six positives and many unknowns cannot support the same evaluation claim as a common, fully audited label. Revisit sampling and adjudication before choosing a more complex model. Multi-label metrics need those denominators.
Implementation
labels = ("seal-breach", "water-damage", "address-mismatch")
parcel_annotations = [
("P941", {"seal-breach": 1, "water-damage": 0, "address-mismatch": None}),
("P942", {"seal-breach": 0, "water-damage": 1, "address-mismatch": 1}),
("P943", {"seal-breach": 1, "water-damage": None, "address-mismatch": 0}),
]
def supervision_counts(rows, vocabulary):
counts = {label: {"positive": 0, "negative": 0, "unknown": 0}
for label in vocabulary}
for parcel_id, annotation in rows:
if set(annotation) != set(vocabulary):
raise ValueError("label vocabulary changed")
for label in vocabulary:
state = annotation[label]
if state not in (0, 1, None):
raise ValueError("invalid target state")
bucket = "unknown" if state is None else ("positive" if state else "negative")
counts[label][bucket] += 1
return counts
report = supervision_counts(parcel_annotations, labels)
assert report["water-damage"] == {"positive": 1, "negative": 1, "unknown": 1}
assert report["seal-breach"]["positive"] == 2Performance and operating cost
Counting N parcels across L labels costs O(NL) time and O(L) summary memory. A dense target matrix costs O(NL) storage; sparse representations save space when most known labels are negative, but must still preserve unknown states. The expensive work is defining and auditing labels consistently.
Common Mistakes
- Do not collapse an uninspected label into a negative.
- Do not use a single multiclass category for conditions that can coexist.
- Do not split related photos or claims across train and test.
Read next
- Binary relevance and label dependence
- Per-label thresholds and action cost
- Multi-label metrics and denominators
- Multi-label parcel review project
Continue the workflow: Multi-task targets and per-head label availability.
Continue the workflow: Span-level evaluation and extraction error taxonomy.
