Binary relevance fits one classifier per label, giving a clear baseline while leaving relationships among labels to the shared features or a later joint model.
Binary relevance and label dependence
Build the independent baseline first
For each parcel condition, train a binary head using only cases whose annotation for that label is known. Share a consistent feature snapshot but keep positive and negative denominators separate. The code shows two fixed linear heads over inspection features; a production training run would fit each head on its own eligible rows. The schema lesson protects against treating missing labels as negatives.
Understand the independence assumption
Separate heads do not directly condition one predicted label on another predicted label. If a torn seal and water damage often co-occur, shared image features may still capture both, but the architecture alone does not encode their dependence. A classifier chain or joint model can use label relationships, yet training-time true labels and serving-time predicted labels create different inputs for downstream heads. Evaluate that mismatch rather than assuming dependence always helps.
Avoid learning a shortcut from the review process
A water-damage label might appear mostly on parcels escalated after a seal alert. A model can then learn reviewer selection rather than water damage. Audit who was inspected for each label, sample some routine parcels and compare error on a held-out fully audited subset. Selected-label workflows have the same evaluation bias.
Compare against joint training fairly
Keep parcel groups, label taxonomy, threshold-selection set and evaluation window fixed when comparing independent and joint models. Report per-label false negatives, exact-set errors and latency. A joint model can share compute but may make a rare label worse. Metric denominators reveal that regression.
Make serving cost explicit
L independent heads over a cached D-dimensional embedding require roughly O(LD) multiply-add work after encoding, while L separate full encoders would multiply the expensive stage. Monitor missing inputs and confidence by label. The frozen-probe lesson explains the representation option; action thresholds are selected separately.
Implementation
from math import exp
condition_heads = {
"seal-breach": {"weights": (1.2, -0.4), "bias": -0.2},
"water-damage": {"weights": (-0.3, 1.4), "bias": -0.3},
}
def condition_probabilities(inspection_features, heads):
probabilities = {}
for label, parameters in heads.items():
score = sum(weight * value for weight, value in
zip(parameters["weights"], inspection_features))
score += parameters["bias"]
probabilities[label] = 1 / (1 + exp(-score))
return probabilities
parcel_features = (1.8, 1.5) # Seal anomaly and moisture signal.
scores = condition_probabilities(parcel_features, condition_heads)
assert scores["seal-breach"] > 0.7
assert scores["water-damage"] > 0.7Performance and operating cost
For L labels and D input features, inference through L linear heads costs O(LD) time and O(LD) stored weights, after feature extraction. Fitting L separate models also multiplies validation work. A shared encoder can dominate latency and memory, so measure the full path rather than head arithmetic alone.
Common Mistakes
- Do not train a head on unknown labels converted to zero.
- Do not claim independent heads prove labels are independent in the world.
- Do not compare model families using different annotation-eligibility rules.
