Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-task head losses with observed-label masks

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A shared encoder can feed separate heads while each head computes loss only on records with an observed target under its own denominator.

Choose a compatible shared input

For inspection photos, an image encoder can serve a replacement classifier and a repair-duration regressor. The heads need different output units and losses; a binary log loss and an absolute-duration loss cannot be added meaningfully until their scales and priorities are defined. The code illustrates the masked per-head arithmetic, not a full neural training loop. Target contracts come first.

Normalize within each head

Divide each head’s observed loss sum by its own number of observed labels, not the total batch size. Otherwise a rare duration label becomes weaker merely because many classification-only cases were batched with it. A head with no observed labels in the batch contributes no loss and should not produce an invalid divide-by-zero. The batch sampler still needs enough duration cases over time.

Do not leak task identity

If duration is observed only after a replacement, its missingness tells the model something about the replacement outcome. Masks may guide training, but must not be supplied as a serving-time feature if unknown at inspection. Nor should a duration head be evaluated on nonreplacement pumps unless the target has been defined for them.

Keep standalone baselines

Train the replacement and duration tasks separately with the same eligible examples, encoders and splits. A shared model is justified only if it improves a declared goal—quality, latency, storage or upkeep—without breaching either task’s minimum. Negative-transfer review compares the results.

Track training and serving costs separately

A shared encoder may save inference when both outputs are requested in one call, but two heads and their losses can complicate training. If duration is rarely requested, always computing it may waste serving resources. Measure the actual request path and report head-specific output validity. Device profiling defines the measurement boundary.

Implementation

python
from math import log

batch = [
    {"replace_y": 1, "replace_p": 0.82, "replace_known": True,
     "minutes_y": 74, "minutes_pred": 69, "minutes_known": True},
    {"replace_y": 0, "replace_p": 0.18, "replace_known": True,
     "minutes_y": None, "minutes_pred": 52, "minutes_known": False},
    {"replace_y": None, "replace_p": 0.55, "replace_known": False,
     "minutes_y": None, "minutes_pred": 61, "minutes_known": False},
]

def observed_head_losses(records):
    classification = [-row["replace_y"] * log(row["replace_p"])
                      - (1 - row["replace_y"]) * log(1 - row["replace_p"])
                      for row in records if row["replace_known"]]
    duration = [abs(row["minutes_y"] - row["minutes_pred"])
                for row in records if row["minutes_known"]]
    return {"replacement": sum(classification) / len(classification) if classification else None,
            "duration_minutes": sum(duration) / len(duration) if duration else None}

head_losses = observed_head_losses(batch)
assert round(head_losses["replacement"], 3) == 0.198
assert head_losses["duration_minutes"] == 5

Performance and operating cost

Loss aggregation costs O(B) time and O(B) storage as written for batch size B; streaming sums reduce auxiliary memory to O(1). Full training adds encoder and head forward/backward passes. Inference savings depend on whether both heads run for the same request and whether feature preparation dominates latency.

Common Mistakes

  • Do not average a sparse head over records with no label for that head.
  • Do not add losses with different units without a documented weighting rule.
  • Do not expose a training-only observed-label mask as an inspection-time feature.

Read next

ai-data
machine-learning
Storage details