Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-task targets and per-head label availability

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A multi-task model shares some representation across distinct prediction targets, but every head needs its own label definition, maturity clock and missing-label mask.

Define each decision separately

A pump-inspection model may predict whether a seal needs replacement and estimate repair minutes. The first target is a binary decision; the second is a duration observed only after a completed job. They differ from a multi-label classifier that emits several tags under one label scheme. A replacement label can be known while duration is missing, and a missing duration is not zero minutes. Unknown targets require an explicit state.

Name the serving moment

Both heads should use the same inspection-time inputs if they share an encoder. Repair minutes recorded after dispatch cannot enter either head’s input. If the duration prediction is requested only after replacement approval, that later decision may justify a different feature clock and perhaps a separate model. Prediction-time availability decides whether sharing is legitimate.

Keep masks with labels

Store a boolean observed flag per task and per record. A mask of false means no training loss for that head on that record; it does not say the outcome was negative. The code builds a small status packet and rejects a duration observed without a positive, completed repair under this local data contract. Real systems may define different duration populations, so validate that rule first.

Split by physical pump and time

Multiple photographs of one pump and its later work order belong to one split. For a future release, hold out later inspections and wait until target windows mature. Do not put an early photo in training and a later view in test. Identity and time splits protect both heads.

Report support per task

A model may have 8,000 replacement labels but only 470 repair durations. State each denominator by depot, camera and pump family, along with the intersection having both labels. The intersection may be a selective subset of difficult jobs, creating bias in the shared representation. Masked loss uses those counts.

Implementation

python
inspection_cases = [
    {"pump": "station-47", "replace": 1, "replace_observed": True,
     "repair_minutes": 74, "duration_observed": True},
    {"pump": "station-62", "replace": 0, "replace_observed": True,
     "repair_minutes": None, "duration_observed": False},
    {"pump": "station-83", "replace": None, "replace_observed": False,
     "repair_minutes": None, "duration_observed": False},
]

def task_support(cases):
    for case in cases:
        if case["duration_observed"] and (case["replace"] != 1 or case["repair_minutes"] is None):
            raise ValueError("duration outside completed replacement contract")
    return {"replacement": sum(case["replace_observed"] for case in cases),
            "duration": sum(case["duration_observed"] for case in cases)}

assert task_support(inspection_cases) == {"replacement": 2, "duration": 1}

Performance and operating cost

Checking N records costs O(N) time and O(1) auxiliary memory. Storing separate masks and label provenance increases data work but prevents silent zero-imputation. Shared inference may save encoder compute, yet label maturity and outcome adjudication remain per-head costs.

Common Mistakes

  • Do not convert missing duration to zero.
  • Do not assume both tasks share a prediction clock just because they share an image.
  • Do not report one training-row count for heads with different label coverage.

Read next

ai-data
machine-learning
Storage details