Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-task negative transfer and per-task baselines

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Negative transfer occurs when shared training harms a target relative to a credible single-task baseline under the same evaluation contract.

Make the comparison paired

For each later-period pump, compare the shared model and replacement-only baseline on the same replacement label. For completed repairs, compare the shared duration head and duration-only baseline on the same duration records. A different test population can manufacture an apparent gain or loss. Split design must stay fixed.

Name the noninferiority margins

A joint model may save serving time while losing a small amount of accuracy. Decide beforehand whether any loss is acceptable for each head and slice. Replacement recall on rare severe faults may have a zero-regression rule, while duration mean absolute error may allow a one-minute change. The example gate uses local margins; these are not universal values.

Inspect where sharing helps

A shared visual encoder could help duration prediction on rare pump families if damage appearance is informative. It could also make the replacement head rely on post-repair documentation patterns if target clocks were mixed. Break results down by camera, depot, pump family and label support. Task contracts explain the clock risk.

Measure the resource gain

Compute full-path latency and peak memory for joint and separate serving, with both heads requested and with replacement-only requests. A shared model is not automatically cheaper if it activates an expensive duration head on every request. Serving budgets use measured costs.

Retain a split option

If the duration head improves but replacement recall regresses, keep separate models or share fewer layers. Release decisions are not a vote on architectural elegance. The project holds the joint model unless all required task and operating gates pass.

Implementation

python
paired_results = {
    "replacement_recall_single": 0.88,
    "replacement_recall_joint": 0.84,
    "duration_mae_single": 13.4,
    "duration_mae_joint": 12.7,
    "joint_p95_ms": 63,
    "separate_p95_ms": 89,
}

def compare_joint_model(results, minimum_recall, max_duration_mae):
    return {
        "replacement_ok": results["replacement_recall_joint"] >= minimum_recall,
        "duration_ok": results["duration_mae_joint"] <= max_duration_mae,
        "serving_gain_ms": results["separate_p95_ms"] - results["joint_p95_ms"],
        "recall_change": results["replacement_recall_joint"] - results["replacement_recall_single"],
    }

review = compare_joint_model(paired_results, minimum_recall=0.87, max_duration_mae=14)
assert review["replacement_ok"] is False
assert review["duration_ok"] is True
assert review["serving_gain_ms"] == 26
assert round(review["recall_change"], 2) == -0.04

Performance and operating cost

The gate is O(1) once metrics exist. Obtaining paired metrics costs inference over the same held-out records for joint and single-task models, plus full-path device benchmarks. Training separate baselines increases compute but is necessary to identify whether parameter sharing helps or harms.

Common Mistakes

  • Do not claim a joint win from one improved head while another breaches its minimum.
  • Do not compare models on different head-specific test cohorts.
  • Do not infer serving savings from parameter count alone.

Read next

ai-data
machine-learning
Storage details