Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multi-task loss balance and shared-gradient conflict

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Loss weights set the training tradeoff between tasks; opposing gradients in shared parameters can reveal an optimization conflict but do not alone prove worse generalization.

State which task may yield

A replacement false negative may trigger a missed repair, while a duration error affects scheduling. Those costs are not expressed by raw log loss and minutes error. Choose weights on development data against explicit task guardrails. A scalar total loss can fall while the high-cost task gets worse. Decision costs must be evaluated after training.

Compare gradient direction carefully

For shared encoder parameters, compute each head’s gradient before combining them. A negative dot product means the two local updates point in partly opposing directions on that batch. It may be temporary noise or a systematic conflict. The code computes cosine for small illustrative vectors; it does not implement a gradient-surgery optimizer. Repeat across batches, sites and training stages.

Control scale before interpreting conflict

A duration loss measured in minutes may have a far larger numeric magnitude than a binary log loss. Normalize targets and inspect gradient norms, not just loss numbers. Altering weights changes the optimization objective and can move one task below its accepted quality. Per-head loss establishes the denominator first.

Separate conflict from negative transfer

Even aligned training gradients can produce a joint model that generalizes worse than separate models because of limited shared capacity or label selection. Conversely, occasional negative cosine need not harm held-out quality. The decisive evidence is each task’s untouched-test result with the same target cohort and serving contract. The comparison defines that evidence.

Keep the intervention small

If one task harms the other, try a measured change: fewer shared layers, a different sampling ratio, or separate models. Recheck latency and maintainability. Do not add elaborate optimization machinery because one batch showed disagreement; diagnose labels and task compatibility first.

Implementation

python
from math import sqrt

replacement_gradient = (0.8, -0.2, 0.4)
duration_gradient = (-0.5, 0.1, -0.2)

def gradient_cosine(first, second):
    if len(first) != len(second):
        raise ValueError("shared-parameter gradients must align by shape")
    dot = sum(left * right for left, right in zip(first, second))
    first_norm = sqrt(sum(value * value for value in first))
    second_norm = sqrt(sum(value * value for value in second))
    if first_norm == 0 or second_norm == 0:
        return None
    return dot / (first_norm * second_norm)

agreement = gradient_cosine(replacement_gradient, duration_gradient)
assert agreement is not None and agreement < -0.98
assert gradient_cosine(replacement_gradient, replacement_gradient) == 1.0

Performance and operating cost

A cosine over D shared parameters costs O(D) time and O(1) scalar storage. Obtaining separate per-task gradients can require extra backward passes or retained computation graphs. More sophisticated balancing adds training cost; justify it with held-out task gains, not the diagnostic alone.

Common Mistakes

  • Do not infer negative transfer solely from one opposing-gradient batch.
  • Do not compare loss magnitudes expressed in different units as if they were priorities.
  • Do not choose weights from the final untouched test.

Read next

ai-data
machine-learning
Storage details