Loss weights set the training tradeoff between tasks; opposing gradients in shared parameters can reveal an optimization conflict but do not alone prove worse generalization.
Multi-task loss balance and shared-gradient conflict
State which task may yield
A replacement false negative may trigger a missed repair, while a duration error affects scheduling. Those costs are not expressed by raw log loss and minutes error. Choose weights on development data against explicit task guardrails. A scalar total loss can fall while the high-cost task gets worse. Decision costs must be evaluated after training.
Compare gradient direction carefully
For shared encoder parameters, compute each head’s gradient before combining them. A negative dot product means the two local updates point in partly opposing directions on that batch. It may be temporary noise or a systematic conflict. The code computes cosine for small illustrative vectors; it does not implement a gradient-surgery optimizer. Repeat across batches, sites and training stages.
Control scale before interpreting conflict
A duration loss measured in minutes may have a far larger numeric magnitude than a binary log loss. Normalize targets and inspect gradient norms, not just loss numbers. Altering weights changes the optimization objective and can move one task below its accepted quality. Per-head loss establishes the denominator first.
Separate conflict from negative transfer
Even aligned training gradients can produce a joint model that generalizes worse than separate models because of limited shared capacity or label selection. Conversely, occasional negative cosine need not harm held-out quality. The decisive evidence is each task’s untouched-test result with the same target cohort and serving contract. The comparison defines that evidence.
Keep the intervention small
If one task harms the other, try a measured change: fewer shared layers, a different sampling ratio, or separate models. Recheck latency and maintainability. Do not add elaborate optimization machinery because one batch showed disagreement; diagnose labels and task compatibility first.
Implementation
from math import sqrt
replacement_gradient = (0.8, -0.2, 0.4)
duration_gradient = (-0.5, 0.1, -0.2)
def gradient_cosine(first, second):
if len(first) != len(second):
raise ValueError("shared-parameter gradients must align by shape")
dot = sum(left * right for left, right in zip(first, second))
first_norm = sqrt(sum(value * value for value in first))
second_norm = sqrt(sum(value * value for value in second))
if first_norm == 0 or second_norm == 0:
return None
return dot / (first_norm * second_norm)
agreement = gradient_cosine(replacement_gradient, duration_gradient)
assert agreement is not None and agreement < -0.98
assert gradient_cosine(replacement_gradient, replacement_gradient) == 1.0Performance and operating cost
A cosine over D shared parameters costs O(D) time and O(1) scalar storage. Obtaining separate per-task gradients can require extra backward passes or retained computation graphs. More sophisticated balancing adds training cost; justify it with held-out task gains, not the diagnostic alone.
Common Mistakes
- Do not infer negative transfer solely from one opposing-gradient batch.
- Do not compare loss magnitudes expressed in different units as if they were priorities.
- Do not choose weights from the final untouched test.
