A multi-task release decision joins per-head label support, separate baselines, error guardrails and measured shared-serving cost.
Project: release review for shared pump-inspection heads
Freeze both targets
The depot system predicts seal replacement and repair duration from inspection evidence. Record the replacement horizon, duration population, prediction clock, observed-label masks and exclusion rules. Keep photos of one physical pump in one split and reserve later inspections for the final test. Target availability is the input gate.
Train three comparable candidates
Use one replacement-only model, one duration-only model and a shared-encoder two-head model. Keep preprocessing and eligible records matched. Tune task weights and shared-layer depth on development data only. Log masked loss denominators and gradient diagnostics but select by held-out task quality. Head loss and gradient conflict support diagnosis.
Score each task and slice
For replacement, report severe-fault recall, calibration and false review volume. For duration, report absolute error among completed repairs and the fraction without a duration label. Include depot, camera and pump-family slices with sample counts. Compare each head against its single-task counterpart on paired cases. Negative-transfer review supplies the comparison.
Measure the real serving path
Benchmark p95 time, peak memory and fallback behavior for replacement-only calls and calls that request both outputs. Ensure a duration prediction is never presented for a pump outside its target population without a defined interpretation. Keep separate-model artifacts ready if the shared candidate breaches a critical head guardrail.
Make the release call
The sample packet below holds a shared model with a replacement recall regression despite better duration error and latency. A shadow test can verify logging before action, but it cannot override the predeclared safety minimum. Preserve versions and rollback steps.
Implementation
release_packet = {
"replacement_labels_mature": True,
"duration_mask_audited": True,
"replacement_recall": 0.84,
"minimum_replacement_recall": 0.87,
"duration_mae_minutes": 12.7,
"maximum_duration_mae_minutes": 14.0,
"p95_full_path_ms": 63,
"maximum_p95_ms": 78,
"separate_baselines_kept": True,
"rollback_tested": True,
}
def multitask_release_gate(packet):
holds = []
if not packet["replacement_labels_mature"] or not packet["duration_mask_audited"]:
holds.append("label contract")
if packet["replacement_recall"] < packet["minimum_replacement_recall"]:
holds.append("replacement recall")
if packet["duration_mae_minutes"] > packet["maximum_duration_mae_minutes"]:
holds.append("duration error")
if packet["p95_full_path_ms"] > packet["maximum_p95_ms"] or not packet["separate_baselines_kept"] or not packet["rollback_tested"]:
holds.append("serving or rollback")
return "hold: " + ", ".join(holds) if holds else "eligible for shadow review"
assert multitask_release_gate(release_packet) == "hold: replacement recall"Performance and operating cost
The gate is O(1). Its inputs require three model trainings, mature head-specific labels, paired held-out scoring and device benchmarks. A shared encoder may lower inference cost when both heads run, but training and monitoring can become more complex. The example thresholds are local operating rules, not general standards.
Common Mistakes
- Do not infer that the shared model is better because total loss fell.
- Do not present a duration head outside its defined population.
- Do not release after a critical head regresses merely because another head or latency improves.
