Negative transfer is a target-domain regression caused by reusing a source representation or update strategy, measured against a valid target baseline on the same cases.
Negative transfer by target slice
Define the comparison before looking at results
A parcel-damage model can use a frozen pretrained encoder, a tuned encoder or target-only features. Evaluate each on the same untouched target parcels with the same label definition and action threshold. Compare absolute error counts and paired differences, not scores from different depots. The target contract prevents a label or preprocessing change from masquerading as transfer failure.
Report the harm that matters
The code counts paired mistakes for an operational baseline and a transferred candidate by camera site. The candidate improves the overall error count in the tiny illustration but worsens one site. For damage review, false negatives may cost more than extra manual inspections, so inspect those separately before deployment. Threshold metrics and decision costing determine the right report.
Avoid a false cause claim
A site gap can arise from lighting, packaging mix, label quality or sample noise. An error table does not establish which. Inspect representative failure cases with consent and blinding where possible, then test a targeted change on a later cohort. Source-task similarity is a hypothesis; target-domain paired outcomes are the evidence. Grouped shift evaluation helps keep near-duplicate parcels together.
Keep model selection out of the final slice report
If a developer tries many checkpoints and chooses the one with the smallest final-test site gap, that test is now training information. Use development slices for design and a sealed later target period for confirmation. Small slices need support counts and uncertainty; a zero miss count on a handful of photos is not a guarantee. Paired intervals can describe some uncertainty when resampling respects parcel groups.
Decide whether to transfer at all
The simplest target-only baseline may win once inference cost, errors and maintenance are included. A source model that helps one camera but hurts another may be restricted to a well-supported domain, sent to review, or rejected. The release project records that disposition explicitly.
Implementation
# Site, true damage label, baseline decision, transferred decision.
target_cases = [
("Harbor", 1, 0, 1), ("Harbor", 0, 1, 0),
("Harbor", 1, 0, 1), ("Harbor", 0, 0, 0),
("Inland", 1, 1, 0), ("Inland", 0, 0, 1),
("Inland", 1, 1, 1), ("Inland", 0, 0, 0),
]
def paired_site_errors(rows):
report = {}
for site, truth, baseline, transferred in rows:
counts = report.setdefault(site, {"support": 0, "baseline_errors": 0,
"transfer_errors": 0, "transfer_misses": 0})
counts["support"] += 1
counts["baseline_errors"] += baseline != truth
counts["transfer_errors"] += transferred != truth
counts["transfer_misses"] += truth == 1 and transferred == 0
return report
site_report = paired_site_errors(target_cases)
assert site_report["Harbor"]["baseline_errors"] == 3
assert site_report["Harbor"]["transfer_errors"] == 0
assert site_report["Inland"]["transfer_errors"] == 2Performance and operating cost
Auditing N paired target cases costs O(N) time and O(S) counters for S sites. A credible slice comparison needs new target labels and correct parcel grouping; these collection and review costs can dominate the metric calculation.
Common Mistakes
- Do not compare candidates on different target populations.
- Do not hide a costly false-negative increase inside average accuracy.
- Do not call a small subgroup gap a proven causal effect of pretraining.
Read next
- Pretrained encoder and target-task contract
- Frozen embeddings and a linear probe
- Staged fine-tuning and checkpoint selection
- Transfer learning release review project
Continue the workflow: Non-IID clients, local steps and update disagreement.
Continue the workflow: Multi-task negative transfer and per-task baselines.
