When client distributions differ, additional local steps can move models toward site-specific objectives and make their returned updates disagree.
Non-IID clients, local steps and update disagreement
Measure heterogeneity before tuning
Repair depots may handle different pump types, inspection cameras and maintenance schedules. Compare matured outcome prevalence, feature ranges and model errors by depot. A low pooled loss can coexist with a high missed-failure rate at a smaller site. Group error gaps make that failure visible.
Read disagreement with context
Two client deltas pointing in opposite directions may signal data heterogeneity, label inconsistency, or noisy small samples. The code computes cosine agreement for toy updates; it is a diagnostic, not proof of a particular cause. Pair it with contract checks, eligible sample counts and held-out site metrics. The shared target must be checked first.
Control the local work budget
More local epochs reduce communication rounds but can increase drift and overfitting at small depots. Compare one, three and five local passes under a fixed total compute or communication budget, keeping sampling and evaluation cohorts constant. Select the schedule on development sites, then report untouched later-period performance. Do not tune per site on its test records.
Separate personalization from one global model
If one global model cannot meet each depot’s minimum performance, a calibrated local head or site-specific threshold may help. That adds serving, monitoring and rollback complexity. Define whether local adaptation uses private training data or leaked later outcomes. Threshold choice can sometimes address action costs without retraining the representation.
Consider who is absent
Update disagreement among successful clients says nothing about clients that failed to participate. Slow links, old devices or low-volume depots may be exactly where the model needs improvement. The participation audit must run alongside drift diagnostics.
Implementation
from math import sqrt
depot_deltas = {
"north-47": (0.24, -0.04),
"west-62": (0.17, -0.02),
"east-83": (-0.19, 0.03),
}
def cosine_agreement(first, second):
numerator = sum(left * right for left, right in zip(first, second))
first_norm = sqrt(sum(value * value for value in first))
second_norm = sqrt(sum(value * value for value in second))
if first_norm == 0 or second_norm == 0:
raise ValueError("zero update needs separate handling")
return numerator / (first_norm * second_norm)
assert cosine_agreement(depot_deltas["north-47"], depot_deltas["west-62"]) > 0.99
assert cosine_agreement(depot_deltas["north-47"], depot_deltas["east-83"]) < -0.99Performance and operating cost
One pairwise cosine costs O(D) time for D parameters. All C-by-C client comparisons cost O(C squared times D) time; sampled comparisons or update summaries are cheaper. Extra local epochs increase device compute and can reduce rounds while worsening drift. Choose settings on site-level validation rather than minimizing communication alone.
Common Mistakes
- Do not infer the cause of opposing updates from cosine alone.
- Do not improve the pooled metric by sacrificing a small high-risk depot.
- Do not compare local-epoch settings under different participation cohorts without reporting the change.
Read next
- Federated client data and target contract
- Federated averaging and client weighting
- Client participation, dropouts and stale updates
- Federated evaluation by site and denominator
Continue the workflow: Multi-task loss balance and shared-gradient conflict.
