Federated evaluation needs both example-weighted and site-weighted views, with denominators for mature outcomes, participation and rare failures.
Federated evaluation by site and denominator
Name the measure before aggregating
A depot-level missed-failure rate uses confirmed follow-up events as its denominator; calibration at 37 days needs mature outcomes. Do not average client accuracy values when their denominators and class prevalences differ and then call the result a network-wide example metric. Rare-event metrics show why prevalence matters.
Compute micro and macro separately
A micro average weights every evaluated pump equally. A macro average weights each depot equally. The code computes both from toy correct-prediction counts; their difference is information, not a rounding problem. Neither replaces the minimum site-specific standard when deployment happens at all depots. Group gaps should be shown next to the averages.
Evaluate an untouched period at every site
Select rounds, aggregation weights and local-step budgets using development periods. Then run the selected checkpoint on later local holdouts without sending raw records to the server. Record the model and feature versions and the local result schema. If a client has too few mature positives, report an interval or “insufficient support” rather than a precise recall.
Treat participation as part of the result
Only reporting sites are visible in the result table. Publish the number invited and evaluated, and compare reporting sites with missing ones. A high pooled score from the largest two depots does not license a rollout to six. Participation records close that denominator.
Protect the evaluation channel
Even aggregate metrics can reveal sensitive local properties when groups are small or reported repeatedly. Set minimum cell sizes, retention and access rules. This is separate from the training objective. Safe release discusses disclosure controls; the project uses the audited scorecard.
Implementation
site_results = [
{"depot": "north-47", "evaluated": 200, "correct": 180},
{"depot": "west-62", "evaluated": 50, "correct": 35},
{"depot": "east-83", "evaluated": 20, "correct": 12},
]
def accuracy_views(results):
if not results or any(row["evaluated"] <= 0 or not 0 <= row["correct"] <= row["evaluated"] for row in results):
raise ValueError("invalid local denominator")
micro = sum(row["correct"] for row in results) / sum(row["evaluated"] for row in results)
macro = sum(row["correct"] / row["evaluated"] for row in results) / len(results)
worst = min(row["correct"] / row["evaluated"] for row in results)
return micro, macro, worst
micro, macro, worst = accuracy_views(site_results)
assert round(micro, 3) == 0.841
assert round(macro, 3) == 0.733
assert worst == 0.6Performance and operating cost
Aggregating C site summaries is O(C) time and O(1) auxiliary memory. The expensive part is local holdout scoring and maturity tracking. Transferring only metrics reduces raw-data movement but does not automatically protect small cohorts; suppression and access controls may be needed.
Common Mistakes
- Do not label a mean of site accuracies as example-weighted accuracy.
- Do not infer quality at sites that did not report.
- Do not publish unstable rare-event rates from tiny local denominators.
