Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Federated averaging and client weighting

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A federated round combines compatible local model updates using declared client weights; the weights determine whose data dominates the global model.

Start from the same checkpoint

Each participating repair depot receives the same global model version and preprocessing contract, trains locally for the agreed step budget, and returns an update. Reject packets from an older checkpoint or mismatched tensor shape. Averaging weights from different starting versions does not represent the intended round. The client contract defines those versions.

Compute an explicit weighted delta

A common baseline multiplies each client delta by its eligible example count, sums the results, then divides by total eligible examples. The code uses a two-parameter toy model so the weighting is inspectable. It is not a full optimizer. Equal-client weighting produces a different answer, and a capped count can prevent one large site from overwhelming small sites if that matches the objective.

Distinguish participation from population

Only clients that finish a round contribute to that round. If slow depots systematically drop out, the weighted average may drift toward fast, well-connected sites. Log invited, started, completed and rejected clients with reasons; a large total example count does not prove geographic coverage. Participation audit checks that selection.

Constrain update scale

Local training can produce a large delta from data shift, a bad sensor conversion, excessive local epochs or malicious input. Inspect norms and nonfinite values before aggregation. Clipping can bound influence but changes optimization; it is not a substitute for secure aggregation or differential privacy. Contribution bounds have their own privacy meaning.

Evaluate the result, not the averaging step

A numerically valid average may worsen rare-failure recall at one depot. Compare local and global validation by site, and use a later holdout after choosing rounds. Site evaluation makes the outcome visible; the project converts it into a release decision.

Implementation

python
client_updates = [
    {"depot": "north-47", "base": "round-8", "examples": 180, "delta": (0.20, -0.04)},
    {"depot": "west-62", "base": "round-8", "examples": 90, "delta": (-0.10, 0.08)},
    {"depot": "east-83", "base": "round-8", "examples": 30, "delta": (0.50, -0.02)},
]

def weighted_delta(updates, expected_base):
    if not updates or any(update["base"] != expected_base or update["examples"] <= 0 for update in updates):
        raise ValueError("incompatible client update")
    dimensions = len(updates[0]["delta"])
    if any(len(update["delta"]) != dimensions for update in updates):
        raise ValueError("parameter shapes differ")
    total = sum(update["examples"] for update in updates)
    return tuple(sum(update["examples"] * update["delta"][dimension] for update in updates) / total
                 for dimension in range(dimensions))

round_delta = weighted_delta(client_updates, "round-8")
assert tuple(round(value, 3) for value in round_delta) == (0.14, -0.002)

Performance and operating cost

Aggregating C clients with D parameters costs O(C times D) arithmetic and O(D) accumulator memory if updates stream securely into the server. Upload traffic is roughly proportional to participating clients times model size before compression or protocol overhead. This toy code omits transport, secure aggregation, training and adversarial checks.

Common Mistakes

  • Do not average updates from different starting checkpoints.
  • Do not hide the choice between example-weighted and equal-client objectives.
  • Do not treat clipping alone as a privacy or poisoning solution.

Read next

ai-data
machine-learning
Storage details