Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Federated client data and target contract

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Federated training keeps examples at participating sites while exchanging model updates; it still needs a common label, feature and prediction-time contract.

Define the unit that stays local

A network of repair depots wants to predict which inspected pumps will need follow-up within 37 days. Photos and work orders stay inside each depot. A client is one depot, not an arbitrary training shard. The server can coordinate rounds and receive approved update packets, but that arrangement does not by itself prevent information leakage from those packets. Aggregation risk is a separate security decision.

Align labels before averaging models

Every depot must use the same event definition, observation window, censoring rule and exclusion policy. A work order reopened after 39 days is not a positive for the 37-day target, and a pump observed for only 12 days is not a verified negative. If local labels disagree, averaging updates cannot create a coherent global objective. Feature availability also needs one shared clock.

Version the feature transformation

Image crop, sensor units, missing-value codes and category mapping must be compatible across clients. A local calibration step may be acceptable if it is part of the serving contract, but silently fitting one global scaler on a central sample defeats the claimed data boundary. Record the preprocessing version with each update. Fold-local preprocessing offers the analogous single-site discipline.

State the population objective

Example-weighted averaging tends toward the distribution represented by the participating examples. Equal-client averaging gives a tiny depot the same influence as a large one. Neither is automatically fair or operationally correct. Define whether the target is a random pump, a random depot, or minimum acceptable quality at each depot before selecting weights. Aggregation turns this into arithmetic.

Keep local and global holdouts

Reserve a later-period test at each depot, untouched by round selection. Include newly added depots in a separate portability assessment if deployment will expand. Report depot counts, example counts, positive counts and maturity by split; a pooled metric alone hides whether one depot fails. Site evaluation specifies the denominators.

Implementation

python
depot_contracts = [
    {"depot": "north-47", "target_days": 37, "sensor_unit": "celsius", "crop_version": "v4", "matured": 182},
    {"depot": "west-62", "target_days": 37, "sensor_unit": "celsius", "crop_version": "v4", "matured": 94},
    {"depot": "east-83", "target_days": 37, "sensor_unit": "celsius", "crop_version": "v4", "matured": 57},
]

def validate_client_contracts(contracts):
    if not contracts:
        raise ValueError("at least one depot required")
    expected = (contracts[0]["target_days"], contracts[0]["sensor_unit"], contracts[0]["crop_version"])
    for depot in contracts:
        actual = (depot["target_days"], depot["sensor_unit"], depot["crop_version"])
        if actual != expected or depot["matured"] <= 0:
            raise ValueError("incompatible or empty depot contract")
    return sum(depot["matured"] for depot in contracts)

assert validate_client_contracts(depot_contracts) == 333

Performance and operating cost

Contract validation is O(C) time for C clients and O(1) auxiliary memory. The real cost lies in local feature preparation, outcome maturation and version coordination. Moving only updates can reduce raw-data transfer, but it does not remove communication, security review or the need for representative evaluation.

Common Mistakes

  • Do not call a short-observed pump a verified negative for a longer horizon.
  • Do not aggregate updates produced from incompatible feature units or label definitions.
  • Do not equate local data residence with a formal privacy guarantee.

Read next

ai-data
machine-learning
Storage details