A routed model needs both predictive checks and systems checks: expert collapse, token overflow, device skew and inference-time batching can each undo the expected gain.
Expert load balance, overflow slices and dispatch consistency
Measure offered and accepted load
Offered load counts the router’s first-choice tokens. Accepted load counts tokens an expert actually processes after capacity is enforced. Reporting only the latter can hide an expert receiving many requests and silently dropping the excess. Keep per-expert histograms by batch, input category and sequence position, then inspect whether rare service-event types are overrepresented in overflow. The router lesson provides the capacity arithmetic.
Balance without forcing sameness
A balance term can encourage traffic across experts, but uniform use is a proxy for healthy training, not the objective itself. Compare task loss and expert utilization as the coefficient changes. Expert collapse can leave some weights nearly untrained; too much balancing can send a token to an unhelpful expert. Freeze a few reviewed examples and trace router choices through training. An auxiliary loss needs its own unit, weighting and checkpoint metadata.
Make dispatch deterministic where required
For a fixed checkpoint and token batch, router scores, tie handling and overflow priority should yield a repeatable output within numeric tolerance. Random token shuffling during training can reduce order bias but must not be an unexplained inference dependency. Test a single input alone, in a crowded batch and after changing batch order. When output changes because capacity was exceeded, document it as a model limitation rather than calling it harmless batching variation.
Account for communication
With experts spread over devices, tokens must travel to the device holding each expert and their outputs must return. A model may reduce multiply-adds while increasing transfer volume and synchronization. Profile dispatch bytes, all-to-all time, expert idle time, peak memory and p95 request latency at the actual serving mix. Compare with a dense model at the same accuracy target. Distributed checkpoints must also preserve the expert-to-device mapping or a portable remap.
Gate using both quality and operation
Report accuracy by event class, overflow by event class, accepted load skew and full-path latency. A rollout gate can require no class with rising overflow, bounded expert skew, and a meaningful gain over a dense baseline at the same cost envelope. The applied project uses these conditions before a shadow release. If a gate fails, use the dense model; adding experts blindly can amplify the failure.
Implementation
from collections import Counter
def audit_dispatch(offered, accepted, expert_count):
if len(offered) != len(accepted):
raise ValueError("offered and accepted streams must align")
if any(expert not in range(expert_count) for expert in offered):
raise ValueError("unknown offered expert")
if any(expert is not None and expert not in range(expert_count)
for expert in accepted):
raise ValueError("unknown accepted expert")
offered_counts = Counter(offered)
accepted_counts = Counter(expert for expert in accepted if expert is not None)
return [(offered_counts[index], accepted_counts[index])
for index in range(expert_count)]
offered = [0, 0, 0, 0, 1, 1, 2, 3]
accepted = [0, 0, 0, None, 1, 1, 2, 3]
counts = audit_dispatch(offered, accepted, expert_count=4)
assert counts == [(4, 3), (2, 2), (1, 1), (1, 1)]
assert sum(requested - processed for requested, processed in counts) == 1Performance and operating cost
The audit traverses T routing records in O(T + E) time and O(E) counter space for E experts. A production trace may add token IDs, sequence positions and device IDs, increasing storage linearly with T. Dense baselines avoid expert all-to-all exchange but execute more active weights per token; sparse performance therefore depends on expert placement, batch composition and overflow rules rather than parameter count alone.
Common Mistakes
- Do not report accepted load without offered load and overflow.
- Do not interpret uniform routing as proof of useful expert specialization.
- Do not omit batch composition from latency and output-parity tests.
