An expert router saves per-token computation only if routing, overflow handling and expert utilization are measured as part of the model contract.
Sparse expert routers and per-batch capacity accounting
Route tokens, not whole deployments
A mixture-of-experts layer contains several parameter sets called experts. The router scores each token and sends it to a selected subset, often one or two experts, while the surrounding attention and other layers may remain shared. Sparse activation lowers arithmetic per token relative to executing every expert, but all expert weights still have to live somewhere and token exchange across devices can dominate latency. Record the expert count, selected experts per token and router score normalization with the checkpoint. Attention masks remain a separate sequence contract.
Calculate capacity from the actual batch
Suppose 47 tokens route to four experts, one expert per token and capacity factor 1.25. A typical per-expert limit is the ceiling of 47 times 1.25 divided by four, or 15 tokens. If the router chooses the first expert for 21 tokens, six exceed its limit despite empty slots elsewhere. The code computes these counts with a deterministic router. Define whether overflow is dropped, sent to a backup expert or processed by a shared path; each choice changes output quality and cost.
Avoid an invisible order effect
A naive capacity loop admits tokens in their input order. Reordering the same batch can then alter which tokens get expert computation. Some systems prioritize stronger router scores or use a specified token-dispatch rule. Record the rule and test both original and reordered batches. A full sentence might be split across batches at inference, so a batch-dependent overflow policy can create inconsistent answers. Capacity is a serving behavior, not merely a training-memory switch.
Watch concentration before tuning loss
Count routed and accepted tokens per expert, overflow rate, entropy of router choices and accuracy by input slice. A router that sends nearly everything to one expert defeats the intended parallelism and can starve the others of gradients. Auxiliary balancing objectives may reduce concentration, yet an excessive coefficient can force artificial uniformity when tasks are genuinely uneven. The balance lesson compares utilization with task loss and device cost.
Choose a dense baseline
A dense feed-forward block with comparable active computation gives a meaningful reference. Compare quality, total parameter memory, accelerator communication, p95 latency and training stability under the same token budget. If a sparse layer wins only at a batch size unavailable in production, it has not met the serving requirement. The project tests service-event classification with bounded overflow and a rollback to its dense baseline.
Implementation
from math import ceil
def dispatch_experts(router_choices, expert_count, capacity_factor):
if expert_count < 1 or capacity_factor <= 0:
raise ValueError("expert count and capacity factor must be positive")
if not router_choices:
return [], [0] * expert_count, 0
capacity = ceil(len(router_choices) * capacity_factor / expert_count)
accepted = [0] * expert_count
assignments = []
for expert_id in router_choices:
if not 0 <= expert_id < expert_count:
raise ValueError("router selected an unknown expert")
if accepted[expert_id] < capacity:
assignments.append(expert_id)
accepted[expert_id] += 1
else:
assignments.append(None)
return assignments, accepted, capacity
choices = [0] * 21 + [1] * 9 + [2] * 9 + [3] * 8
assignments, utilization, per_expert_limit = dispatch_experts(choices, 4, 1.25)
assert per_expert_limit == 15
assert utilization == [15, 9, 9, 8]
assert assignments.count(None) == 6Performance and operating cost
Computing a selected expert for each of T tokens is O(T) routing decisions after router scores exist; scoring E experts per token adds O(TE) work and score storage unless selection is fused. Expert forward work depends on accepted tokens and active width, while parameters scale with all E experts. On multiple devices, all-to-all token exchange and imbalance can dominate. The example counts top-one capacity only and intentionally leaves overflow unprocessed so the loss is visible.
Common Mistakes
- Do not assume sparse activation means low total parameter memory.
- Do not drop overflow tokens without measuring which tasks they contain.
- Do not compare expert throughput to a dense model at a different token budget.
