Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: test sparse experts for service-event classification

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Compare a small expert layer with a dense classifier on the same time-split event data, then gate on class quality, overflow, device cost and batch parity.

Prepare the event ledger

Collect service-event text and structured attributes with event time, source system, class label and incident group. Split by incident and time before constructing token batches. Remove identifiers that encode the answer or appear after the classification decision. Include rare incident classes and quiet periods; a sparse model might appear strong on frequent routine events while discarding the very class that matters to operators. The target boundary applies to event sequences as well as language generation.

Set a dense reference

Train a compact shared encoder and dense classifier with the same input tokens, optimizer budget and held-out incidents. Its class recall, calibration, memory and p95 latency form the baseline. The snippet routes four synthetic token features through two experts without a capacity limit; it verifies the tensor path and finite loss only. The full project must add capacity accounting and overflow handling before making a performance claim. The router lesson gives the operational contract.

Inspect expert behavior

Track first-choice and accepted tokens per expert, overflow by event class and how often experts receive gradients. Compare alone-versus-batched predictions for reviewed incidents at serving batch sizes. A load-balancing term can move traffic while worsening rare-class recall, so publish both curves. The balance lesson ties the histograms to the release gate.

Profile the real path

Include tokenization, routing, expert dispatch, output gathering and final thresholding. If experts occupy separate accelerators, measure communication bytes and idle time; if they share one device, measure total resident weights. Repeat with the intended concurrency and sequence-length distribution. Compare accuracy at a fixed compute and latency envelope, not only parameter totals. A model that misses the deadline for long events is not an operational improvement.

Release only a bounded shadow

Score incoming events without changing routing decisions, then review disagreements and overflow. Require no deterioration in severe-incident recall, a published overflow ceiling for every class, stable batching parity and an agreed p95 latency. Package the router, experts, tokenization, capacity and fallback as one versioned artifact. Retain the dense model for immediate rollback; record why any event took a fallback path.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
event_features = torch.rand(4, 6)
reviewed_classes = torch.tensor([0, 1, 2, 1])
router = nn.Linear(6, 2)
experts = nn.ModuleList([nn.Linear(6, 8), nn.Linear(6, 8)])
classifier = nn.Linear(8, 3)
route_logits = router(event_features)
expert_choice = route_logits.argmax(dim=1)
expert_outputs = torch.stack([expert(event_features) for expert in experts], dim=1)
selected = expert_outputs[torch.arange(len(event_features)), expert_choice]
class_logits = classifier(torch.relu(selected))
loss = functional.cross_entropy(class_logits, reviewed_classes)
loss.backward()
assert class_logits.shape == (4, 3) and torch.isfinite(loss)

Performance and operating cost

This tiny demonstration evaluates every expert and then selects one, so its compute is O(TEF) rather than truly sparse; it is a safe tensor-path check, not a throughput demonstration. A production dispatcher computes only selected expert outputs but adds grouping, exchange and overflow costs. Resident weights still grow with expert count, while active expert arithmetic grows with selected experts per token. Measure the full path against the dense reference before claiming any gain.

Common Mistakes

  • Do not cite this all-experts toy forward as proof of sparse serving speed.
  • Do not hide severe-event overflow inside an aggregate average.
  • Do not release the expert layer without the exact router and capacity version.

Read next

ai-data
deep-learning
Storage details