Sharing inference workers saves idle capacity only when one model cannot consume the memory, queue and latency budget of another.
Multi-model serving: admission control for noisy neighbors
Measure each tenant separately
A shared worker may host a receipt triage model and an invoice classifier. Track request rate, queue wait, inference time, memory resident, load time and fallback by model digest and client. Aggregate endpoint latency hides the small model whose tail latency doubles when a larger model reloads. Tail observation should retain the model key, while alerts should name the affected owner rather than page everyone for total endpoint utilization.
Reserve before dispatch
Set per-model in-flight and queued-request limits based on measured service time and memory. Reject or route overflow to a named fallback before the shared queue grows without bound. A high-priority model may need reserved slots; fairness can use weighted queues, but weights should reflect business harm and service-level objectives. One model that loads a large artifact can evict another from cache and trigger repeated cold starts. Artifact loading is part of the serving path, not free background work.
Isolate the expensive boundary
Keep models with strict tail latency or very high traffic on dedicated workers when sharing makes isolation too costly. Separate GPU memory pools or deployment groups where an out-of-memory event can restart every tenant. Requests and limits at the container level help but do not guarantee per-model fairness inside one process; the dispatch queue still needs its own control. Queue saturation rules describe how an overloaded model should fail without taking down unrelated traffic.
Prove protection under burst
Load test both models together, then burst one while holding the other at a steady baseline. Record its p95/p99 latency, rejected count and model reloads. A single-model benchmark cannot reveal interference. The project forces a surge in invoice traffic while checking whether receipt triage retains reserved slots. If a shared configuration fails the test, move the strict model to dedicated capacity and measure the extra cost.
Implementation
def admit_model_request(model_key, active, limits):
if model_key not in limits:
return "reject:unknown-model"
if active.get(model_key, 0) >= limits[model_key]:
return "reject:model-capacity"
return "admit"
limits = {"receipt-risk": 47, "invoice-classifier": 20}
active = {"receipt-risk": 18, "invoice-classifier": 20}
assert admit_model_request("receipt-risk", active, limits) == "admit"
assert admit_model_request("invoice-classifier", active, limits) == "reject:model-capacity"
assert admit_model_request("unknown", active, limits) == "reject:unknown-model"
Performance and operating cost
Admission is O(1) expected time and space per request. A production dispatcher must update counters atomically to prevent concurrent over-admission. Reserving slots lowers peak utilization, while sharing lowers idle infrastructure cost. Measure the combined bill and the induced tail latency; a cheaper shared endpoint may be unacceptable for the strict model.
Common Mistakes
- Watching only endpoint-wide average latency.
- Assuming container CPU limits enforce per-model fairness.
- Letting artifact reloads consume the strict model’s reserved memory.
- Setting queue limits without a caller-visible overload response.
Read next
- Model cache warmup and memory planning for shared inference
- Project: protect receipt inference in a shared model pool
- Serving overload: bound queues and choose a fallback before time runs out
- Inference latency budgets: measure queue, feature and model time
- Model alerts: page on customer symptoms with a named owner
Continue the workflow: Inference cost ledger: allocate shared capacity without false precision.
Continue the workflow: Project: release a receipt-quality and risk-model cascade.
