An inference service must limit work in flight and define what happens when it cannot score within the caller deadline.
Serving overload: bound queues and choose a fallback before time runs out
Budget capacity in work units
A queue of 470 tiny requests may be manageable, while 470 image embeddings may exhaust memory. Define a concurrency limit, queue length or weighted work budget for the deployed model and hardware. Admit only work that has a credible chance of finishing before its deadline. When capacity is exhausted, return a declared overloaded state or an approved deterministic fallback. Silently waiting past the caller deadline wastes compute and hides the cause. Latency budgets make the decision measurable.
Separate safe fallback from a guess
For a receipt-risk service, a timeout might route a transaction to manual review; it should not return a fabricated low-risk score. The fallback must have an explicit customer and operations contract, with its own volume limit and audit trail. If an old model is the fallback, check its feature contract and availability; the old model may share the failed feature store and thus fail at the same time. Schema compatibility matters during degradation as much as during promotion.
Plan retry behavior across boundaries
A client retry can multiply load during an outage. Set caller and server deadlines, bounded retries with jitter where appropriate, and idempotent request identity for side-effecting workflows. Do not retry a model score blindly if it triggers a downstream action. Count initial attempts separately from retry attempts so dashboards reveal amplification. A circuit breaker can stop calls to a known-failing dependency, but it should not conceal persistent failures from incident owners.
Exercise the failure path
Load test at normal and burst rates, then slow the feature store and model independently. Verify that rejected work leaves the queue promptly, admitted work respects deadlines, and fallback counts are visible. Check that a rollout does not consume all spare capacity with shadow traffic. The serving project includes overload, dependency failure and rollback drills with an explicit customer outcome for each.
Implementation
def admission_decision(active_work, queued_work, deadline_ms,
estimated_wait_ms, model_ms):
if min(active_work, queued_work, deadline_ms,
estimated_wait_ms, model_ms) < 0:
raise ValueError("capacity inputs must be nonnegative")
if active_work >= 24 and queued_work >= 47:
return "manual-review-overload"
if estimated_wait_ms + model_ms > deadline_ms:
return "manual-review-deadline"
return "admit"
assert admission_decision(12, 8, 180, 31, 62) == "admit"
assert admission_decision(24, 47, 180, 31, 62) == "manual-review-overload"
assert admission_decision(12, 8, 80, 31, 62) == "manual-review-deadline"
Performance and operating cost
This local admission decision is O(1) time and space. A production counter must update atomically across workers, and estimated wait time can be wrong during bursts. Model execution and fallback routing have different costs; reserve capacity for manual review rather than shifting an unbounded queue into another service.
Common Mistakes
- Returning a low-risk score on timeout without evidence.
- Allowing retries to multiply an already overloaded queue.
- Using a queue limit without a request deadline.
- Assuming a fallback model has independent dependencies.
Read next
- Inference latency budgets: measure queue, feature and model time
- Project: operate receipt scoring with a deadline and overload path
- Feature schema evolution: keep producers and rollback models compatible
- Shadow and canary rollout: compare a candidate without losing a rollback
- Model monitoring: separate input drift, data faults and delayed outcomes
Continue the workflow: Multi-model serving: admission control for noisy neighbors.
