Place receipt triage beside an invoice classifier, inject cold loads and a traffic burst, then choose shared or dedicated capacity from measured harm and cost.
Project: protect receipt inference in a shared model pool
Define the serving envelope
The receipt model has a strict tail-latency budget and a safe manual-review fallback. The invoice classifier has a looser deadline and a larger artifact. Measure both artifacts’ resident and peak memory, load time and steady inference time on the target worker. Reserve slots for receipt requests before sharing a queue. The warm set names exact approved digests, and per-model admission prevents invoice traffic from taking every slot.
Run three comparative loads
First test each model alone to establish a baseline. Next hold receipt traffic steady while invoice traffic surges to several times its usual rate. Finally restart one worker and force both artifacts to load again. Capture p50, p95 and p99 latency, rejected requests, queue wait, cold starts, evictions and memory pressure for each model separately. Aggregate endpoint latency may look acceptable while receipt p99 fails. Keep the same input mix and concurrency for all comparisons.
Apply the protection policy
When invoice traffic exhausts its allocation, reject or delay invoice requests before they enter the receipt reserve. Ensure overload uses a documented response rather than an ordinary score. If the receipt model is evicted or exceeds its deadline during the test, pin it, increase ready replicas or move it to dedicated workers. Fallback policy describes what the receipt caller sees; a container restart alone does not make a safe decision.
Record the capacity decision
Compare infrastructure cost per hour with p99 receipt latency, missed requests and operator burden. Document chosen worker count, memory reserve, model limits, warmup time and when to retest. A shared pool may be cheaper at low traffic but lose that advantage once enough idle capacity is reserved for isolation. Link the result to the receipt service-level target and to model-specific alerts so a future invoice surge is detected quickly.
Implementation
def serving_choice(shared, dedicated, receipt_p99_limit_ms=180):
if shared["receipt_p99_ms"] > receipt_p99_limit_ms:
return "dedicated:latency"
if shared["receipt_failed_requests"] > 0:
return "dedicated:failed-requests"
if shared["hourly_cost"] < dedicated["hourly_cost"]:
return "shared"
return "dedicated:cost"
shared = {"receipt_p99_ms": 214, "receipt_failed_requests": 0,
"hourly_cost": 4.7}
dedicated = {"receipt_p99_ms": 136, "receipt_failed_requests": 0,
"hourly_cost": 6.2}
assert serving_choice(shared, dedicated) == "dedicated:latency"
assert serving_choice({**shared, "receipt_p99_ms": 150}, dedicated) == "shared"
Performance and operating cost
The final comparison is O(1) time and space, but the evidence requires representative multi-model load tests, cold-start trials and cost measurement. The threshold and prices are fictional planning values. A single p99 estimate from too few requests is unstable; measure over enough traffic and repeat under a worker restart before making a hosting decision.
Common Mistakes
- Deciding from a single-model benchmark.
- Calling shared capacity cheaper without counting reserved idle replicas.
- Ignoring cold starts after scale-out and restart.
- Accepting an endpoint-wide SLO while receipt-specific p99 fails.
Read next
- Multi-model serving: admission control for noisy neighbors
- Model cache warmup and memory planning for shared inference
- Project: operate receipt scoring with a deadline and overload path
- Serving overload: bound queues and choose a fallback before time runs out
- Model alerts: page on customer symptoms with a named owner
