A model that computes quickly can still miss the service deadline when requests wait in a queue or feature lookup stalls.
Inference latency budgets: measure queue, feature and model time
Allocate the request deadline
A receipt-scoring endpoint has a 240 ms service deadline. Reserve time for ingress, feature lookup, queue wait, model execution and response serialization. The budget is a design constraint, not a promise that each part always finishes exactly on schedule. Instrument each stage with a bounded set of labels: route, model version, outcome class and region. Never put request IDs into metric labels; retain request-level correlation in traces with privacy controls. Inference logs distinguish useful join keys from unnecessary payloads.
Measure tails and failed requests
A low mean can hide an overloaded slice. Record duration histograms and count timeouts, cancellations, rejections and fallback responses in the same evaluation window. A p95 value calculated only from successful requests excludes some of the slowest calls. Compare candidate and incumbent under comparable traffic and hardware. Canary release needs this operational comparison in addition to model quality.
Use one end-to-end clock
Start a monotonic timer when the service accepts a request and stop it when a response or terminal failure is committed. Service stages can have nested timers, but their measured durations may overlap under asynchronous work; do not sum overlapping spans and call that user latency. Include retries and queue wait in the end-to-end result. Set a downstream deadline that leaves time to return an honest fallback before the client gives up.
Test load shapes, not just steady throughput
Run bursts, cold starts, large feature payloads and a slow feature store. Track latency by model version and outcome, but keep label cardinality bounded. The serving SLO project asks whether the system can reject excess work promptly instead of letting a growing queue make every customer wait. A fast model without admission control is still an unreliable service.
Implementation
def request_budget(stages_ms, deadline_ms):
if deadline_ms <= 0 or any(value < 0 for value in stages_ms.values()):
raise ValueError("durations and deadline must be valid")
spent = sum(stages_ms.values())
return {"within_budget": spent <= deadline_ms,
"spent_ms": spent, "remaining_ms": max(0, deadline_ms - spent)}
receipt_path = {"ingress": 18, "feature_lookup": 71, "queue": 33,
"model": 62, "serialization": 19}
assert request_budget(receipt_path, 240) == {
"within_budget": True, "spent_ms": 203, "remaining_ms": 37}
assert not request_budget({**receipt_path, "queue": 82}, 240)["within_budget"]
Performance and operating cost
Combining s nonoverlapping stage durations takes O(s) time and O(1) auxiliary space. Histograms add bounded metric storage per label combination; request identifiers as labels would cause unbounded cardinality. The arithmetic example assumes sequential stages, so real traces must account for overlap and record end-to-end latency directly.
Common Mistakes
- Reporting a success-only p95 that omits timeouts.
- Summing overlapping spans as if they were sequential.
- Ignoring queue time because model execution is fast.
- Using request IDs as metric labels.
Read next
- Serving overload: bound queues and choose a fallback before time runs out
- Project: operate receipt scoring with a deadline and overload path
- Shadow and canary rollout: compare a candidate without losing a rollback
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Streaming Analytics Tutorial
Continue the workflow: Test partial failure and deadline exhaustion in model chains.
Continue the workflow: Live inference batching: spend queue time inside a request deadline.
