A live inference batch combines compatible requests, but its queue wait must fit inside each caller’s remaining deadline.
Live inference batching: spend queue time inside a request deadline
Account for the whole request
A parcel-damage model runs on a shared accelerator. The service can process several photographs together, so batching may improve device use, yet a request spends time in admission, preprocessing, queueing, inference, postprocessing and network delivery. Set the caller deadline before choosing a batch wait. Reserve a measured margin for everything after the queue. The latency budget separates those stages; a fast model kernel does not rescue a caller whose queue already consumed the deadline.
Release on size or age
Use a bounded maximum batch size and a maximum wait measured from the oldest queued request. A full compatible batch can leave immediately. At low traffic, release a partial batch when the oldest request reaches its wait limit. A scheduler that waits for every batch to fill makes off-peak callers pay indefinitely for throughput that never arrives. Check the oldest request’s absolute deadline before adding a new one; the queue must not admit work that cannot plausibly complete.
Measure the trade with realistic arrivals
Compare single-request serving with several batch limits under both steady and bursty arrivals. Report requests per second, achieved batch-size distribution, queue p50 and p99, end-to-end p95 and p99, deadline misses, rejected work and device utilization. A larger batch can reduce compute per request while increasing queue age or memory pressure. Measure the actual shape mix: photographs with different decode and resize costs can shift the bottleneck before inference begins. Warmup and capacity alter the result; shape control prevents one large payload from dominating a group.
Keep an escape route
If the queue is near saturation, reject early or use a separately approved lower-cost route rather than accepting a doomed request. Pin scheduler settings and model digest together for canary comparison. Roll back on sustained deadline misses even if average throughput rises. Canary rules need both tail latency and error rate; the parcel project tests a burst after the initial benchmark.
Implementation
def release_batch(queue, now_ms, max_size, max_wait_ms, reserve_ms):
if not queue:
return [], "empty"
if max_size < 1 or max_wait_ms < 0 or reserve_ms < 0:
raise ValueError("invalid scheduler budget")
oldest = queue[0]
expiry_ms = oldest["deadline_ms"] - reserve_ms
if now_ms >= expiry_ms:
return [], "reject:deadline"
if len(queue) >= max_size or now_ms - oldest["arrival_ms"] >= max_wait_ms:
return queue[:max_size], "release"
return [], "wait"
waiting = [{"request_id": "parcel-47", "arrival_ms": 100,
"deadline_ms": 190}]
assert release_batch(waiting, 121, 4, 20, 38)[1] == "release"
assert release_batch(waiting, 160, 4, 20, 38)[1] == "reject:deadline"
Performance and operating cost
Inspecting the oldest entry and slicing a released batch costs O(b) time and space for b released requests. A queue with constant-time append and removal avoids repeatedly shifting a Python list in a real server. Batching can improve device throughput while adding queue delay and temporary activation memory; profile the whole path at representative concurrency.
Common Mistakes
- Measuring kernel time while omitting queue and preprocessing time.
- Waiting for a full batch when traffic falls below the expected rate.
- Accepting a request whose remaining deadline cannot cover execution.
- Promoting a throughput gain after p99 deadline misses rise.
Read next
- Batch shape compatibility: prevent large requests from setting every caller’s cost
- Project: tune parcel-damage inference for bursts and mixed image sizes
- Inference latency budgets: measure queue, feature and model time
- Model cache warmup and memory planning for shared inference
- Shadow and canary rollout: compare a candidate without losing a rollback
