Choose a bounded batch policy, measure queue tails and reject a configuration that gains throughput by missing caller deadlines.
Project: tune parcel-damage inference for bursts and mixed image sizes
Specify the serving contract
A sorting-center camera calls a parcel-damage classifier and needs a result within 190 milliseconds. Keep 38 milliseconds for preprocessing, postprocessing and transport after the queue. Each request carries arrival time, deadline, image size, model digest and preprocessing revision. A scan above the approved pixel limit is rejected before scheduling. Payload admission keeps the queue from becoming a storage area for unbounded work.
Benchmark candidate schedulers
Run a single-request baseline, a small batch with a short wait and a larger batch with a longer wait. Replay both steady traffic and a scanner-restart burst. For each policy record achieved batch sizes, useful versus padded pixels, queue p99, end-to-end p99, deadline misses and GPU memory. The long-wait policy wins requests per second during a steady run but fails the restart burst. Keep the shorter wait even though its median device utilization is lower. Deadline accounting decides this trade.
Handle mixed sizes
Separate normal photographs from the panoramic scan shape bucket. The two normal requests can batch; the panorama has an explicit size cap and, if allowed, its own queue age. Test a cancellation while a batch is forming, then verify response IDs and retry tokens still identify the original caller. Compatibility policy avoids processing several small images at panorama cost.
Canary and handoff
Pin scheduler settings alongside the model and resize revision. Send a small traffic share to the candidate and compare per-bucket tail latency, deadline misses, rejection reasons and classification quality against the incumbent. Set a rollback owner and a hard miss-rate gate before traffic rises. Deliver the benchmark packet, selected policy, rejected alternative and rollback evidence. The canary workflow treats this as a serving release, not a mere configuration tweak.
Implementation
def batching_release_gate(candidate, incumbent, limits):
if candidate["quality_recall"] < limits["minimum_recall"]:
return "hold:quality"
if candidate["deadline_miss_rate"] > limits["maximum_miss_rate"]:
return "hold:deadline"
if candidate["queue_p99_ms"] > incumbent["queue_p99_ms"] + limits["queue_regression_ms"]:
return "hold:queue-tail"
return "canary:bounded-batch"
limits = {"minimum_recall": 0.84, "maximum_miss_rate": 0.012,
"queue_regression_ms": 9}
incumbent = {"queue_p99_ms": 24}
slow = {"quality_recall": 0.88, "deadline_miss_rate": 0.026,
"queue_p99_ms": 42}
fast = {"quality_recall": 0.88, "deadline_miss_rate": 0.008,
"queue_p99_ms": 31}
assert batching_release_gate(slow, incumbent, limits) == "hold:deadline"
assert batching_release_gate(fast, incumbent, limits) == "canary:bounded-batch"
Performance and operating cost
The aggregate gate is O(1) time and space once benchmark reports exist. The expensive part is load replay across arrival patterns and image-size buckets; production batching also consumes queue memory. A faster steady-state throughput figure is not worth a missed camera deadline during predictable restart bursts.
Common Mistakes
- Benchmarking only a smooth arrival stream.
- Selecting the largest batch from utilization alone.
- Ignoring rejection and cancellation paths in response-ID mapping.
- Rolling out a scheduler change without a separate tail-latency rollback gate.
Read next
- Live inference batching: spend queue time inside a request deadline
- Batch shape compatibility: prevent large requests from setting every caller’s cost
- Inference latency budgets: measure queue, feature and model time
- Shadow and canary rollout: compare a candidate without losing a rollback
- Inference ingress: bound payload size and compute before model work
