Batching works only when requests share a valid input contract and the cost of padding or preprocessing remains bounded.
Batch shape compatibility: prevent large requests from setting every caller’s cost
Define what can share execution
The parcel model accepts photographs after resizing, but one client also sends panoramic scans. A batch must share model digest, preprocessing revision, tensor layout and any state constraint. It is unsafe to combine requests merely because they reached the same endpoint. A stateful sequence may need ordered affinity, while a stateless classifier can often group independent requests. Request contracts should make these differences explicit before scheduler policy is tuned.
Control the largest-item penalty
A text model may pad a batch to its longest sequence. If three short claims enter alongside a long document, all four slots can pay for the long length. Group by a small number of length or shape buckets, cap accepted payload size and measure useful tokens or pixels against processed tokens or pixels. Buckets add fragmentation: a rare shape may wait longer. Set a maximum queue age per bucket, then let a partial batch proceed instead of violating a caller deadline. Admission checks bound the worst single request.
Track fairness across request classes
A single first-in queue can let expensive scans delay small image requests. Separate priority lanes only for an approved service contract, with capacity reserved so lower-priority work does not starve. Measure queue time and deadline misses per size bucket and client class with bounded metric labels. Do not include raw client IDs in high-cardinality monitoring. Metric identity rules keep dashboards usable; the deadline policy remains the governing constraint.
Test failure and ordering
Exercise a malformed payload, an oversize scan, a canceled request and a batch that partially fails. Every result must be mapped back to its request ID; batch order is not a safe substitute when stages can retry or cancel independently. Record which requests were rejected before execution and which failed after execution. Retry controls prevent a timeout from multiplying device load. In the project, a panoramic scan is isolated without quietly dropping small parcels.
Implementation
def compatible_groups(requests, max_pixels):
groups = {}
rejected = []
for request in requests:
if request["pixels"] > max_pixels:
rejected.append(request["request_id"])
continue
shape_bucket = "large" if request["pixels"] > 350_000 else "small"
key = (request["model_digest"], request["preprocess_rev"], shape_bucket)
groups.setdefault(key, []).append(request["request_id"])
return groups, rejected
requests = [
{"request_id": "parcel-47", "pixels": 210_000,
"model_digest": "damage-r7", "preprocess_rev": "resize-3"},
{"request_id": "parcel-48", "pixels": 890_000,
"model_digest": "damage-r7", "preprocess_rev": "resize-3"},
{"request_id": "scan-49", "pixels": 2_100_000,
"model_digest": "damage-r7", "preprocess_rev": "resize-3"},
]
groups, rejected = compatible_groups(requests, 1_000_000)
assert sorted(len(group) for group in groups.values()) == [1, 1]
assert rejected == ["scan-49"]
Performance and operating cost
Grouping n requests is O(n) expected time and O(n) storage. A real scheduler also pays queue bookkeeping and may lose batching efficiency when shape buckets are too fine. The benefit must be weighed against padding cost, deadline misses and memory use under the actual traffic distribution.
Common Mistakes
- Grouping different preprocessing or model revisions because the endpoint name matches.
- Letting one long or large item determine compute for every small request.
- Creating so many buckets that rare shapes wait past their deadlines.
- Returning results by batch position after cancellations or retries change order.
Read next
- Live inference batching: spend queue time inside a request deadline
- Project: tune parcel-damage inference for bursts and mixed image sizes
- Inference ingress: bound payload size and compute before model work
- Inference API contracts: version the decision, not only the payload
- Model monitoring dimensions without metric-cardinality failure
