Validate input shape and resource cost before parsing or model execution so one request cannot exhaust a shared serving pool.
Inference ingress: bound payload size and compute before model work
Limit the earliest expensive boundary
A receipt endpoint may accept a scan image and metadata. Reject oversized bodies at the gateway before copying them into application memory; then check decoded dimensions, file count, page count and permitted format before extraction. A compressed file can expand far beyond its upload size, so both compressed and decoded limits matter. The API contract must return a stable rejection code. Do not let a malformed payload reach a GPU worker simply because its HTTP body passed a size check.
Estimate work by request class
A one-page image and a 47-page document can have very different extraction and inference cost. Define bounded request classes with a maximum estimated work unit, then admit against the model-specific concurrency and queue budget. Separate authentication, authorization and customer quota from model capacity; a valid customer can still send a request that exceeds the endpoint limit. Per-model admission protects neighboring models; this ingress gate prevents a single allowed request from consuming unreasonable resources.
Fail clearly and do not retry blindly
Return an explicit client error for invalid shape or unsupported size; retriable overload needs a different code and bounded retry advice. A caller that retries an invalid 47-page document will only repeat the waste. Log a reason and bounded size bucket, not the document bytes, in operational telemetry. Metric-cardinality rules keep document IDs out of labels, while restricted logs retain a joinable request ID if investigation is needed.
Test bypass and expansion cases
Exercise tiny files with huge declared dimensions, compressed payloads with high expansion ratios, malformed metadata, arrays with too many elements and concurrent near-limit requests. Ensure the gateway and application enforce compatible limits; one layer may be bypassed during internal calls. The project demonstrates a request rejected before extraction and a permitted request that still meets its latency budget. A 413 response alone does not prove the decoder never allocated memory.
Implementation
def admit_receipt_payload(request, max_bytes=4_700_000,
max_pages=8, max_pixels=18_000_000):
checks = (request["body_bytes"], request["pages"], request["pixels"])
if any(value < 0 for value in checks):
raise ValueError("negative payload measure")
if request["body_bytes"] > max_bytes:
return "reject:body-size"
if request["pages"] > max_pages or request["pixels"] > max_pixels:
return "reject:decode-work"
return "admit"
receipt = {"body_bytes": 470_000, "pages": 2, "pixels": 8_200_000}
assert admit_receipt_payload(receipt) == "admit"
assert admit_receipt_payload({**receipt, "pages": 47}) == "reject:decode-work"
Performance and operating cost
Checking declared bounds is O(1) time and space. Decoding still costs work proportional to expanded pixels and pages, so verify actual decoded size before large allocations and enforce runtime memory limits. Tight limits may reject legitimate long receipts; choose them from measured use and a documented alternate review path.
Common Mistakes
- Checking upload bytes but not decoded pixel count.
- Treating invalid shape as retriable server overload.
- Placing the first size check after expensive decoding.
- Assuming valid authentication makes resource use safe.
Read next
- Inference retries: bound repeated work and preserve one decision
- Project: protect receipt inference from expensive and repeated input
- Inference API contracts: version the decision, not only the payload
- Multi-model serving: admission control for noisy neighbors
- Model monitoring dimensions without metric-cardinality failure
Continue the workflow: Batch shape compatibility: prevent large requests from setting every caller’s cost.
