Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Inference ingress: bound payload size and compute before model work

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Validate input shape and resource cost before parsing or model execution so one request cannot exhaust a shared serving pool.

Limit the earliest expensive boundary

A receipt endpoint may accept a scan image and metadata. Reject oversized bodies at the gateway before copying them into application memory; then check decoded dimensions, file count, page count and permitted format before extraction. A compressed file can expand far beyond its upload size, so both compressed and decoded limits matter. The API contract must return a stable rejection code. Do not let a malformed payload reach a GPU worker simply because its HTTP body passed a size check.

Estimate work by request class

A one-page image and a 47-page document can have very different extraction and inference cost. Define bounded request classes with a maximum estimated work unit, then admit against the model-specific concurrency and queue budget. Separate authentication, authorization and customer quota from model capacity; a valid customer can still send a request that exceeds the endpoint limit. Per-model admission protects neighboring models; this ingress gate prevents a single allowed request from consuming unreasonable resources.

Fail clearly and do not retry blindly

Return an explicit client error for invalid shape or unsupported size; retriable overload needs a different code and bounded retry advice. A caller that retries an invalid 47-page document will only repeat the waste. Log a reason and bounded size bucket, not the document bytes, in operational telemetry. Metric-cardinality rules keep document IDs out of labels, while restricted logs retain a joinable request ID if investigation is needed.

Test bypass and expansion cases

Exercise tiny files with huge declared dimensions, compressed payloads with high expansion ratios, malformed metadata, arrays with too many elements and concurrent near-limit requests. Ensure the gateway and application enforce compatible limits; one layer may be bypassed during internal calls. The project demonstrates a request rejected before extraction and a permitted request that still meets its latency budget. A 413 response alone does not prove the decoder never allocated memory.

Implementation

python
def admit_receipt_payload(request, max_bytes=4_700_000,
                          max_pages=8, max_pixels=18_000_000):
    checks = (request["body_bytes"], request["pages"], request["pixels"])
    if any(value < 0 for value in checks):
        raise ValueError("negative payload measure")
    if request["body_bytes"] > max_bytes:
        return "reject:body-size"
    if request["pages"] > max_pages or request["pixels"] > max_pixels:
        return "reject:decode-work"
    return "admit"

receipt = {"body_bytes": 470_000, "pages": 2, "pixels": 8_200_000}
assert admit_receipt_payload(receipt) == "admit"
assert admit_receipt_payload({**receipt, "pages": 47}) == "reject:decode-work"

Performance and operating cost

Checking declared bounds is O(1) time and space. Decoding still costs work proportional to expanded pixels and pages, so verify actual decoded size before large allocations and enforce runtime memory limits. Tight limits may reject legitimate long receipts; choose them from measured use and a documented alternate review path.

Common Mistakes

  • Checking upload bytes but not decoded pixel count.
  • Treating invalid shape as retriable server overload.
  • Placing the first size check after expensive decoding.
  • Assuming valid authentication makes resource use safe.

Read next

Continue the workflow: Batch shape compatibility: prevent large requests from setting every caller’s cost.

ai-data
mlops
Storage details