Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: protect receipt inference from expensive and repeated input

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Place shape limits, tenant quotas and idempotency ahead of extraction, then prove rejected and retried requests do not consume duplicate model work.

Write the ingress contract

Specify maximum upload bytes, decoded pixels, pages and accepted formats for receipt scans. Define a bounded alternate review path for legitimate larger submissions. Add tenant rate and in-flight limits, a stable idempotency key and a response code for size, conflict and temporary overload. Payload admission should run before extraction; the public contract explains each rejection to callers.

Inject resource abuse cases

Send a compressed scan whose decoded dimensions exceed the pixel limit, a 47-page document and a burst of small valid receipts. Verify oversized requests never start the extractor or GPU worker. The burst should hit a tenant limit while unrelated tenants retain their model slots. Record gateway request count, admission result, expensive stage invocations and caller-visible response for each case. Shared model isolation adds protection against one model consuming the entire pool.

Test retries after uncertain completion

Allow one valid receipt to commit a decision, then drop its response. Retry with the original tenant, key and payload: return the existing decision ID without a second model invocation. Retry with the same key but changed bytes: reject as a conflict. The retry rule must use an atomic shared record in the deployed system. A local unit example proves branch logic, not cross-replica safety; run the integration test across two workers.

Publish the load and safety result

Report accepted, rejected, replayed and conflicting requests; extraction and model invocation counts; p99 latency for unaffected tenants; quota state and deduplication record lifetime. Identify the owner of each limit and the procedure for a legitimate oversized document. A guard that silently drops requests or converts malformed input to a low-risk score fails this exercise. Link the result to the serving target and to complete telemetry counts so the protection is observable.

Implementation

python
def guard_result(payload_bytes, pages, replay_state, tenant_active):
    if payload_bytes > 4_700_000 or pages > 8:
        return "reject:payload"
    if replay_state == "conflict":
        return "reject:key-conflict"
    if replay_state == "completed":
        return "return:prior-decision"
    if tenant_active >= 23:
        return "retry:tenant-capacity"
    return "start:inference"

assert guard_result(470_000, 2, "new", 18) == "start:inference"
assert guard_result(470_000, 47, "new", 0) == "reject:payload"
assert guard_result(470_000, 2, "completed", 23) == "return:prior-decision"

Performance and operating cost

The illustrated gate is O(1) time and space per request. A real decoder needs resource limits after expansion, and deduplication requires a shared atomic store. Quotas trade some peak throughput for isolation. The integration check should confirm that rejected requests never enter expensive stages and that replayed responses do not consume fresh inference capacity.

Common Mistakes

  • Enforcing payload limits only after extraction.
  • Returning a fresh decision for a retry whose first response was lost.
  • Letting one tenant consume reserved capacity for others.
  • Using a non-atomic per-worker retry map in a multi-worker deployment.

Read next

ai-data
mlops
Storage details