A serving API must specify request validation, response meaning, model identity and stable failure behavior for every caller.
Inference API contracts: version the decision, not only the payload
Separate public and internal contracts
The HTTP request belongs to the caller; the feature vector belongs to the serving pipeline. A public receipt request may contain a receipt ID and capture time while the internal model consumes normalized merchant age and currency signals. Changing the feature contract does not require every caller to change. Conversely, renaming a public field is breaking even if the model bytes are identical. Feature schema evolution covers internal consumers; this page covers the endpoint boundary and its visible decision semantics.
Specify acceptance and rejection
Write the accepted media type, required fields, units, length limits, idempotency key and timestamp rules. Define a stable error code for invalid input, unsupported currency and temporary feature unavailability. Do not reply with a successful score when admission fails. A response should include a decision ID, route, model digest or approved alias, policy revision and a machine-readable reason for fallback. Admission gates decide whether the feature state is fit to score; the API translates that state into a safe caller response.
Protect old clients during change
Additive optional fields can often remain within one API version if old clients ignore unknown response fields and new required fields are not introduced. A changed score meaning, route enum or error code requires deliberate migration even when the JSON shape stays the same. Keep the old contract available until every registered consumer has moved, or provide an explicit version adapter with contract tests. Client migration sets the evidence needed to retire that adapter.
Test against real caller expectations
Use captured redacted requests and caller-owned assertions to test valid input, absent fields, extra fields, repeated idempotency keys and unavailable features. Exercise the actual deployed endpoint; a schema test alone cannot catch a model digest unexpectedly promoted behind an alias. The project migrates two receipt consumers while preserving their response behavior. Keep response identity in restricted decision logs so an incident can join an API response to the served model.
Implementation
def validate_receipt_request(payload):
required = {"receipt_id", "currency", "captured_at"}
if not isinstance(payload, dict) or not required <= payload.keys():
return {"ok": False, "code": "INVALID_REQUEST"}
if not all(isinstance(payload[key], str) and payload[key]
for key in required):
return {"ok": False, "code": "INVALID_REQUEST"}
if payload["currency"] not in {"INR", "SGD"}:
return {"ok": False, "code": "UNSUPPORTED_CURRENCY"}
return {"ok": True, "code": "ACCEPTED"}
valid = {"receipt_id": "r-47", "currency": "INR",
"captured_at": "2026-10-06T09:00:00Z"}
assert validate_receipt_request(valid)["ok"]
assert validate_receipt_request({**valid, "currency": "EUR"})["code"] == "UNSUPPORTED_CURRENCY"
Performance and operating cost
Validation is O(f) expected time for f required fields and O(1) extra space here. A full endpoint also pays for parsing, authentication, feature reads and model inference. The example intentionally omits timestamp parsing and size limits; a real ingress should validate those at its trust boundary before feature lookup.
Common Mistakes
- Treating a model artifact revision as an API version.
- Returning HTTP success with an unmarked fallback score.
- Changing a response enum without testing old callers.
- Logging raw receipts to diagnose contract failures.
Read next
- Migrate inference clients with compatibility tests and usage evidence
- Project: migrate receipt inference callers without breaking decisions
- Feature schema evolution: keep producers and rollback models compatible
- Feature contracts: admit only usable inference records
- Inference logs: keep diagnostic joins without copying sensitive payloads
Continue the workflow: Inference ingress: bound payload size and compute before model work.
Continue the workflow: Prediction API exposure: define what a client may learn from scores.
Continue the workflow: Synthetic inference probes: test the full route without polluting outcomes.
