Separate a planned client replay from structured probing, then apply scoped, reversible response controls.
Project: audit a supplier-risk prediction API after a query surge
Map two client contracts
A purchasing system needs only approve, review or decline plus a reason category. An internal risk desk is approved for calibrated probabilities. Record both scopes, authenticated client IDs, request budgets, idempotency rules and retention period. A new endpoint accidentally exposes full probability precision and a hidden feature vector to the purchasing client. Fix the response contract before interpreting any observed traffic. Output scope is the first control.
Classify the surge
Client A sends 119 distinct requests in a short window, with 63 one-field variations and repeated decision-boundary crossings. Client B sends 390 requests after a documented overnight outage, but most are exact retries with stable request IDs. Collapse retries, compare each client with its registered workload and open a review case for A. Do not call A malicious yet, and do not block B’s approved replay. Query review provides the structured evidence packet.
Apply and test limited controls
Remove unauthorized fields from the purchasing response, cap A’s fresh requests temporarily and provide an escalation route. Preserve the same route for the same accepted request ID. Keep B on its recovery allowance. Re-run a legitimate sample from both clients and verify that the risk desk still receives the probability its workflow needs. If A’s case is explained by a sanctioned test, expire the cap and record the conclusion.
Report outcome and limits
Deliver a case packet with client scope, unique-versus-retry counts, response-field exposure, variation signal, rule revision, human decision, hold expiry and service impact. State clearly that the logs show a query pattern, not proof that a model was copied. Track future misuse reports and incident decisions, and connect API changes to request-response contract tests.
Implementation
def api_client_action(client, scope, fresh_requests, planned_limit):
if scope not in client["approved_scopes"]:
return "reject:scope"
if fresh_requests > planned_limit:
return "review:temporary-rate-cap"
return "serve:approved-response"
purchasing = {"approved_scopes": {"route-reader"}}
assert api_client_action(purchasing, "risk-probability-reader", 47, 82) == "reject:scope"
assert api_client_action(purchasing, "route-reader", 119, 82) == "review:temporary-rate-cap"
assert api_client_action(purchasing, "route-reader", 72, 82) == "serve:approved-response"
Performance and operating cost
The contract decision is O(1) expected time and space for a small scope set. Window aggregation scales with fresh request volume and uses per-client expiring state. An overly strict cap can delay legitimate purchasing, while overexposed probabilities reveal unnecessary model behavior; measure both service errors and exposure reductions.
Common Mistakes
- Confusing exact retries with new boundary probes.
- Removing an approved probability field from the internal risk desk.
- Claiming a copied model from traffic patterns alone.
- Keeping a temporary client cap after the investigation closes.
Read next
- Prediction API exposure: define what a client may learn from scores
- Prediction API query review: distinguish probing from legitimate bursts
- Inference API contracts: version the decision, not only the payload
- Inference retries: bound repeated work and preserve one decision
- Model incidents: build a release and evidence timeline before rollback
