Query-pattern alerts need context, an investigation trail and reversible controls before they affect a client.
Prediction API query review: distinguish probing from legitimate bursts
Look for structure in a bounded window
A supplier-risk client sends hundreds of similar requests with one feature changed at a time. Another client legitimately replays a nightly batch after an outage. Count distinct request IDs, feature-family digests, response scopes, route transitions and per-minute rate under one authenticated identity. Compare against the client’s registered workload and prior windows. Exact retries should collapse under the same idempotency key. The exposure contract defines which answers the client should receive.
Make alerts specific and reversible
A sudden surge with many one-field variations can open a case. Step down to a lower request rate or route-only response when the client contract permits it; keep a clear expiry and emergency exception for critical workflows. Do not silently change numeric results for the same request, and do not block a known recovery batch solely because it is large. Evasion containment is a neighboring workflow but asks whether inputs manipulate decisions, not whether outputs reveal model behavior.
Investigate with limited records
Preserve pseudonymous caller ID, rule revision, request count, distinct-input ratio, response scope and sampled request IDs. An investigator can ask the client about a known batch replay and compare timing with deployment or outage records. Record the conclusion as confirmed misuse, explained burst or unresolved. Avoid inferring theft of weights from a pattern alone. Incident evidence keeps who decided what and when.
Measure the controls themselves
Track rate-limit errors, blocked legitimate calls, investigation precision, case age and client appeal outcomes. If a rule blocks many ordinary batch jobs, reduce its scope or change the baseline, then version the rule. Keep model revisions and probability scopes in the case because a release can change query behavior without a caller changing intent. The project exercises both a probing-like sequence and a planned replay.
Implementation
def query_case(window, client_plan):
if window["unique_requests"] <= client_plan["planned_requests"]:
return "observe:within-plan"
if window["one_field_variations"] < client_plan["variation_alert"]:
return "review:unexpected-volume"
if window["response_scope"] != client_plan["approved_scope"]:
return "hold:scope-mismatch"
return "review:structured-query-pattern"
plan = {"planned_requests": 82, "variation_alert": 47,
"approved_scope": "route-reader"}
window = {"unique_requests": 119, "one_field_variations": 63,
"response_scope": "route-reader"}
assert query_case(window, plan) == "review:structured-query-pattern"
assert query_case({**window, "unique_requests": 72}, plan) == "observe:within-plan"
Performance and operating cost
The aggregate rule is O(1) time and space; producing its window features is O(n) time over n requests and needs bounded per-client state. Review time and false-positive client disruption are material. Sampled evidence lowers storage but may miss a particular query sequence, so the sampling policy belongs in the case packet.
Common Mistakes
- Equating high request volume with confirmed extraction.
- Ignoring a client’s documented outage replay.
- Changing response values instead of applying a visible contract or rate policy.
- Leaving a temporary rate cap active without expiry or owner.
Read next
- Prediction API exposure: define what a client may learn from scores
- Project: audit a supplier-risk prediction API after a query surge
- Model incidents: build a release and evidence timeline before rollback
- Inference retries: bound repeated work and preserve one decision
- Inference evasion response: rate limits, review and reversible containment
