Treat suspicious inference patterns as evidence for review, with proportional containment and an explicit release path.
Inference evasion response: rate limits, review and reversible containment
Choose a bounded detection window
Group submissions by actor and parcel over a short window. Count distinct valid image digests, threshold crossings and attempts with inconsistent metadata. A repeated request with the same digest is a retry, not a new candidate image. A device with poor connectivity can produce many retries, so combine distinctness with score behavior and review context. Idempotent retries keep the serving record honest; the threat model states what the actor controls.
Separate alert from adverse action
A suspicious pattern can trigger step-up review, a temporary rate cap or a hold on automated clearance. Keep a human path for urgent shipments and record the decision owner. Do not ban an account because a heuristic crossed one cutoff. Set expiry and appeal conditions for each hold, and record the model, rule and window revision in the case. The override ledger keeps later corrections from disappearing.
Measure operational harm
Track cases confirmed harmful, cases cleared as legitimate retakes, review queue age, rate-limit impact and damage misses by scanner cohort. If alerts concentrate on one camera generation after a software update, investigate drift before assigning attacker intent. Compare before and after using the same observation frame. Delayed-label monitoring helps separate an actual outcome change from an early proxy signal.
Rehearse rollback and evidence handoff
Run an authorized test where a parcel is resubmitted after a harmless lighting change, then a separate test with repeated score flips across distinct images. Ensure the first is not permanently blocked and the second reaches review. If a new rule floods the queue, disable that rule revision and retain cases for inspection. Incident timelines capture the sequence; the project ties alerting to a reversible response.
Implementation
def submission_window_action(events, maximum_distinct, score_span):
distinct = {event["image_digest"] for event in events}
if len(distinct) <= maximum_distinct:
return "serve:normal"
scores = [event["score"] for event in events]
if max(scores) - min(scores) < score_span:
return "observe:retries-or-retakes"
return "review:temporary-automation-hold"
history = [{"image_digest": "frame-a", "score": 0.72},
{"image_digest": "frame-b", "score": 0.68},
{"image_digest": "frame-c", "score": 0.31}]
assert submission_window_action(history, 2, 0.27) == "review:temporary-automation-hold"
assert submission_window_action(history[:2], 2, 0.27) == "serve:normal"
Performance and operating cost
The window function is O(n) time and O(n) space for n events; a keyed streaming implementation should expire old actor-parcel state. Review costs rise with false alerts, while an aggressive hold can delay legitimate parcels. Tune thresholds using labeled cases and monitor queue capacity; this heuristic cannot prove an attack.
Common Mistakes
- Counting identical retries as distinct attack attempts.
- Turning a review signal into an indefinite account ban.
- Ignoring camera drift that changes scores for innocent retakes.
- Deploying a new rule without a disable switch and queue budget.
Read next
- Inference-time evasion: threat model and input evidence
- Project: investigate repeated parcel-scanner score flips
- Inference retries: bound repeated work and preserve one decision
- Model incidents: build a release and evidence timeline before rollback
- Model monitoring: separate input drift, data faults and delayed outcomes
