An inference evasion attempt changes submitted inputs to obtain a favorable prediction without changing the real-world condition.
Inference-time evasion: threat model and input evidence
Separate adversarial behavior from drift
A parcel scanner model routes damaged packages for review. An operator can submit altered images or metadata to push a damaged parcel below the review threshold. That is an inference-time manipulation concern, distinct from contaminated training data and ordinary camera drift. Define attacker control over pixels, crop, device metadata, retries and account identity. Record the decision the attacker wants and what evidence the service can actually observe. Training-source admission addresses a different stage of the lifecycle.
Preserve a minimal decision trail
Retain a governed request ID, account or device pseudonym, input digest, model revision, preprocessing revision, score band, decision and review outcome where permitted. Keep raw images only under a documented retention policy; an input digest alone cannot explain visual changes. The operational question is whether repeated near-identical submissions from one actor cross the threshold after small edits. The response contract should pin the decision revision and prevent clients from silently dropping review outcomes.
Use multiple weak signals carefully
Repeated submissions, rapid score flips for nearly identical images, inconsistent device metadata and a burst of attempts near the threshold can justify a review queue. None alone proves intent: a damaged camera can produce similar patterns. Restrict rate per identity, enforce input schema and image validity, and route uncertain cases to human review. Do not label every threshold crossing as an attack or claim input validation defeats adversarial examples. Containment policy ties signals to bounded action.
Test the same boundary in release reviews
Include controlled perturbations of legitimate parcel images in offline evaluation: crop changes, brightness changes and metadata inconsistencies that remain within the allowed capture workflow. Check whether the decision changes when the actual damage condition does not. Report both false negatives and false alerts by camera and parcel type. Limit experiments to authorized test data. The project demonstrates that a benign retake can look suspicious and must not be silently blocked.
Implementation
def classify_submission_pair(previous, current):
if previous["model_revision"] != current["model_revision"]:
return "incomparable:model-change"
if previous["image_digest"] == current["image_digest"]:
return "duplicate"
if previous["parcel_id"] != current["parcel_id"]:
return "different-parcel"
if abs(previous["score"] - current["score"]) >= 0.27:
return "review:score-flip"
return "observe"
first = {"parcel_id": "shipment-82", "model_revision": "scan-r47",
"image_digest": "sha-a", "score": 0.71}
second = {**first, "image_digest": "sha-b", "score": 0.39}
assert classify_submission_pair(first, second) == "review:score-flip"
assert classify_submission_pair(first, first) == "duplicate"
Performance and operating cost
Pair classification is O(1) time and space after candidate pairs are selected. Finding pairs across a high-volume stream requires keyed state and bounded retention; a naive all-pairs comparison is O(n²). A review queue imposes human cost, so measure false alerts and time-to-decision before broad enforcement.
Common Mistakes
- Calling every changed image an evasion attempt.
- Comparing scores from different model revisions.
- Keeping raw images indefinitely without a retention rule.
- Assuming schema checks prevent small visual perturbations.
Read next
- Inference evasion response: rate limits, review and reversible containment
- Project: investigate repeated parcel-scanner score flips
- Training source admission: provenance, trust tiers and quarantine
- Inference API contracts: version the decision, not only the payload
- Score thresholds are release policy, not model metadata
