Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Inference-time evasion: threat model and input evidence

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An inference evasion attempt changes submitted inputs to obtain a favorable prediction without changing the real-world condition.

Separate adversarial behavior from drift

A parcel scanner model routes damaged packages for review. An operator can submit altered images or metadata to push a damaged parcel below the review threshold. That is an inference-time manipulation concern, distinct from contaminated training data and ordinary camera drift. Define attacker control over pixels, crop, device metadata, retries and account identity. Record the decision the attacker wants and what evidence the service can actually observe. Training-source admission addresses a different stage of the lifecycle.

Preserve a minimal decision trail

Retain a governed request ID, account or device pseudonym, input digest, model revision, preprocessing revision, score band, decision and review outcome where permitted. Keep raw images only under a documented retention policy; an input digest alone cannot explain visual changes. The operational question is whether repeated near-identical submissions from one actor cross the threshold after small edits. The response contract should pin the decision revision and prevent clients from silently dropping review outcomes.

Use multiple weak signals carefully

Repeated submissions, rapid score flips for nearly identical images, inconsistent device metadata and a burst of attempts near the threshold can justify a review queue. None alone proves intent: a damaged camera can produce similar patterns. Restrict rate per identity, enforce input schema and image validity, and route uncertain cases to human review. Do not label every threshold crossing as an attack or claim input validation defeats adversarial examples. Containment policy ties signals to bounded action.

Test the same boundary in release reviews

Include controlled perturbations of legitimate parcel images in offline evaluation: crop changes, brightness changes and metadata inconsistencies that remain within the allowed capture workflow. Check whether the decision changes when the actual damage condition does not. Report both false negatives and false alerts by camera and parcel type. Limit experiments to authorized test data. The project demonstrates that a benign retake can look suspicious and must not be silently blocked.

Implementation

python
def classify_submission_pair(previous, current):
    if previous["model_revision"] != current["model_revision"]:
        return "incomparable:model-change"
    if previous["image_digest"] == current["image_digest"]:
        return "duplicate"
    if previous["parcel_id"] != current["parcel_id"]:
        return "different-parcel"
    if abs(previous["score"] - current["score"]) >= 0.27:
        return "review:score-flip"
    return "observe"

first = {"parcel_id": "shipment-82", "model_revision": "scan-r47",
         "image_digest": "sha-a", "score": 0.71}
second = {**first, "image_digest": "sha-b", "score": 0.39}
assert classify_submission_pair(first, second) == "review:score-flip"
assert classify_submission_pair(first, first) == "duplicate"

Performance and operating cost

Pair classification is O(1) time and space after candidate pairs are selected. Finding pairs across a high-volume stream requires keyed state and bounded retention; a naive all-pairs comparison is O(n²). A review queue imposes human cost, so measure false alerts and time-to-decision before broad enforcement.

Common Mistakes

  • Calling every changed image an evasion attempt.
  • Comparing scores from different model revisions.
  • Keeping raw images indefinitely without a retention rule.
  • Assuming schema checks prevent small visual perturbations.

Read next

ai-data
mlops
Storage details