Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: investigate repeated parcel-scanner score flips

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a small evidence ledger, distinguish retakes from suspicious submissions and exercise a reversible review rule.

Define an authorized investigation

A parcel scanner clears low-risk packages automatically. In a test cohort, one parcel receives several distinct images with scores on both sides of the damage-review threshold. Pin the model and preprocessing revision, request IDs, image digests, timestamps, camera class and outcome labels. The team has access to test images only under a short retention policy. The threat model limits claims to what these records support; intent is not directly observed.

Sort retries, retakes and candidate manipulation

First collapse exact-image retries by digest and request identity. Then group distinct images by parcel and actor within a fixed window. A camera with a faulty flash creates a benign retake and a score shift; a separate sequence with repeated threshold crossings deserves review. Put both into the evidence ledger with reasons, not an attack verdict. Check input validity and metadata consistency without assuming that passing those checks makes an image safe.

Apply a proportional policy

When the distinct-submission count and score span both exceed approved cutoffs, hold automated clearance for that parcel and send it to a staffed review queue. Keep ingestion open under a rate cap so an operator can provide a better capture. Set a hold expiry and route for appeal. The benign flash retake must clear after review; the unresolved repeated flips stay on temporary hold. Containment is a reversible workflow, not an automatic accusation.

Publish the case packet

Report the count of exact retries, distinct images, alert rule revision, review results, false alerts and queue delay. Re-run the case after a model update but keep scores separated by model revision. Rehearse disabling the rule if review age exceeds the service budget. Connect the case to incident evidence and threshold change control before any permanent model or policy adjustment.

Implementation

python
def parcel_case(events, max_distinct, required_span):
    current_revisions = {event["model_revision"] for event in events}
    if len(current_revisions) != 1:
        return "hold:mixed-model-evidence"
    by_digest = {event["image_digest"]: event for event in events}
    scores = [event["score"] for event in by_digest.values()]
    if len(by_digest) > max_distinct and max(scores) - min(scores) >= required_span:
        return "review:parcel-only"
    return "serve:normal"

cases = [{"model_revision": "scan-r47", "image_digest": "a", "score": 0.72},
         {"model_revision": "scan-r47", "image_digest": "b", "score": 0.69},
         {"model_revision": "scan-r47", "image_digest": "c", "score": 0.34}]
assert parcel_case(cases, 2, 0.27) == "review:parcel-only"
assert parcel_case(cases[:2], 2, 0.27) == "serve:normal"

Performance and operating cost

The case pass is O(n) time and O(n) space for n events; keyed retention bounds live state. Human review and temporary holds add latency, so the packet must record both investigation value and legitimate-service impact. Keeping only digests reduces raw-image storage but limits later visual analysis.

Common Mistakes

  • Mixing scores from two model revisions in one score-span calculation.
  • Treating an exact retry as a distinct image.
  • Issuing a permanent ban from a threshold-crossing heuristic.
  • Omitting hold expiry and the result of human review.

Read next

ai-data
mlops
Storage details