Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: stress-test receipt decisions under bounded image changes

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Audit a receipt-quality classifier under numeric perturbations and realistic capture changes, then decide whether a stress-trained candidate actually improves the deployed decision.

Freeze the deployment path

Pin image decode, crop, normalization, model revision, class map, confidence threshold and review routing. Group all captures from one physical receipt before forming training, development and final-test populations. Collect clean held-out examples across devices and rare clipped-edge cases. Save baseline predictions and accepted-case counts before generating any perturbation. The quality project supplies the original classifier contract.

Create two distinct stress families

First, apply a bounded input-gradient perturbation with radii specified in zero-to-one pixel space, recording clean correctness and changed decisions. Second, use physically plausible capture changes such as measured blur, low exposure, perspective shift and partial edge occlusion. Do not call the second group adversarial examples unless its generation actually follows a search objective. Review altered images to ensure their ground-truth defect label still applies. The numeric test has a narrower claim than a camera stress test.

Train and evaluate a candidate

Train a clean reference and a candidate with a declared clean-plus-stress objective, equal real training identities and a documented update budget. Choose radius and loss weight on development data only. Test several search strengths, not merely the training radius; inspect gradients for finite values. Measure clean, numeric and natural-stress recall and false accepts by class and device. The training lesson explains why one improved stress score can conceal a clean regression.

Apply a predeclared gate

Set minimum clipped-edge recall and maximum false accepts separately for clean and stressed traffic, plus maximum manual-review volume. The code below evaluates illustrative measured slice records; it does not claim a trained model achieved them. If a slice has too few examples, report event counts and postpone a strong pass claim rather than hiding it in a global average. A confidence threshold may route uncertain images to review; count that workload explicitly. The review audit makes the denominator visible.

Ship a reversible decision

Package stress generator version, radii and input units, natural-corruption parameters, slice manifests, model revision, calibration and release-gate table. Keep the previous classifier available. Monitor real-world capture mix, false accepts and manual review after staged rollout; a lab perturbation suite cannot anticipate every camera failure. If a candidate passes numeric stress but harms clean or natural-stress decisions, do not release it. State the remaining untested threat boundaries in the artifact.

Implementation

python
from dataclasses import dataclass

@dataclass(frozen=True)
class DecisionSlice:
    name: str
    receipt_count: int
    clipped_recall: float
    false_accepts: int
    review_count: int

def failed_gates(slices: list[DecisionSlice]) -> list[str]:
    failures = []
    for result in slices:
        if result.receipt_count < 47:
            failures.append(result.name + ": insufficient evidence")
        if result.clipped_recall < 0.91:
            failures.append(result.name + ": clipped recall")
        if result.false_accepts > 3:
            failures.append(result.name + ": false accepts")
        if result.review_count > 19:
            failures.append(result.name + ": review capacity")
    return failures

audit_slices = [DecisionSlice("clean", 83, 0.95, 2, 12),
                DecisionSlice("low-exposure", 53, 0.90, 3, 17)]
assert failed_gates(audit_slices) == ["low-exposure: clipped recall"]

Performance and operating cost

For N held-out images and A attack settings, one-step gradient tests add roughly A model backward passes plus prediction forwards, while natural-corruption generation adds image-processing cost. The release-gate scan is O(S) for S slice records, but reliable slice measurement requires enough independent physical receipts. Storing all perturbed images can be expensive; storing seeds, source identities and transform parameters enables replay when the transform is deterministic. Budget manual review for label preservation and high-confidence wrong cases.

Common Mistakes

  • Do not merge clean and stressed scores into one average that hides a rare-class failure.
  • Do not leave input units or perturbation radius out of the audit artifact.
  • Do not call a one-step numeric stress pass a guarantee for physical camera changes.

Read next

ai-data
deep-learning
Storage details