Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: route uncertain receipt defects to human review

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Compare a single classifier, independent ensemble and stochastic-dropout pass under identical receipt splits and a fixed review budget.

Build a defect ledger

Keep receipt ID, store, scanner, capture session, reviewed defect class and review outcome. Split by store and acquisition period before augmentation, then hold back a new scanner or later store group for final evaluation. Include both severe and cosmetic damage, reflective paper, blur and partial scans. The split rule prevents near-identical images from appearing on both sides of the final comparison.

Fit and compare candidates

Train independent ensemble members with separate initialization seeds and documented shared data. Compare against the strongest single model and a model explicitly trained with dropout. The code aggregates three synthetic probability vectors and computes a disagreement score; it is a wiring check, not a trained detector. The ensemble lesson explains what member variation can and cannot reveal.

Choose a review rule once

On development data, set a score threshold that fits the daily review capacity. Report the resulting auto-approval count, severe-defect recall, risk among auto-approved cases and errors caught by reviewers. Then freeze the rule for untouched stores and scanners. If a review queue exceeds staffing, the model must abstain or slow its auto-approval path according to an agreed policy, not lower the threshold after seeing final labels.

Replay full latency

Measure image decode, preprocessing, all model forwards, aggregation and queue submission at intended concurrency. A seven-pass dropout path may be slower than an ensemble batched on a GPU. Keep batch normalization fixed for stochastic inference and test request isolation. The mode lesson prevents a serving worker from changing model state during another request.

Release with explicit failure handling

Gate on severe-defect recall, auto-approved risk, review capacity and p95 response time by scanner slice. Route unreadable images and missing metadata to manual review rather than to a confident negative class. Package model member revisions, score calibration, threshold, preprocessing and rollback pointer. Review disagreements after launch and retrain only after a new independent evaluation split is prepared.

Implementation

python
from math import log

receipt_votes = [(.73, .18, .09), (.61, .28, .11), (.24, .65, .11)]
mean_vote = tuple(sum(vote[class_id] for vote in receipt_votes) / len(receipt_votes)
                  for class_id in range(3))

def uncertainty(distribution):
    return -sum(probability * log(probability) for probability in distribution
                if probability > 0)

predictive = uncertainty(mean_vote)
member_average = sum(map(uncertainty, receipt_votes)) / len(receipt_votes)
disagreement = predictive - member_average
review_threshold = 0.08
route_to_review = disagreement >= review_threshold
assert len(mean_vote) == 3 and abs(sum(mean_vote) - 1) < 1e-9
assert disagreement >= -1e-12
assert isinstance(route_to_review, bool)

Performance and operating cost

Aggregating M member predictions for C classes takes O(MC) arithmetic and O(C) accumulator space. Training and serving M full models are the dominant costs; dropout uses one model but P stochastic forwards. A review queue adds human time that can exceed compute cost, so measure errors caught per reviewed receipt and daily queue volume. The synthetic vectors demonstrate aggregation only, not a validated threshold or detector.

Common Mistakes

  • Do not tune a review threshold on the untouched store test set.
  • Do not approve malformed images as clean because every model agrees.
  • Do not omit human review capacity from the operating cost.

Read next

ai-data
deep-learning
Storage details