Compare a single classifier, independent ensemble and stochastic-dropout pass under identical receipt splits and a fixed review budget.
Project: route uncertain receipt defects to human review
Build a defect ledger
Keep receipt ID, store, scanner, capture session, reviewed defect class and review outcome. Split by store and acquisition period before augmentation, then hold back a new scanner or later store group for final evaluation. Include both severe and cosmetic damage, reflective paper, blur and partial scans. The split rule prevents near-identical images from appearing on both sides of the final comparison.
Fit and compare candidates
Train independent ensemble members with separate initialization seeds and documented shared data. Compare against the strongest single model and a model explicitly trained with dropout. The code aggregates three synthetic probability vectors and computes a disagreement score; it is a wiring check, not a trained detector. The ensemble lesson explains what member variation can and cannot reveal.
Choose a review rule once
On development data, set a score threshold that fits the daily review capacity. Report the resulting auto-approval count, severe-defect recall, risk among auto-approved cases and errors caught by reviewers. Then freeze the rule for untouched stores and scanners. If a review queue exceeds staffing, the model must abstain or slow its auto-approval path according to an agreed policy, not lower the threshold after seeing final labels.
Replay full latency
Measure image decode, preprocessing, all model forwards, aggregation and queue submission at intended concurrency. A seven-pass dropout path may be slower than an ensemble batched on a GPU. Keep batch normalization fixed for stochastic inference and test request isolation. The mode lesson prevents a serving worker from changing model state during another request.
Release with explicit failure handling
Gate on severe-defect recall, auto-approved risk, review capacity and p95 response time by scanner slice. Route unreadable images and missing metadata to manual review rather than to a confident negative class. Package model member revisions, score calibration, threshold, preprocessing and rollback pointer. Review disagreements after launch and retrain only after a new independent evaluation split is prepared.
Implementation
from math import log
receipt_votes = [(.73, .18, .09), (.61, .28, .11), (.24, .65, .11)]
mean_vote = tuple(sum(vote[class_id] for vote in receipt_votes) / len(receipt_votes)
for class_id in range(3))
def uncertainty(distribution):
return -sum(probability * log(probability) for probability in distribution
if probability > 0)
predictive = uncertainty(mean_vote)
member_average = sum(map(uncertainty, receipt_votes)) / len(receipt_votes)
disagreement = predictive - member_average
review_threshold = 0.08
route_to_review = disagreement >= review_threshold
assert len(mean_vote) == 3 and abs(sum(mean_vote) - 1) < 1e-9
assert disagreement >= -1e-12
assert isinstance(route_to_review, bool)Performance and operating cost
Aggregating M member predictions for C classes takes O(MC) arithmetic and O(C) accumulator space. Training and serving M full models are the dominant costs; dropout uses one model but P stochastic forwards. A review queue adds human time that can exceed compute cost, so measure errors caught per reviewed receipt and daily queue volume. The synthetic vectors demonstrate aggregation only, not a validated threshold or detector.
Common Mistakes
- Do not tune a review threshold on the untouched store test set.
- Do not approve malformed images as clean because every model agrees.
- Do not omit human review capacity from the operating cost.
