Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit receipt confidence and review handoff

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Calibrate a receipt-quality classifier, choose an abstention threshold under a review budget and test whether its visual sensitivity follows receipt evidence.

Freeze the evaluation populations

Split by physical receipt before producing multiple crops or captures. Keep training, calibration, threshold-selection and final test populations separate. Include readable, blurred and cut-off receipts from more than one capture device. Record class counts and device counts; a good aggregate score can hide a small device group with dangerous confident mistakes. The quality classifier supplies logits while this project evaluates decisions made from them.

Fit one positive temperature

Freeze model weights and fit a single log-temperature parameter on calibration logits only. Verify top-class identities do not change and compare negative log likelihood, reliability bins and high-confidence error counts before and after. A six-example demonstration is not enough to establish calibration; use a population with enough cases in each important slice. Temperature scaling adjusts sharpness, not classification accuracy.

Choose a review threshold

On a separate threshold-selection group, evaluate a grid of maximum-probability cutoffs. Choose the lowest cutoff meeting a predeclared accepted-case error target while keeping review volume within staffed capacity. On the final test population report coverage, selective error, class recall and manual reviews per hundred receipts. If no cutoff meets both constraints, do not invent a number; return the capacity or model-quality conflict as the project result.

Audit image evidence

Probe a sample of confident correct and confident wrong receipts with two plausible occlusion baselines. Measure target-logit change when replacing defect regions and equally sized background regions. Repeat after randomizing model weights as a sanity check. A heatmap that looks aligned with text but barely changes when weights change is not adequate evidence. The attribution protocol specifies patch size, target logit and baseline.

Ship a reversible decision artifact

Package model weights, preprocessing, class map, temperature, review threshold, model revision and a fallback for nonfinite logits or missing image fields. Reload the package in a clean process and compare fixed-batch decisions. Monitor coverage, review queue, high-confidence wrong cases and source-device mix. The outcome may be a request for better labels or a larger review team rather than an automatic release; preserve the previous artifact until the new policy meets the stated limits.

Implementation

python
import torch
from torch.nn import functional as functional

quality_logits = torch.tensor([[2.8, 0.3, -0.4], [0.5, 2.1, 0.1],
                               [0.9, 1.1, 0.3], [0.2, 0.8, 1.7],
                               [1.6, 0.9, 0.1], [0.3, 1.2, 0.6]])
quality_labels = torch.tensor([0, 1, 2, 2, 0, 1])
saved_temperature = torch.tensor(1.43)
probabilities = functional.softmax(quality_logits / saved_temperature, dim=1)
confidence, predictions = probabilities.max(dim=1)
wrong = predictions.ne(quality_labels)

def review_metrics(minimum_confidence: float) -> tuple[float, float, int]:
    accepted = confidence >= minimum_confidence
    accepted_count = int(accepted.sum())
    if accepted_count == 0:
        return 0.0, float("nan"), len(confidence)
    coverage = accepted_count / len(confidence)
    selective_error = float(wrong[accepted].float().mean())
    review_count = len(confidence) - accepted_count
    return coverage, selective_error, review_count

candidate_policies = {cutoff: review_metrics(cutoff)
                      for cutoff in (0.37, 0.53, 0.71)}
assert all(0 <= coverage <= 1 for coverage, _, _ in candidate_policies.values())
assert all(0 <= count <= 6 for _, _, count in candidate_policies.values())

Performance and operating cost

Temperature fitting is cheap compared with classifier training, and serving adds one division and softmax. Review workload can dominate operating cost: report accepted and routed counts at real traffic volume. Occlusion probes multiply inference by the number of tested patches, so keep them in offline audit or an explicitly budgeted diagnostic path. The six-case code illustrates metric definitions and does not justify a production cutoff.

Common Mistakes

  • Do not choose and report a cutoff on the same final test group.
  • Do not report a low selective error while concealing near-zero coverage.
  • Do not treat a visually plausible attribution map as a substitute for held-out decision checks.

Read next

Continue the workflow: Project: train and audit a receipt-background denoiser.

Continue the workflow: Project: release a compact receipt model to an edge device.

Continue the workflow: Project: distill and release a receipt-quality student.

Continue the workflow: Project: stress-test receipt decisions under bounded image changes.

Continue the workflow: Project: detect conveyor jams from timestamped video clips.

Continue the workflow: Project: inspect pallet damage from 3D point scans.

Continue the workflow: Project: route uncertain receipt defects to human review.

ai-data
deep-learning
Storage details