Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: adapt a receipt model to a new capture device

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Compare a frozen visual backbone with a staged late-block update on a new device population, using explicit regression limits and a reloadable serving artifact.

Build the two-domain manifest

Collect 470 scanned receipts and 137 mobile captures with readable, blurred and cut-off labels. Group by physical receipt and merchant layout before splitting; multiple photos of the same paper cannot cross train and validation. Reserve an untouched test set from each domain. Record counts by class and capture device so that a result based on three cutoff examples is not reported as stable evidence. The base project defines the quality task.

Create the frozen reference

Load a pinned pretrained backbone, replace its head and train only that head with the backbone held in eval mode. Use the preprocessing associated with the selected weight identity and persist it. Record training cost, old-domain and new-domain confusion matrices and the chosen threshold. This becomes the rollback candidate, not just a number in a notebook. Frozen versus adapted state defines the reference.

Run one controlled adaptation

Unfreeze only the final block and create separate learning-rate groups for the block and head. Keep the identity split, preprocessing, labels and validation schedule fixed. Declare an acceptance rule before training: for example, improve mobile-capture recall by at least four percentage points while allowing no more than one percentage point loss on scanned-receipt recall. Because the mobile group is small, include counts and uncertainty; the rule is a product decision, not a statistical guarantee.

Audit the hidden state

Before each stage, print the trainable parameter names and check BatchNorm mode for the frozen and unfrozen blocks. Save a full state dictionary and reload it into a newly created model. Compare logits for a fixed scanned and mobile image under eval mode. Break the preprocessing once on purpose, such as swapping channel order, and confirm an input-contract test catches it. Running statistics can change predictions without a trainable gradient.

Publish a decision record

Show both-domain confusion matrices, latency, memory, training-time cost, acceptance-rule result and a list of the worst mobile and scan errors. If the adapted model misses the old-domain limit, ship the frozen candidate and document which slice failed. Package weights, architecture revision, weight identity, preprocessing, label map and threshold. A deployment smoke test should reload from that package and reproduce fixed-batch logits before the traffic switch.

Implementation

python
import torch
from torch import nn

def recall_from_counts(true_positive: int, false_negative: int) -> float:
    denominator = true_positive + false_negative
    if denominator == 0:
        raise ValueError("recall requires at least one positive case")
    return true_positive / denominator

baseline_scan_recall = recall_from_counts(90, 10)
adapted_scan_recall = recall_from_counts(89, 11)
baseline_mobile_recall = recall_from_counts(54, 16)
adapted_mobile_recall = recall_from_counts(59, 11)
mobile_gain = adapted_mobile_recall - baseline_mobile_recall
scan_regression = baseline_scan_recall - adapted_scan_recall
accepted = mobile_gain >= 0.04 and scan_regression <= 0.010001
assert accepted

class QualityHead(nn.Module):
    def __init__(self, input_width: int):
        super().__init__()
        self.classifier = nn.Linear(input_width, 3)

    def forward(self, backbone_features: torch.Tensor) -> torch.Tensor:
        return self.classifier(backbone_features)

artifact = {"label_map": {0: "readable", 1: "blurred", 2: "cut_off"},
            "scan_recall": adapted_scan_recall,
            "mobile_recall": adapted_mobile_recall}
assert set(artifact["label_map"]) == {0, 1, 2}

Performance and operating cost

The frozen run spends compute on backbone forward passes and head updates. Late-block tuning adds backward compute and activation memory for that block. Evaluation now covers two populations and slices within each; its cost scales with the total examples. Keep numerator and denominator beside each rate, because a one-case change can matter more than a rounded percentage in a 137-image domain.

Common Mistakes

  • Do not put two captures of one physical receipt on opposite sides of a split.
  • Do not ship the adapted model because only the new-domain mean improved.
  • Do not save weights without the preprocessing and threshold that produced the accepted metrics.

Read next

ai-data
deep-learning
Storage details