Use paired views of unlabeled receipts to train an encoder, then test whether those features improve scarce-label receipt-quality decisions on untouched physical receipts.
Project: pretrain receipt features from unlabeled captures
Freeze identities and baselines
Create a manifest with physical receipt ID, capture session, device and any available defect label. Split by physical receipt before augmenting, and reserve the final test group from all pretraining, hyperparameter selection and nearest-neighbor inspection. Establish two downstream baselines: a quality model trained from scratch and a frozen generic feature encoder, both with the same labeled examples. The project is about incremental value from receipt-specific unlabeled captures, not about obtaining a low unsupervised loss. Split discipline applies to every unlabeled capture.
Specify the paired-view rule
Generate two mild views of each pretraining image while preserving visible clipped edges, blur and folds. Reject transforms that erase a defect or expose nonreceipt metadata as a shortcut. Sample at most one physical receipt per batch unless a multi-positive objective is implemented. Store fixed-seed view previews for a device and defect audit. The same source image must remain identifiable across views, but two views should not be byte-identical. The view contract defines allowed variation.
Train with an indexed loss
Pass both views through a shared encoder and projection head. Normalize projected vectors, mask self-similarity, map each anchor to its other view and calculate cross entropy. The snippet uses a small dense encoder and one step on synthetic features to demonstrate this contract; actual receipt images require image preprocessing and a suitable convolutional or attention encoder. Log global unique identities, positive similarities, apparent-negative similarities, embedding variance and loss. The loss lesson explains the denominator.
Measure transfer under scarce labels
Train a frozen linear head using the same limited label budget for every encoder. Then fine-tune an encoder separately, keeping the final test untouched. Report class recall, false accepts, confidence coverage and device breakdown. Compare results across at least two label budgets; a representation can help when labels are scarce and make little difference when labels are plentiful. Inspect nearest neighbors for false similarity based on paper color or camera border rather than quality. Do not select the winning temperature by the final test.
Deliver a reproducible artifact
Save the source manifest hash, view generator configuration, batch identity rule, encoder and head revisions, projection-head handling and downstream split IDs. The projection head used for the contrastive objective need not be the downstream feature API; document which tensor is exported. Add a fixed-batch reload check and a table of paired baselines. If held-out recall does not improve, publish the negative result as a boundary of the method rather than substituting an easier metric. The confidence audit catches cost shifted into human review.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
receipt_vectors = torch.randn(6, 16)
view_a = receipt_vectors + 0.03 * torch.randn_like(receipt_vectors)
view_b = receipt_vectors + 0.05 * torch.randn_like(receipt_vectors)
encoder = nn.Sequential(nn.Linear(16, 24), nn.ReLU(), nn.Linear(24, 12))
projection = nn.Linear(12, 8)
optimizer = torch.optim.AdamW(list(encoder.parameters()) +
list(projection.parameters()), lr=0.0007)
projected = functional.normalize(projection(encoder(torch.cat((view_a, view_b)))), dim=1)
scores = projected @ projected.T / 0.27
scores.fill_diagonal_(float("-inf"))
positive_index = (torch.arange(12) + 6) % 12
optimizer.zero_grad(set_to_none=True)
loss = functional.cross_entropy(scores, positive_index)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)
assert projected.shape == (12, 8)Performance and operating cost
The encoder runs once per view and stores both computational graphs for backpropagation. Pairwise scores require O(B²d) multiply work and O(B²) memory for B receipt identities and projection width d. The toy dense network is small; a real image encoder usually dominates compute. Preparing fixed-seed visual audits and downstream label-budget comparisons adds work, but those checks determine whether pretraining helped the product rather than merely optimizing a proxy loss.
Common Mistakes
- Do not pretrain on final-test receipt identities and call the later evaluation untouched.
- Do not export the projection tensor as the downstream feature without checking that choice.
- Do not use a view transform that removes the rare defect whose recall is being measured.
