A self-supervised image pair is useful only when both views preserve the signal the downstream model must recognize and the split respects the original item.
Contrastive view identity and augmentation contracts
Name the unit of identity
Two crops from one photographed receipt are two views of one physical receipt, not two independent examples. Store a stable receipt identity before generating any crop, rotation, exposure change or compression artifact. Group the train and held-out split by that identity and, where possible, by capture session. A validation image copied into training under another filename will make representation quality appear better than it is. The augmentation boundary applies before labels are considered, because a self-supervised objective can memorize near-duplicates too.
Define which changes preserve meaning
A mild change in brightness may preserve a visible fold or cut-off edge. A crop that removes the edge does not. If the downstream task detects clipped receipts, making an original and a crop-without-the-clipping into a positive pair teaches the encoder to ignore the very defect of interest. Write a task-specific invariance table: allowed exposure shift, small perspective warp, blur range and crop coverage. Reject a proposed view transform when a reviewer cannot still identify the same defect. Generic image augmentation defaults are not a contract for document quality.
Separate pairing from class labels
In a basic instance-discrimination task, the two views of the same source are positive; views of other sources in the batch act as negatives. That rule says nothing about whether two different receipts share a quality class. Pushing all other receipts apart may waste capacity or create false negatives, especially if multiple captures of one physical receipt have inconsistent identifiers. Track source and session IDs so obvious false negatives can be masked or the batch can be sampled differently. The contrastive loss needs a precise positive index for every anchor.
Audit the view generator
Sample and save paired previews at fixed seeds across lighting, device and defect slices. Calculate how often the generator produces an empty crop, saturates the image, erases a cut-off edge or makes both views nearly identical. Keep the generator configuration and code revision with the pretrained checkpoint. A model can optimize a shortcut in view-generation artifacts instead of learning receipt structure; paired visual audits and downstream ablations are more informative than one training loss curve. Record source ID, seed and transform parameters for failure replay.
Evaluate transfer after pretraining
Freeze the encoder and train a small quality head on the labeled training population; then test a fine-tuned encoder under the same split. Compare against an encoder trained from scratch with equal downstream labels and compute budget. Report each defect slice, device slice and calibration, not just an average score. A contrastive objective can produce well-separated training embeddings yet fail to improve the decision the product makes. The project ties view design to that release test.
Implementation
from collections import defaultdict
receipt_records = [
{"receipt_id": "r-047", "capture_id": "front", "split": "train"},
{"receipt_id": "r-047", "capture_id": "tilted", "split": "train"},
{"receipt_id": "r-083", "capture_id": "front", "split": "holdout"},
{"receipt_id": "r-083", "capture_id": "low-light", "split": "holdout"},
]
splits_by_receipt = defaultdict(set)
for record in receipt_records:
splits_by_receipt[record["receipt_id"]].add(record["split"])
assert all(len(assigned_splits) == 1 for assigned_splits in splits_by_receipt.values())
def pair_captures(records):
grouped = defaultdict(list)
for record in records:
grouped[record["receipt_id"]].append(record["capture_id"])
return {receipt_id: tuple(captures) for receipt_id, captures in grouped.items()
if len(captures) >= 2}
positive_sources = pair_captures(receipt_records)
assert positive_sources["r-047"] == ("front", "tilted")Performance and operating cost
Generating two image views approximately doubles image preprocessing and encoder forward/backward work relative to one-view training. Storage can stay near one copy of each source if views are generated on demand, but deterministic replay requires seeds and transform records. Pair lookup and split auditing are O(N) in source records. Downstream evaluation is additional cost, not optional overhead: without it the pretrained representation has no verified application value.
Common Mistakes
- Do not split two captures of one physical receipt across train and holdout.
- Do not crop away a target defect and call the altered image an equivalent positive view.
- Do not interpret other instances in a batch as guaranteed semantic negatives.
Read next
- Image augmentation: split originals first and preserve the label
- Contrastive temperature and negative-mask accounting
- Project: pretrain receipt features from unlabeled captures
- Project: localize receipt defects with a residual CNN
- Frozen features versus fine-tuning a pretrained backbone
Continue the workflow: Audio time-frequency masks and label preservation.
