A two-view batch needs an exact positive index, a self-comparison mask and an explicit policy for duplicate identities before its loss means what the team thinks it means.
Contrastive temperature and negative-mask accounting
Build the similarity matrix
Encode B source images twice, producing 2B projected embeddings. Normalize each vector, then form pairwise dot products; with normalized vectors those are cosine similarities. Divide by a positive temperature before applying a stable softmax cross entropy. The matched view for the first B anchors appears B positions later, while the second B anchors match B positions earlier. The diagonal is not a negative: an image has perfect similarity to itself and must be excluded. View identity determines those indices before tensor arithmetic.
Account for negatives correctly
For each anchor, the denominator contains all other views except itself. It includes the matched positive once and the remaining 2B minus two views as apparent negatives. If two batch entries are captures of the same physical receipt, those apparent negatives can contradict the positive objective. Choose a batch sampler that keeps one source identity per batch, or construct an explicit multi-positive/masked loss. Merely hiding the duplicate from the positive index leaves it in the denominator and still pushes it away.
Understand temperature without folklore
A smaller positive temperature makes a fixed similarity difference affect the softmax more strongly; gradients concentrate around difficult comparisons. Too small a value can make optimization brittle or encourage attention to mislabeled negatives. A larger value softens those comparisons but may produce weak separation. Select temperature on a development set using downstream transfer and representation diagnostics, not the lowest pretraining loss alone. Keep the projection head and encoder roles distinct: the head can absorb objective-specific geometry while the encoder supplies downstream features.
Watch batch and distributed effects
The similarity matrix is quadratic in the number of views. Larger batches expose more negatives but increase memory and the chance of false negatives. In distributed training, gathering embeddings across ranks changes both the denominator and gradient flow; an implementation that gathers detached peer embeddings is not equivalent to one with fully differentiable gathering. Count unique source identities globally, define whether repeated sampler tail items contribute, and test a tiny reference batch. Sampler tails alter this objective in a less obvious way than a supervised mean loss.
Check representation collapse and shortcuts
Track embedding variance per dimension, nearest-neighbor identities and positive-versus-negative similarity distributions. If every vector becomes almost identical, the objective or masking may be broken; if similarity separates only device type, the encoder may exploit capture metadata instead of receipt quality. Evaluate a frozen linear probe on labeled training data and a held-out physical-receipt set. The code below checks the exact two-view index algebra and computes one stable cross entropy. It is not a complete training system.
Implementation
import torch
from torch.nn import functional as functional
torch.manual_seed(47)
source_embeddings = torch.randn(5, 12)
first_views = source_embeddings + 0.04 * torch.randn_like(source_embeddings)
second_views = source_embeddings + 0.04 * torch.randn_like(source_embeddings)
projected_views = functional.normalize(
torch.cat((first_views, second_views), dim=0), dim=1)
temperature = 0.23
similarities = projected_views @ projected_views.T / temperature
self_mask = torch.eye(len(projected_views), dtype=torch.bool)
similarities = similarities.masked_fill(self_mask, float("-inf"))
positive_indices = (torch.arange(len(projected_views)) + 5) % 10
contrastive_loss = functional.cross_entropy(similarities, positive_indices)
assert torch.isfinite(contrastive_loss)
assert positive_indices.tolist() == [5, 6, 7, 8, 9, 0, 1, 2, 3, 4]Performance and operating cost
For B identities and d-dimensional projections, the similarity multiply costs O(B²d) time and O(B²) matrix memory. Two encoder views also require roughly twice the encoding work of one-view training. This small dense implementation is useful for checking indexing, but a large multiworker job may need chunked similarities or another negative-sampling design. Temperature changes arithmetic, not asymptotic cost; a larger global batch can dominate memory even if the encoder itself fits.
Common Mistakes
- Do not include the diagonal self-similarity as a negative or positive.
- Do not allow duplicate physical receipts to become unexamined false negatives.
- Do not compare losses at different temperatures as if they were on one fixed scale.
Read next
- Contrastive view identity and augmentation contracts
- Project: pretrain receipt features from unlabeled captures
- Distributed sampler tails and global loss weighting
- Batching and class sampling: know the population the optimizer sees
- Frozen features versus fine-tuning a pretrained backbone
Continue the workflow: Two-tower embeddings and in-batch negative contracts.
