Contrastive training pulls accepted views together and contrasts other views, but a batch member can be a false negative when it shows the same entity or condition.
Contrastive pairs, batch negatives and identity collisions
Define positives with care
For warehouse images, two valid views of one parcel can be positive. Record their parcel identity and capture interval. A later image may show a newly created tear; that view is not automatically equivalent to an earlier undamaged image. The positive rule follows the view contract, not the filename pattern.
Audit the denominator
Contrastive losses often treat other batch items as negatives. Another frame of the same parcel can therefore become a false negative and be pushed away from its sibling. Different parcels with identical damage may also be semantically close, even though instance discrimination treats them as distinct. The code removes known same-parcel collisions; it cannot discover unknown semantic collisions.
Do not reward scanner identity
If positive images share a barcode or scanner watermark, the encoder can solve the pretext task using those marks while ignoring damage. Crop or redact operational identifiers when safe, and test a barcode-only baseline. If redaction removes real defect evidence, change the capture process or objective rather than claiming clean invariance.
Use difficult negatives sparingly
Selecting only nearest neighbors as negatives increases the chance of choosing unlabeled examples that genuinely share the target condition. Review hard pairs and monitor false-negative rates by damage type. Hard-negative review supplies an analogous retrieval failure analysis.
Compare downstream slices
A useful contrastive encoder should help a frozen probe or a carefully tuned downstream model on later parcels, including rare crushed-corner cases and new depots. Evaluate against no-pretraining and masked-reconstruction baselines with the same labeled split. Probe design makes the comparison accountable.
Implementation
batch = [
{"view": "47-a", "parcel_id": "dock-47"},
{"view": "47-b", "parcel_id": "dock-47"},
{"view": "62-a", "parcel_id": "dock-62"},
{"view": "83-a", "parcel_id": "dock-83"},
]
def safe_negative_views(anchor, candidates):
return [candidate["view"] for candidate in candidates
if candidate["view"] != anchor["view"]
and candidate["parcel_id"] != anchor["parcel_id"]]
assert safe_negative_views(batch[0], batch) == ["62-a", "83-a"]
assert "47-b" not in safe_negative_views(batch[0], batch)Performance and operating cost
For B views, a naive all-pairs similarity matrix costs O(B²) time and memory before encoder work. Filtering known identity collisions can be O(B²) if done per anchor as shown; grouping by identity reduces bookkeeping cost. Large negative pools also raise review and storage costs.
Common Mistakes
- Do not treat another view of the same entity as a negative.
- Do not assume an instance-negative is semantically different.
- Do not hide the cost of larger batches or memory queues.
