A shared embedding can rank related image and text records only when pair identity and comparison candidates are trustworthy.
Cross-modal representation: positive pairs, negatives and shared space
Choose positives
An image and its verified receipt text form a positive pair. Two crops from one receipt may share the parent identity but should not automatically be treated as distinct independent examples. Track duplicate scans and near-identical documents so a train-test split does not place copies on both sides. Pair identity precedes any contrastive objective.
Build plausible negatives
A random unrelated receipt is an easy negative. A useful evaluation asks whether a model can distinguish two receipts from the same merchant with similar totals and layout. Do not label a candidate negative if it could be another valid description of the same asset. Record the candidate-generation rule because retrieval scores change with the pool.
Control representation versions
An image encoder and text encoder must write vectors into one compatible space under a known model version. Replacing only the text encoder can make existing image vectors incomparable. Normalize vectors consistently if cosine similarity is the chosen measure. A vector with no provenance is unsafe for debugging or deletion, even if the nearest-neighbor search still returns a result.
Test a ranking
Pair image I-47 with text T-47, then include T-48 from the same merchant and T-93 from another merchant. The correct text should rank first under the tested model, but the evaluation should also show the margin to T-48. Repeat with a blurred image; a low margin should be visible rather than presented as a certain match.
Implementation
from math import sqrt
def cosine_score(image_vector, text_vector):
if len(image_vector) != len(text_vector) or not image_vector:
return None
image_norm = sqrt(sum(value * value for value in image_vector))
text_norm = sqrt(sum(value * value for value in text_vector))
if image_norm == 0 or text_norm == 0:
return None
return sum(left * right for left, right in zip(image_vector, text_vector)) / (image_norm * text_norm)Performance and operating cost
One similarity score costs O(D) time and O(1) auxiliary space for vector dimension D. Scoring C candidates directly costs O(C times D); an index reduces search work but adds build, versioning and recall checks.
Common Mistakes
- Do not treat every unpaired item as a verified negative.
- Do not compare vectors produced by incompatible encoder versions.
- Do not split duplicate assets across training and evaluation.
Read next
- Multimodal data contract: paired records, identity and consent
- Cross-modal retrieval: candidate pools, ranking and versioned indexes
- Retrieval-Based AI Tutorial
- Vision evaluation splits: group captures and test acquisition shift
Continue the workflow: Triplet margin and normalized embedding geometry.
