Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Cross-modal representation: positive pairs, negatives and shared space

Last updated: 7 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A shared embedding can rank related image and text records only when pair identity and comparison candidates are trustworthy.

Choose positives

An image and its verified receipt text form a positive pair. Two crops from one receipt may share the parent identity but should not automatically be treated as distinct independent examples. Track duplicate scans and near-identical documents so a train-test split does not place copies on both sides. Pair identity precedes any contrastive objective.

Build plausible negatives

A random unrelated receipt is an easy negative. A useful evaluation asks whether a model can distinguish two receipts from the same merchant with similar totals and layout. Do not label a candidate negative if it could be another valid description of the same asset. Record the candidate-generation rule because retrieval scores change with the pool.

Control representation versions

An image encoder and text encoder must write vectors into one compatible space under a known model version. Replacing only the text encoder can make existing image vectors incomparable. Normalize vectors consistently if cosine similarity is the chosen measure. A vector with no provenance is unsafe for debugging or deletion, even if the nearest-neighbor search still returns a result.

Test a ranking

Pair image I-47 with text T-47, then include T-48 from the same merchant and T-93 from another merchant. The correct text should rank first under the tested model, but the evaluation should also show the margin to T-48. Repeat with a blurred image; a low margin should be visible rather than presented as a certain match.

Implementation

python
from math import sqrt

def cosine_score(image_vector, text_vector):
    if len(image_vector) != len(text_vector) or not image_vector:
        return None
    image_norm = sqrt(sum(value * value for value in image_vector))
    text_norm = sqrt(sum(value * value for value in text_vector))
    if image_norm == 0 or text_norm == 0:
        return None
    return sum(left * right for left, right in zip(image_vector, text_vector)) / (image_norm * text_norm)

Performance and operating cost

One similarity score costs O(D) time and O(1) auxiliary space for vector dimension D. Scoring C candidates directly costs O(C times D); an index reduces search work but adds build, versioning and recall checks.

Common Mistakes

  • Do not treat every unpaired item as a verified negative.
  • Do not compare vectors produced by incompatible encoder versions.
  • Do not split duplicate assets across training and evaluation.

Read next

Continue the workflow: Triplet margin and normalized embedding geometry.

ai-data
multimodal-ai
Storage details