Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Cross-modal retrieval: candidate pools, ranking and versioned indexes

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Image-to-text or text-to-image retrieval is a ranking task whose result depends on candidate eligibility, index freshness and duplicate control.

Define the query

A reviewer may use an image crop to find the matching receipt transcript. Query and candidate are different media types, but the answer still needs a parent receipt ID, page alignment and a permitted-use check. A withdrawn image or text asset must be filtered before ranking, even if its embedding remains in a reproducibility snapshot. Retrieval design separates candidate supply from final ranking.

Build a hard pool

Measure recall against candidates from similar merchants, layouts and totals, not only random unrelated receipts. If the true transcript is absent from the index, the ranker cannot recover it. Report index coverage and candidate recall before top-K accuracy. Deduplicate multiple OCR versions or define which version is eligible; otherwise several near-copies can occupy the entire result list.

Version the index

Record both encoder versions, vector normalization, index build cutoff and deletion-filter version. A new text encoder requires compatible image vectors or a planned dual-index transition. Shared-space compatibility is measurable: replay fixed query pairs and compare rank and score distributions before switching traffic.

Inspect one query

For image I-47, return the correct transcript at rank two behind a same-merchant near-duplicate. The report should show candidate IDs, scores, duplicate group and why the top result was wrong. Rebuild the index after deleting that near-duplicate and confirm the deletion gate acts immediately, even before full reindexing finishes.

Implementation

python
def recall_at_k(ranked_candidate_ids, relevant_ids, limit):
    if limit < 1:
        raise ValueError("limit must be positive")
    relevant = set(relevant_ids)
    if not relevant:
        return None
    return len(set(ranked_candidate_ids[:limit]) & relevant) / len(relevant)

Performance and operating cost

Top-K recall costs O(K + R) time and space for R relevant candidates. Exact scoring over C D-dimensional vectors costs O(C times D); approximate indexes trade build cost and possible recall loss for faster search.

Common Mistakes

  • Do not score a ranker when the correct item was never indexed.
  • Do not serve a withdrawn asset because its vector still exists.
  • Do not compare embeddings from incompatible model versions.

Read next

ai-data
multimodal-ai
Storage details