Image-to-text or text-to-image retrieval is a ranking task whose result depends on candidate eligibility, index freshness and duplicate control.
Cross-modal retrieval: candidate pools, ranking and versioned indexes
Define the query
A reviewer may use an image crop to find the matching receipt transcript. Query and candidate are different media types, but the answer still needs a parent receipt ID, page alignment and a permitted-use check. A withdrawn image or text asset must be filtered before ranking, even if its embedding remains in a reproducibility snapshot. Retrieval design separates candidate supply from final ranking.
Build a hard pool
Measure recall against candidates from similar merchants, layouts and totals, not only random unrelated receipts. If the true transcript is absent from the index, the ranker cannot recover it. Report index coverage and candidate recall before top-K accuracy. Deduplicate multiple OCR versions or define which version is eligible; otherwise several near-copies can occupy the entire result list.
Version the index
Record both encoder versions, vector normalization, index build cutoff and deletion-filter version. A new text encoder requires compatible image vectors or a planned dual-index transition. Shared-space compatibility is measurable: replay fixed query pairs and compare rank and score distributions before switching traffic.
Inspect one query
For image I-47, return the correct transcript at rank two behind a same-merchant near-duplicate. The report should show candidate IDs, scores, duplicate group and why the top result was wrong. Rebuild the index after deleting that near-duplicate and confirm the deletion gate acts immediately, even before full reindexing finishes.
Implementation
def recall_at_k(ranked_candidate_ids, relevant_ids, limit):
if limit < 1:
raise ValueError("limit must be positive")
relevant = set(relevant_ids)
if not relevant:
return None
return len(set(ranked_candidate_ids[:limit]) & relevant) / len(relevant)Performance and operating cost
Top-K recall costs O(K + R) time and space for R relevant candidates. Exact scoring over C D-dimensional vectors costs O(C times D); approximate indexes trade build cost and possible recall loss for faster search.
Common Mistakes
- Do not score a ranker when the correct item was never indexed.
- Do not serve a withdrawn asset because its vector still exists.
- Do not compare embeddings from incompatible model versions.
