Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Embedding retrieval recall and collapse checks

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Evaluate an embedding against a frozen gallery and unseen queries, measuring correct-neighbor retrieval and whether the representation keeps useful spread.

Choose the query and gallery roles

A held-out seal image is the query. The gallery contains eligible, versioned reference photos available when the query arrives. Define success as at least one correct defect-family reference in the first K results, not merely a close vector. Exclude the same physical component from the gallery when the product must generalize to new components. Identity splits make that requirement testable.

Report both hit rate and false match cost

Recall at K counts queries with a relevant neighbor in the first K. For an operational decision, also count high-scoring wrong repairs, no-match abstentions and performance by camera, defect rarity and part type. A model may raise recall by returning many uncertain neighbors while increasing manual review. Selective prediction frames the action threshold.

Look for collapse

If vectors converge to nearly the same location, triplet selection and retrieval become uninformative. Track nonduplicate pair distances, per-dimension spread and the fraction of query vectors with near-tied nearest neighbors. Low variance alone is not conclusive for every architecture; inspect it alongside recall and a random-encoder baseline.

Keep exact search as an oracle

Before selecting an approximate index, brute-force distances over a manageable frozen gallery provide a reference neighbor list. The encoder can still be wrong; exact search only answers whether the index reproduces the encoder’s choices. The index audit compares those lists and full-path latency.

Preserve version compatibility

Query and gallery embeddings must use the same encoder and preprocessing. A rolling model deployment can mix versions if index swap and application rollout are not coordinated. Stamp each vector, query and result with encoder and gallery versions; reject incompatible combinations or route to a matching index. Index swap lifecycle covers the adjacent data path.

Implementation

python
from math import dist

gallery = {
    "cut-reference": {"vector": (0.12, 0.08), "defect": "edge-cut", "component": "seal-83"},
    "pit-reference": {"vector": (0.86, 0.81), "defect": "surface-pit", "component": "seal-91"},
    "wear-reference": {"vector": (0.54, 0.31), "defect": "abrasion", "component": "seal-62"},
}
queries = [
    {"vector": (0.15, 0.10), "defect": "edge-cut", "component": "seal-47"},
    {"vector": (0.82, 0.78), "defect": "surface-pit", "component": "seal-24"},
]

def exact_recall_at_one(test_queries, references):
    hits = 0
    for query in test_queries:
        eligible = [(name, item) for name, item in references.items()
                    if item["component"] != query["component"]]
        if not eligible:
            raise ValueError("empty eligible gallery")
        winner = min(eligible, key=lambda pair: dist(query["vector"], pair[1]["vector"]))
        hits += winner[1]["defect"] == query["defect"]
    return hits / len(test_queries)

assert exact_recall_at_one(queries, gallery) == 1.0

Performance and operating cost

Exact retrieval costs O(Q times G times D) for Q queries, G gallery vectors and D dimensions; storage is O(G times D). An approximate index can reduce query latency but introduces its own build, memory and neighbor-recall costs. The tiny example proves only the arithmetic path, not model quality on real defect images.

Common Mistakes

  • Do not leave duplicates of the query component in an unseen-component gallery.
  • Do not use index agreement with exact search as the only quality metric.
  • Do not mix encoder versions between query and gallery vectors.

Read next

ai-data
machine-learning
Storage details