A good embedding model can serve the wrong recommendations when the item index is stale, the vector schema changes or eligibility is applied too late.
Retrieval index versions, eligibility filters and candidate parity
Treat the index as a model artifact
An item vector is valid only for the item encoder, tokenizer, feature schema and normalization rule that produced it. Save those fingerprints with the index. Deploying a new query tower against an old article index can change scores without any code error. Publish the index build time, source article revision, item count and tombstone count. The two-tower lesson defines the matching vector space.
Filter before promising a candidate
Remove unpublished, restricted, duplicate and already-completed articles using the eligibility state available at request time. Some approximate search systems retrieve a wide set and filter afterward; if most results fail eligibility, the final list may be short or skewed. Measure eligible yield and over-fetch needs by reader group. A score is not permission to display an item. The recommendation contract handles business eligibility beyond embedding similarity.
Compare exact and indexed neighbors
For sampled queries, compute exact dot products against the same frozen eligible corpus and compare their top candidates with the approximate index. Report index recall at K, latency and memory under several parameter settings. If the exact model misses the relevant lesson, rebuilding the index will not fix it; if the exact result is good but the index drops it, retraining the towers is premature. The code performs a tiny exact ranking and an eligibility filter as a deterministic oracle.
Handle publication and deletion explicitly
New articles should not wait indefinitely for a full nightly rebuild. Define incremental addition, replacement and deletion behavior, then verify that unpublished or withdrawn items disappear from every shard. A tombstoned item can survive in a cached candidate list after the index is updated. Record freshness delay from publish or withdraw event to served result, and add a last-mile eligibility check at response construction.
Gate on real request shapes
Measure p95 query encoding, index search, eligibility filtering and ranking separately, then measure the end-to-end request. Test sparse reader histories, new content, long-tail topics and a sudden catalog update. The project compares indexed retrieval with a topic baseline and checks the same versioned result in a live-like replay.
Implementation
def eligible_top_k(query_vector, indexed_articles, allowed_ids, limit):
if len(query_vector) == 0 or limit < 1:
raise ValueError("query width and limit must be positive")
ranked = []
for article_id, article_vector in indexed_articles.items():
if len(article_vector) != len(query_vector):
raise ValueError("embedding width changed")
if article_id in allowed_ids:
score = sum(query * item for query, item in zip(query_vector, article_vector))
ranked.append((score, article_id))
ranked.sort(key=lambda result: (-result[0], result[1]))
return [article_id for _, article_id in ranked[:limit]]
article_index = {47: (0.8, 0.2), 58: (0.1, 0.9), 69: (0.6, 0.4)}
visible_articles = {47, 69}
assert eligible_top_k((0.7, 0.3), article_index, visible_articles, 2) == [47, 69]Performance and operating cost
Exact scoring of M indexed items at width D uses O(MD) arithmetic, and sorting all eligible candidates uses O(M log M) time and O(M) temporary storage. A heap can reduce top-K selection cost when K is small, while an approximate index trades some neighbor recall for faster search. Eligibility filtering and index freshness still require explicit work. The tiny oracle is for parity tests, not a large-catalog serving implementation.
Common Mistakes
- Do not pair a new query encoder with stale item vectors without a version check.
- Do not let restricted or withdrawn items survive in a cached result.
- Do not report index recall against a corpus with different eligibility rules.
Read next
- Two-tower embeddings and in-batch negative contracts
- Project: retrieve learning articles with two neural towers
- Candidate retrieval: separate broad discovery from hard eligibility
- Offline recommendation evaluation: replay catalog and learner state in time
- Inference contracts: preserve preprocessing and measure tail latency
