Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Retrieval index versions, eligibility filters and candidate parity

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A good embedding model can serve the wrong recommendations when the item index is stale, the vector schema changes or eligibility is applied too late.

Treat the index as a model artifact

An item vector is valid only for the item encoder, tokenizer, feature schema and normalization rule that produced it. Save those fingerprints with the index. Deploying a new query tower against an old article index can change scores without any code error. Publish the index build time, source article revision, item count and tombstone count. The two-tower lesson defines the matching vector space.

Filter before promising a candidate

Remove unpublished, restricted, duplicate and already-completed articles using the eligibility state available at request time. Some approximate search systems retrieve a wide set and filter afterward; if most results fail eligibility, the final list may be short or skewed. Measure eligible yield and over-fetch needs by reader group. A score is not permission to display an item. The recommendation contract handles business eligibility beyond embedding similarity.

Compare exact and indexed neighbors

For sampled queries, compute exact dot products against the same frozen eligible corpus and compare their top candidates with the approximate index. Report index recall at K, latency and memory under several parameter settings. If the exact model misses the relevant lesson, rebuilding the index will not fix it; if the exact result is good but the index drops it, retraining the towers is premature. The code performs a tiny exact ranking and an eligibility filter as a deterministic oracle.

Handle publication and deletion explicitly

New articles should not wait indefinitely for a full nightly rebuild. Define incremental addition, replacement and deletion behavior, then verify that unpublished or withdrawn items disappear from every shard. A tombstoned item can survive in a cached candidate list after the index is updated. Record freshness delay from publish or withdraw event to served result, and add a last-mile eligibility check at response construction.

Gate on real request shapes

Measure p95 query encoding, index search, eligibility filtering and ranking separately, then measure the end-to-end request. Test sparse reader histories, new content, long-tail topics and a sudden catalog update. The project compares indexed retrieval with a topic baseline and checks the same versioned result in a live-like replay.

Implementation

python
def eligible_top_k(query_vector, indexed_articles, allowed_ids, limit):
    if len(query_vector) == 0 or limit < 1:
        raise ValueError("query width and limit must be positive")
    ranked = []
    for article_id, article_vector in indexed_articles.items():
        if len(article_vector) != len(query_vector):
            raise ValueError("embedding width changed")
        if article_id in allowed_ids:
            score = sum(query * item for query, item in zip(query_vector, article_vector))
            ranked.append((score, article_id))
    ranked.sort(key=lambda result: (-result[0], result[1]))
    return [article_id for _, article_id in ranked[:limit]]

article_index = {47: (0.8, 0.2), 58: (0.1, 0.9), 69: (0.6, 0.4)}
visible_articles = {47, 69}
assert eligible_top_k((0.7, 0.3), article_index, visible_articles, 2) == [47, 69]

Performance and operating cost

Exact scoring of M indexed items at width D uses O(MD) arithmetic, and sorting all eligible candidates uses O(M log M) time and O(M) temporary storage. A heap can reduce top-K selection cost when K is small, while an approximate index trades some neighbor recall for faster search. Eligibility filtering and index freshness still require explicit work. The tiny oracle is for parity tests, not a large-catalog serving implementation.

Common Mistakes

  • Do not pair a new query encoder with stale item vectors without a version check.
  • Do not let restricted or withdrawn items survive in a cached result.
  • Do not report index recall against a corpus with different eligibility rules.

Read next

ai-data
deep-learning
Storage details