A new index must pass relevance, coverage and latency gates before an atomic pointer sends production queries to it.
Retrieval release gates: shadow queries and reversible index cutover
Keep candidate and serving results distinct
Send a recorded query sample to both old and new indexes while only the old result reaches users. Keep the query text, client context, eligible-document filters and timestamp fixed for each paired comparison. Do not reuse old embeddings against the new index. Measure result IDs and grades, not raw similarity scores: scores from different embedding spaces are not directly comparable. The index contract binds each query to its own encoder.
Choose a useful relevance frame
Curate repair questions with judged relevant bulletin IDs, including rare machine variants, newly issued instructions and withdrawn bulletins. Measure recall at the displayed depth, filter correctness, zero-result rate and p95 retrieval latency by cohort. An improvement in average recall cannot excuse returning an obsolete safety step. Include source-revision lag and missing-document counts as independent gates. Slice gates give the release a way to reject a harmful minority regression.
Move one pointer after evidence passes
Once the candidate is complete and the update replay has reached the agreed watermark, switch a serving alias or versioned pointer atomically. Keep the old collection available through a soak window, and pin each request to the collection revision it used. Watch relevance proxies, empty results, latency and failures after cutover. If a gate fails, restore the old alias while preserving the new collection for diagnosis. A reversible pointer makes the rollback explicit.
Avoid a false sense of safety
Shadow traffic is not a random sample of all future questions. Protected or infrequent filters may have no shadow coverage, so the judged set must include them. A live user query may reference a bulletin created after the frozen evaluation set; freshness checks cover this gap. Record which cohorts remain untested and use a limited canary before full traffic. The project exercises a stale delete and a relevance regression that an average metric hides.
Implementation
def retrieval_release_gate(report, limits):
if report["missing_eligible"] or report["stale_deletes"]:
return "hold:corpus-integrity"
if report["rare_variant_recall"] < limits["rare_variant_recall"]:
return "hold:rare-variant"
if report["p95_ms"] > limits["p95_ms"]:
return "hold:latency"
if report["source_lag"] > limits["source_lag"]:
return "hold:freshness"
return "cutover:alias"
limits = {"rare_variant_recall": 0.82, "p95_ms": 47, "source_lag": 2}
report = {"missing_eligible": 0, "stale_deletes": 0,
"rare_variant_recall": 0.87, "p95_ms": 39, "source_lag": 1}
assert retrieval_release_gate(report, limits) == "cutover:alias"
assert retrieval_release_gate({**report, "stale_deletes": 1}, limits) == "hold:corpus-integrity"
Performance and operating cost
The gate is O(1) time and space over an aggregated report. Shadowing adds one candidate query per sampled serving query, so read load grows with sample rate. Keeping the old index through the soak window adds storage, but permits a pointer rollback without a second rebuild.
Common Mistakes
- Comparing similarity scores across unrelated embedding spaces.
- Treating a high average recall as proof that rare filters are safe.
- Deleting the old index before the soak window ends.
- Cutting over with unresolved source lag or stale tombstones.
Read next
- Embedding index migration: bind vectors, queries and source revisions
- Project: migrate a repair-manual retrieval index without losing updates
- Slice quality gates when labels are sparse or delayed
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Generative evaluation gates: grounded claims, schemas and abstention
