A reranker orders retrieved candidates; its useful ceiling is set by upstream recall, eligibility filters and the latency budget of the entire request.
Ranking serving budget and candidate recall
Measure the retrieval ceiling
For each independently judged maintenance query, ask whether at least one grade-2-or-better page entered the candidate pool. If no relevant page is present, no downstream ordering model can place one first. Report candidate recall by equipment family, document age and permission class, with the size of each retrieved set. Candidate retrieval is a separate system decision.
Keep permissions in the path
Some documents may be relevant but not allowed for a given crew. Eligibility checks belong before ranking and again at final rendering if permissions can change. A ranker must not learn from fields a viewer cannot access. Corpus permissions describe the lifecycle requirement; a model score never grants access.
Budget full-path latency
Measure query parsing, candidate retrieval, feature lookup, model scoring, permission recheck and render time on the serving device. A 24-millisecond scorer can still make a 170-millisecond request miss a 150-millisecond target. Record p50, p95, timeout and cold-cache behavior at realistic concurrency. Device profiling provides the measurement discipline.
Bound feature cost
Precompute document-only features when their revision changes. Compute query-document interaction features only for the eligible top candidates, and use a declared fallback when a feature service times out. Do not insert a missing value that carries hidden knowledge of label availability. Validate feature freshness and queueing when deciding whether a larger model is worth deploying.
Design fallback before rollout
If the ranker fails, use a known safe heuristic that still enforces permissions and freshness. Log which path served each request and compare task outcomes separately. The fallback has its own quality and latency profile; it cannot be dismissed because a healthy-model benchmark passes. The project makes this a promotion gate.
Implementation
judged_requests = [
{"request": "pump-47", "eligible_relevant": {"seal-a7"}, "retrieved": ["seal-a7", "parts-b4"]},
{"request": "valve-62", "eligible_relevant": {"isolation-c9"}, "retrieved": ["parts-b4", "check-c2"]},
{"request": "motor-83", "eligible_relevant": {"bearing-d5"}, "retrieved": ["bearing-d5", "motor-e2"]},
]
def candidate_recall(requests):
with_positive = [request for request in requests if request["eligible_relevant"]]
if not with_positive:
raise ValueError("no judged positive requests")
found = sum(bool(request["eligible_relevant"] & set(request["retrieved"]))
for request in with_positive)
return found, len(with_positive), found / len(with_positive)
found, eligible, recall = candidate_recall(judged_requests)
assert (found, eligible) == (2, 3)
assert round(recall, 3) == 0.667
stage_latency_ms = {"retrieval": 48, "features": 31, "scoring": 24, "render": 57}
assert sum(stage_latency_ms.values()) == 160 # Above a 150 ms request budget.Performance and operating cost
Candidate recall calculation is O(total retrieved IDs) expected time with temporary sets. Scoring cost grows with candidate count and feature work, while end-to-end latency includes retrieval and rendering. A timeout fallback, permission recheck and cold-cache measurement can each add cost absent from a model-only benchmark.
Common Mistakes
- Do not claim a reranker fixes missing candidates.
- Do not use model scores to bypass document eligibility or permissions.
- Do not report scorer latency as if it were the user-visible request latency.
