A positive query-passage pair should carry evidence for the requested answer at a known source revision, not merely share terms with the query.
Retrieval positives: query intent, passage identity and revision
Define the relevance judgment
A query about gateway-west rollback time may retrieve a page that mentions the service but discusses alert thresholds. That page is topically related and still not an answer-bearing positive. Record the query ID, intent, service, environment, passage ID, source revision, judged answer span and reviewer decision. A passage can be relevant for one query and irrelevant for another. Hybrid ranking can find candidates, but it cannot supply gold relevance labels by itself.
Keep source identity stable
Store the canonical passage ID separately from its current text and revision. A corrected runbook may keep the same section path while changing a numeric limit; the old positive judgment must not be reused blindly. Preserve prior judgments for audit and mark them stale until a reviewer checks the new revision. Passage revision auditing carries the source change into the index; pair labels must follow the same change.
Distinguish passage support from answer correctness
A relevant passage may require a second passage to answer a compound question. Label direct support, partial support and non-support separately. Attach exact evidence spans and policy context. Do not mark a passage positive because a generated answer sounds plausible; judge the passage before seeing the generated answer. Claim evidence contracts require the final answer to point back to support that actually exists.
Audit the training set
Sample positives by query type, source language, service and revision. Check whether protected or restricted passages entered the training export. Group repeated queries and revised passages into families before splitting. Measure disagreement on borderline judgments and how often a passage is labeled positive for one revision but not another. The negative review uses these positive labels to avoid training against a passage that also answers the question.
Implementation
def reviewed_positive_pair(query, passage, evidence_start, evidence_end):
if query["service_id"] != passage["service_id"]:
raise ValueError("query and passage service differ")
if not passage["revision_id"] or not passage["passage_id"]:
raise ValueError("source identity is required")
if not (0 <= evidence_start < evidence_end <= len(passage["text"])):
raise ValueError("invalid evidence span")
return {"query_id": query["query_id"],
"passage_id": passage["passage_id"],
"revision_id": passage["revision_id"],
"evidence": passage["text"][evidence_start:evidence_end]}
query = {"query_id": "rollback-47", "service_id": "gateway-west"}
passage = {"passage_id": "runbook-82", "revision_id": "r8",
"service_id": "gateway-west", "text": "Wait 82 minutes before rollback."}
pair = reviewed_positive_pair(query, passage, 5, 15)
assert pair["evidence"] == "82 minutes"
assert pair["revision_id"] == "r8"
Performance and operating cost
Bounds and identity checks are O(1) for fixed-size metadata; copying an evidence span of s characters costs O(s) time and space. Human relevance review and passage retrieval dominate real collection cost. This function validates a reviewed span but cannot determine whether that span actually answers the query.
Common Mistakes
- Calling a passage positive because it contains the same service name.
- Reusing an old relevance label after a threshold changed.
- Taking a model-generated answer as proof of passage support.
- Training on restricted passage text without access review.
