An evidence-reference check verifies that every record ID attached to a generated decision belongs to the current authorized evidence bundle. It is a structural gate, not a proof that the cited passage supports the decision. Keep that distinction explicit. A policy ID retrieved for another account or an expired version may be real yet still outside this request's permitted set. Build the set after authorization and version filtering, then validate the output before a final decision reaches a downstream system.
Code lab: reject unknown evidence IDs
Decision in practice
A claims assistant receives policy record POL-47 and receipt RC-83 for case CL-291. The first response cites both records and passes the reference gate. The second cites a plausible but absent receipt RC-99, so it is rejected even though the prose sounds confident. A separate reviewer still has to check whether POL-47 actually authorizes the requested outcome. If a record changes between retrieval and review, pin the evidence version or rerun the check against the new bundle.
allowed_evidence_ids = {"POL-47", "RC-83"}
responses = [
{"case_id": "CL-291", "decision": "approve", "evidence_ids": ["POL-47", "RC-83"]},
{"case_id": "CL-292", "decision": "approve", "evidence_ids": ["POL-47", "RC-99"]},
]
for response in responses:
cited_ids = set(response["evidence_ids"])
unknown_ids = cited_ids - allowed_evidence_ids
valid = bool(cited_ids) and not unknown_ids
print(f"{response['case_id']}: {'pass' if valid else 'reject'}")
Expected output: CL-291: pass; CL-292: reject
Performance and operating cost
Converting citations to a set and comparing them with the already-built allowed set takes O(C) expected time and O(C) extra space for C cited IDs. Building the allowed set from E records takes O(E) time and space. This check is cheap relative to retrieval or model generation, but it cannot establish semantic support. Test duplicate IDs, empty citations, stale versions, and records from a different tenant. Keep the actual evidence text and access decision available for a later claim-level review.
Common Mistakes
- Do not treat an existing ID as proof that the cited claim is true.
- Do not build the allowed set before tenant and version filtering.
- Do not allow a final decision with an empty evidence list.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Evidence IDs: make generated claims auditable against supplied records
- Retrieved evidence: reconcile versions and conflicting facts
- Code lab: score answered cases and abstentions
- Code lab: find pairwise judge order changes
- Code lab: block a critical prompt regression
- Prompt evaluation code labs
Continue with: API docs release: test links, examples, and live version.
Continue with: Retrieval prompts: pack evidence and map claims.
