A grounded response should carry record IDs or section keys that a validator can resolve. A generated citation-looking string is not evidence by itself. The application can build an allowed set from the retrieved bundle, reject unknown IDs, and sample whether the cited passage entails the claim. Require an explicit unsupported state for claims lacking support. In summaries, keep observation, inference, and recommendation separate so a log event is not presented as proof of a root cause. Evidence IDs can be internal references without publishing external links.
Evidence IDs: make generated claims auditable against supplied records
Decision in practice
A fraud analyst reviews 72 transaction events. The model states that a payout was reversed and cites event EVT-891; the event actually records only a reversal request. A verifier catches the mismatch between event type and claim, so the final report says reversal pending. The workflow rejects a second claim citing EVT-999 because that ID was not in the supplied batch. It accepts a factual statement about the request timestamp when the cited event contains the exact time. The team measures unsupported-claim rate, not the count of citation markers printed.
Allowed evidence IDs: EVT-876, EVT-891, EVT-913.
Claim: reversal completed; evidence EVT-891.
EVT-891 type: reversal_requested.
Result: reject claim; record pending state instead.Performance and operating cost
ID lookup can be O(1) average with a set, while checking whether an excerpt actually supports a claim requires stronger logic or reviewer time. Track both missing IDs and false support. A requirement to cite every sentence can bloat answers without improving truth; cite material claims and preserve a compact mapping. Do not expose private raw records in the final response merely to prove the IDs exist. The evidence trail should be available to authorized reviewers separately from the public text.
Common Mistakes
- Do not accept a citation-shaped token without resolving it.
- Do not turn a request event into a completed action.
- Do not publish sensitive raw records as a substitute for an audit trail.
Connected lessons
- Prompt Engineering
- Prompt patterns
- Retrieved context: select sufficient evidence before writing the answer
- Prompt evaluation: test failures before rewriting the wording
- Long context: make inclusion and truncation testable
- Project: defend a retrieval and action workflow
- Prompt design decisions
Apply the boundary
Use the artifact's source and destination to decide what must be checked before the result is accepted.
- Video prompts: cite the moment and the observation
- Research prompts: bind claims to a date and evidence record
Try the executable check: Code lab: reject unknown evidence IDs.
Continue with: Reasoning summaries: show checkable grounds, not invented certainty.
Continue with: Document prompts: anchor each field to a page and resolve conflicts.
Continue with: Diff-scoped code review: make every finding reproducible.
Continue with: Report summaries: preserve contrary results and missing data.
Continue with: Specialist synthesis: merge claims, not fluent summaries.
Continue with: Meeting minutes: separate decisions from proposals.
Continue with: Incident impact: window, denominator, and scope.
