A retrieved passage may contain instructions aimed at the model; its text must not gain authority to change tool use or access policy.
Retrieved text as untrusted data: keep instructions and tools separate
Name the trust levels
The user request and application instructions have different roles from handbook passages. A document can contain an imperative such as “ignore prior rules” as ordinary data, or as an attack planted by someone who edits the corpus. Do not promote that text into system instructions. Permission checks] must happen before a passage reaches any model or tool.
Restrict actions outside language
A prompt can ask a model to ignore a passage, but prompt wording alone cannot enforce authorization or tool limits. Give the answer component only read capabilities needed for its task. Validate tool arguments and destination allowlists in application code. A retrieved passage should never decide which user, tenant or document scope the server searches.
Preserve suspicious evidence safely
Log passage ID and a bounded diagnostic reason when hostile instructions are detected; avoid echoing private text into broad logs. Keep an incident fixture with a forged policy update and a request to expose another department’s handbook. The expected result is no unauthorized retrieval or action, regardless of answer fluency.
Test in stages
Run permission tests with retrieval alone, then answer tests with a hostile passage inserted, then tool-call tests under a constrained stub. A system that simply refuses every question is not useful; measure legitimate answer completion alongside attack resistance. Layered evaluation] prevents one success measure from hiding another failure.
Implementation
def validate_retrieval_request(principal, requested_scope, allowed_scopes):
if requested_scope not in allowed_scopes.get(principal, set()):
raise PermissionError("retrieval scope is not authorized")
return {"principal": principal, "scope": requested_scope}
def validate_tool_action(action_name, allowed_actions):
if action_name not in allowed_actions:
raise PermissionError("answer component cannot perform that action")Performance and operating cost
Authorization lookup is usually O(1) to O(log P) for P policy entries; comprehensive attack testing costs curated fixtures and review time. External tool side effects have far higher cost than a few checks.
Common Mistakes
- Do not treat retrieved prose as trusted instructions.
- Do not rely on a prompt to enforce authorization.
- Do not grant write tools to a read-only answer feature.
