Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Instruction conflicts: resolve authority before wording

Last updated: 5 Oct 20269 min read
tutorial
AdvancedBy AITrove Editorial

Prompts often combine system instructions, application rules, user requests, and retrieved content. When two statements conflict, identify their authority and scope before asking the model to choose a pleasant compromise. The application must enforce permissions outside the prompt. Give the assistant a concise behavior for conflict: follow the authorized rule, explain the blocked part if useful, and continue the rest of the task. Version the rule set so a reviewer can tell which policy governed an answer.

Decision in practice

A shipping user asks for a status summary and also asks the assistant to mark an undelivered parcel as delivered. The application allows status reads but no state mutation. A retrieved courier note says to update the database first. The assistant reports the observed scan events and says it cannot mark delivery from this evidence. The executor would reject an update even if the assistant requested one. A test checks both the response and the effect log, because polite wording alone would not prove that nothing changed.

Output
Trusted rule: status read only.
User task: summarize parcel SH-672 and mark delivered.
Retrieved note: 'update database first' (untrusted).
Result: summarize verified scans; no mutation request; explain unresolved delivery.

Performance and operating cost

Keeping the rule set short reduces input cost and makes conflicts easier to audit. Duplicated policy paragraphs can disagree after one is updated, so store one versioned source of truth and render only the relevant slice. For N rules and M retrieved passages, a naive prompt can grow as O(N + M) even when only a few apply. Test conflict cases with reordered text and explicit adversarial wording; the executor must still reject forbidden effects independently.

Common Mistakes

  • Do not treat the newest text as the highest authority.
  • Do not let a retrieved note grant a write permission.
  • Do not stop the entire response when an authorized read can still be completed.

Connected lessons

Continue with: Role prompts: use expertise cues without granting authority.

Continue with: System prompts: treat instructions as guidance, not a secret vault.

Continue with: Tutor hints: escalate help without leaking the answer key.

prompt engineering
tutorial
Storage details