Choose the action that respects the task, evidence, and effect boundary. Each answer includes a short reason.
Review the decisions
- Prompt injection: test untrusted content at every boundary
- Evaluation sets: measure the failure cases that matter
- Model judges: calibrate rubrics and swap candidate order
- Prompt releases: version the whole decision path and keep a rollback
- Prompt budgets: trade output quality against cost and tail latency
- Prompt privacy: send only the fields needed for the task
- Multilingual prompts: test policy meaning across languages
- Tool calls: validate intent and arguments before an external effect
- Conversation memory: retain decisions without retaining every private detail
- Evidence IDs: make generated claims auditable against supplied records
- Tool loops: set budgets, state checks, and a stopping condition
Common Mistakes
- Treating readable output as a verified decision.
- Allowing untrusted context to grant an action.
- Ignoring an explicit unknown state.
