A prompt revision is an experiment. Keep a fixed set of representative tasks, expected properties, and failure cases; run the old and new contracts against the same inputs. Separate a formatting error from an unsupported factual claim or an unsafe action. A single polished example is weak evidence because it may be the easiest case. Record the model setting, input version, prompt version, and reviewer decision so a result can be reproduced or explained after the tool changes.
Prompt evaluation: test failures before rewriting the wording
Score the actual job
A team generates incident handoffs. Its evaluation set contains 24 routine alerts, 9 incomplete timelines, and 7 misleading log entries. The revised prompt passes only if it preserves event IDs, marks uncertain causes, stays within five bullets, and avoids external action. Run both prompt versions over all 40 cases. If the new version improves formatting but turns three unknown causes into confident claims, it is not ready. Inspect failures by class, edit the contract to address a specific failure, then rerun the full set. Do not tune solely against the cases that failed last time; that can regress ordinary cases.
Evaluation batch: 40 handoff cases
Routine: 24; incomplete: 9; misleading: 7
Checks: event traceability, uncertainty, five-bullet limit, no external action
Acceptance: zero invented causes and zero unauthorized actions
Review: compare old and new prompt on the same cases
Record: prompt version, input set, model setting, reviewer decisionCost and verification
Evaluating N cases costs O(N) model calls and O(N) review decisions, with higher cost when cases include long documents. Sample generation cannot replace human review of high-impact claims. Track both pass rate and the number of severe failures; an average can hide a dangerous regression. Refresh the evaluation set when production tasks change, but keep a stable core so prompt versions remain comparable.
Common Mistakes
- Do not declare a prompt better after one favorable response.
- Do not change inputs and prompt wording in the same comparison without recording both.
- Do not let formatting scores hide invented facts or unauthorized actions.
Connected lessons
- Prompt Engineering
- Prompt engineering: define a task that can be checked
- Prompt context: separate instructions from retrieved material
- Spring AI evaluations: separate model quality from enforced safety contracts
Apply this next
- Evaluation sets: measure the failure cases that matter
- Prompt releases: version the whole decision path and keep a rollback
- Prompt production decisions
Field guide: Prompt failure diagnosis: find the broken contract first.
