Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Prompt evaluation: test failures before rewriting the wording

Last updated: 2 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

A prompt revision is an experiment. Keep a fixed set of representative tasks, expected properties, and failure cases; run the old and new contracts against the same inputs. Separate a formatting error from an unsupported factual claim or an unsafe action. A single polished example is weak evidence because it may be the easiest case. Record the model setting, input version, prompt version, and reviewer decision so a result can be reproduced or explained after the tool changes.

Score the actual job

A team generates incident handoffs. Its evaluation set contains 24 routine alerts, 9 incomplete timelines, and 7 misleading log entries. The revised prompt passes only if it preserves event IDs, marks uncertain causes, stays within five bullets, and avoids external action. Run both prompt versions over all 40 cases. If the new version improves formatting but turns three unknown causes into confident claims, it is not ready. Inspect failures by class, edit the contract to address a specific failure, then rerun the full set. Do not tune solely against the cases that failed last time; that can regress ordinary cases.

Output
Evaluation batch: 40 handoff cases
Routine: 24; incomplete: 9; misleading: 7
Checks: event traceability, uncertainty, five-bullet limit, no external action
Acceptance: zero invented causes and zero unauthorized actions
Review: compare old and new prompt on the same cases
Record: prompt version, input set, model setting, reviewer decision

Cost and verification

Evaluating N cases costs O(N) model calls and O(N) review decisions, with higher cost when cases include long documents. Sample generation cannot replace human review of high-impact claims. Track both pass rate and the number of severe failures; an average can hide a dangerous regression. Refresh the evaluation set when production tasks change, but keep a stable core so prompt versions remain comparable.

Common Mistakes

  • Do not declare a prompt better after one favorable response.
  • Do not change inputs and prompt wording in the same comparison without recording both.
  • Do not let formatting scores hide invented facts or unauthorized actions.

Connected lessons

Apply this next

Field guide: Prompt failure diagnosis: find the broken contract first.

prompt engineering
AI workflows
Storage details