A useful evaluation set groups cases by decision type, evidence quality, language, document size, and consequence of error. Keep a stable holdout that prompt authors do not tune against directly; use a separate development set for iteration. Define the grading rule before inspecting model output. Binary pass rules work for schema, prohibited effects, and evidence references, while editorial quality needs a rubric and human review. Track both aggregate performance and the worst important slice. A prompt that improves common cases while worsening abstention may be a regression.
Evaluation sets: measure the failure cases that matter
Decision in practice
A claims team samples 73 records: clean approvals, expired coverage, missing receipts, contradictory dates, and scanned forms with weak OCR. It labels each record with the accepted decision and acceptable uncertainty. When a new prompt scores better overall but turns two missing-receipt cases into approvals, the release gate fails. The team keeps those cases in the holdout and adds a fresh development case for the new failure pattern. Evaluators see record IDs and policy versions, so a later policy update can be separated from prompt drift.
Slices: clean, missing evidence, conflict, OCR uncertainty, unauthorized action.
Gate: zero unauthorized effects; no regression on missing-evidence decisions.
Record: prompt version, model version, case ID, verdict, reason.Performance and operating cost
Evaluation cost grows with case count and repeated model runs. If N cases are each run R times, model calls are O(NR); judge calls add another multiplier. Sample where uncertainty is high, but retain a fixed core for comparison. Report confidence intervals or raw counts when slices are small, and avoid claiming that a change from one error to zero proves general reliability. A human adjudication queue should capture disagreements rather than silently changing expected labels.
Common Mistakes
- Do not tune directly on the final holdout.
- Do not let a high average conceal a critical slice failure.
- Do not score a changed policy as if only the prompt changed.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Prompt evaluation: test failures before rewriting the wording
- Model judges: calibrate rubrics and swap candidate order
- Prompt injection: test untrusted content at every boundary
- Project: defend a retrieval and action workflow
- Prompt production decisions
Further prompt decisions
Test the task on changed inputs and record what the application actually accepted or rejected.
- Model settings: change one generation variable against a fixed case set
- Metamorphic tests: verify behavior when harmless details change
- Fairness checks: test equivalent cases across groups and wording
Next decision
Check whether the right facts reached the workflow and whether the result is safe at its destination.
- Evaluation leakage: keep the release test independent
- Project: measure a retrieval-backed answer gate
Continue with: Negative controls: test the answer that should not be produced.
Related implementation
Continue with: Paired prompt evaluation: count changes, then inspect uncertainty.
Field guide: Prompt failure diagnosis: find the broken contract first.
Continue with: Project: select and verify a reasoning pattern.
Continue with: Project: verify a shipment performance report.
Continue with: Cross-locale evaluation: compare decisions, not word-for-word text.
Continue with: Accessible output evaluations: measure failures by artifact and user task.
Continue with: Real-time voice evaluations: test timing and state, not transcript fluency.
Continue with: Synthetic case coverage: count distinct decisions, not rewritten sentences.
Continue with: Forecast prompts: compare against a rolling baseline.
Continue with: Security alert prompts: test benign explanations before escalation.
Continue with: Survey prompts: pilot the instrument before fielding.
Continue with: Accessibility prompts: test keyboard and dialog focus paths.
Continue with: Experiment prompts: freeze the primary metric and guardrails.
