A paired prompt evaluation sends each case through a baseline and a candidate under the same evidence, model settings, and scoring rule. The unit of comparison is the case, not the average of two unrelated batches. Record wins, losses, ties, and critical regressions separately. A small positive net count can be noise, especially when a few cases carry most of the apparent gain. Repeated model calls may reveal unstable cases; they do not repair a flawed label or make a critical failure acceptable. State the decision threshold before reading the result.
Paired prompt evaluation: count changes, then inspect uncertainty
Decision in practice
A claims assistant gets a shorter prompt. On 47 fixed claims, the candidate improves 14, worsens 9, and leaves 24 unchanged against the previous version. Net improvement is five cases, but one loss approves a claim with a missing receipt. The release is blocked by that predeclared critical rule. The team keeps per-case outcomes and the exact case-set revision, checks whether the receipt was missing in both runs, and sends the contested label to review. It does not report only a 14-to-9 headline or swap out the failing case.
Case set EV-47; same evidence and model settings for both versions.
Candidate better: 14; baseline better: 9; tie: 24.
Critical regression: CL-431, missing receipt accepted.
Gate: block any critical false approval; review disputed label.
Next run: retain all 47 cases and record the new candidate revision.Performance and operating cost
For N cases, one paired pass uses 2N model calls. Repeating each side R times uses 2NR calls plus grading and review; the calls usually dominate latency and spend. The result table occupies O(N) records. A confidence interval or resampling calculation can quantify sampling uncertainty, but it cannot fix correlated cases, biased labels, or a changed evidence snapshot. Stratify by the decision that matters, and budget reviewer time for changes near the release threshold. A larger sample is useful only when its cases represent the production failure modes.
Common Mistakes
- Do not compare candidates on different case sets and call the result paired.
- Do not bury a critical false approval in an aggregate win rate.
- Do not treat repeated sampling as a substitute for a valid expected label.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Evaluation sets: measure the failure cases that matter
- Code lab: block a critical prompt regression
- Code lab: find pairwise judge order changes
- Evaluation labels: adjudicate disagreement before scoring a release
- Adversarial case mutations: test the boundary, not a magic phrase
- Online prompt experiments: define exposure and stop rules first
- Evaluation case ledgers: revise labels without erasing history
- Project: govern a claims-assistant evaluation board
- Prompt evaluation governance decisions
