A pairwise judge order audit compares a judge's preference for two fixed candidate responses after their presentation order is swapped. Record verdicts using stable candidate IDs, not first or second position. If the normalized winner changes, presentation order may be influencing the judgment. A changed verdict is a signal for review, not a calibrated estimate of bias from a few examples. Human adjudication should examine the rubric and the original responses.
Code lab: find pairwise judge order changes
Decision in practice
A support team compares prompt candidate A with candidate B on three tickets. For ticket TK-481, A wins in either presentation order. For TK-482, the winner changes from B to A after the swap. The third case ties twice. The team sends TK-482 to an independent reviewer and checks whether the rubric rewards the right behavior. They keep both judge calls and the response text; a single retained score would hide the order effect.
pairwise_reviews = [
{"ticket_id": "TK-481", "winner_ab": "A", "winner_ba": "A"},
{"ticket_id": "TK-482", "winner_ab": "B", "winner_ba": "A"},
{"ticket_id": "TK-483", "winner_ab": "tie", "winner_ba": "tie"},
]
changed = [
review["ticket_id"]
for review in pairwise_reviews
if review["winner_ab"] != review["winner_ba"]
]
print("changed:", ", ".join(changed) if changed else "none")
Expected output: changed: TK-482
Performance and operating cost
The comparison is O(N) time for N paired cases and O(F) space for F flagged case IDs. The expensive work is obtaining two judge calls per case and adjudicating disagreements. If a judge is stochastic, repeat each order before drawing a strong conclusion, then compare disagreement patterns by task slice. Pin the rubric and candidate responses during the audit. A tie in one order and a preference in the other is also a changed verdict worth review.
Common Mistakes
- Do not record the winning screen position as the candidate identity.
- Do not discard the reversed-order judgment after averaging scores.
- Do not treat one changed verdict as a population-level bias estimate.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Model judges: calibrate rubrics and swap candidate order
- Metamorphic tests: verify behavior when harmless details change
- Code lab: reject unknown evidence IDs
- Code lab: score answered cases and abstentions
- Code lab: block a critical prompt regression
- Prompt evaluation code labs
