Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Code lab: find pairwise judge order changes

Last updated: 2 Oct 20269 min read
tutorial
AdvancedBy AITrove Editorial

A pairwise judge order audit compares a judge's preference for two fixed candidate responses after their presentation order is swapped. Record verdicts using stable candidate IDs, not first or second position. If the normalized winner changes, presentation order may be influencing the judgment. A changed verdict is a signal for review, not a calibrated estimate of bias from a few examples. Human adjudication should examine the rubric and the original responses.

Decision in practice

A support team compares prompt candidate A with candidate B on three tickets. For ticket TK-481, A wins in either presentation order. For TK-482, the winner changes from B to A after the swap. The third case ties twice. The team sends TK-482 to an independent reviewer and checks whether the rubric rewards the right behavior. They keep both judge calls and the response text; a single retained score would hide the order effect.

python
pairwise_reviews = [
    {"ticket_id": "TK-481", "winner_ab": "A", "winner_ba": "A"},
    {"ticket_id": "TK-482", "winner_ab": "B", "winner_ba": "A"},
    {"ticket_id": "TK-483", "winner_ab": "tie", "winner_ba": "tie"},
]

changed = [
    review["ticket_id"]
    for review in pairwise_reviews
    if review["winner_ab"] != review["winner_ba"]
]
print("changed:", ", ".join(changed) if changed else "none")

Expected output: changed: TK-482

Performance and operating cost

The comparison is O(N) time for N paired cases and O(F) space for F flagged case IDs. The expensive work is obtaining two judge calls per case and adjudicating disagreements. If a judge is stochastic, repeat each order before drawing a strong conclusion, then compare disagreement patterns by task slice. Pin the rubric and candidate responses during the audit. A tie in one order and a preference in the other is also a changed verdict worth review.

Common Mistakes

  • Do not record the winning screen position as the candidate identity.
  • Do not discard the reversed-order judgment after averaging scores.
  • Do not treat one changed verdict as a population-level bias estimate.

Connected lessons

prompt engineering
evaluation
python
Storage details