Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Evaluation labels: adjudicate disagreement before scoring a release

Last updated: 4 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

A reference label is a decision claim made by a reviewer under a particular policy version. When reviewers disagree, the case may expose unclear instructions, incomplete evidence, or a policy boundary rather than a model defect. Have independent reviewers label without seeing the candidate response where practical. Record their labels and reasons, adjudicate consequential disagreements, and preserve the original votes. A model judge can assist with triage, but it should be checked against human-labeled cases before it becomes a release gate. The adjudicated outcome must be tied to the policy and evidence that produced it.

Decision in practice

Two reviewers label ticket TK-862 differently: one says billing dispute, the other says service incident because the outage caused the duplicate charge. The case packet lacks the incident start time. The board marks the reference as unresolved, requests the event record, and leaves this case outside the binary score until a policy owner decides the precedence rule. It still tracks the case as an ambiguity signal. After the record arrives, both reviewers see the same evidence and the policy owner records the final label with a reason. The prompt is not tuned to whichever answer the current model produced.

Output
Case TK-862; policy LP-12; evidence packet EP-8.
Reviewer A: billing; reviewer B: incident.
Disagreement reason: outage time missing.
State: unresolved; excluded from pass/fail denominator.
Resolution: append event record and policy-owner decision.
Audit: preserve both votes and the final rationale.

Performance and operating cost

Two independent labels cost roughly twice the first-pass human effort, and adjudication adds work to a subset of cases. If N cases each receive K labels, collection requires O(NK) judgments; matching and disagreement counts are O(NK). Reserve extra review for high-impact or unstable slices rather than treating every typo as a policy dispute. An unresolved reference cannot support a precise accuracy claim. Removing it silently changes the denominator, so report both scored and unresolved counts. The label ledger also helps when policy changes require older cases to be reclassified.

Common Mistakes

  • Do not force a label when required evidence is absent.
  • Do not let a candidate answer determine its own reference label.
  • Do not erase original reviewer votes after adjudication.

Connected lessons

Continue with: Synthetic case oracles: reject invalid labels before scoring a model.

Continue with: Search prompts: judge relevance without treating unknown as wrong.

Continue with: Annotation prompts: adjudicate conflicts and version the rule.

prompt engineering
evaluation
Storage details