Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Annotation prompts: measure agreement before adjudication

Last updated: 7 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

Double labeling gives two independent judgments for the same item under the same rubric version. Hide the first label from the second reviewer to avoid anchoring. Compute raw agreement as matching first-pass labels divided by double-labeled items, then inspect which label pairs disagree. Agreement measures consistency, not truth. Imbalanced categories can inflate it; chance-adjusted measures need their own assumptions and a complete contingency table. The prompt can format the ledger but deterministic code should count matches.

Operational case

Meridian's two reviewers agree on 40 of 47 excerpts and disagree on seven. Raw agreement is 40/47, approximately 85.1 percent. That number does not show whether both reviewers misunderstood refund policy. The seven disagreements are listed by item ID and label pair. Three need preceding context and four need a clearer rule for language that could request either refund or replacement. No reviewer sees the other's label before the first pass closes.

Output
double-labeled items: 47
first-pass agreements: 40
disagreements: 7
raw agreement: 40 / 47 = 0.851063...
quality verdict: requires adjudication and class review

Performance and review cost

Comparing N paired labels is O(N) time and O(N) ledger storage; class-pair counts need O(C squared) counters for C labels. Independent passes roughly double annotation labor. That cost exposes a weak rubric before its labels become a hidden dependency in evaluation.

Common Mistakes

  • Do not call 85.1 percent accuracy.
  • Do not let the second reviewer see the first label.
  • Do not hide class-specific disagreement behind one average.

Connected lessons

prompt engineering
annotation quality
Storage details