Annotation quality is a workflow for defining a labelable unit, applying a versioned rubric, recording uncertainty, and resolving disagreement. Meridian has 47 fictional support-ticket excerpts for a refund-versus-replacement routing task. Two reviewers label each excerpt independently. Their first-pass agreement is 40 of 47, about 85.1 percent; that is raw agreement, not proof that the labels are correct.
Project: adjudicate Meridian support labels
Review the packet
Seven excerpts disagree. Three lack earlier thread context, and four contain wording that could support refund or replacement. After context review and owner adjudication, five receive final labels and two remain unknown. The final ledger has 45 labeled excerpts and two explicit unknowns. Do not train or score on the two unknowns as if they were negatives. Lock the rubric, input scope, and holdout allocation before using the examples in a prompt or evaluation set.
47 excerpts = 40 first-pass agreements + 7 disagreements
7 disagreements = 3 missing context + 4 ambiguous wording
5 adjudicated to final label; 2 remain UNKNOWN
final: 45 labeled + 2 UNKNOWN
raw agreement: 40 / 47 = 85.1% (rounded)Performance and review cost
Two independent passes require O(2N) annotation actions for N items; adjudication adds O(D) review for D disagreements. The ledger requires O(N) storage. The extra review cost is justified where a wrong routing label changes customer treatment. Raw agreement alone can look high under an imbalanced label distribution, so inspect confusion by class and keep an unknown state.
Common Mistakes
- Do not report raw agreement as correctness.
- Do not force the two unknown excerpts into a majority class.
- Do not expose private customer text in a prompt when a redacted unit is enough.
Related lessons
- Annotation prompts: define the unit and label taxonomy
- Annotation prompts: minimize private input while preserving context
- Annotation prompts: measure agreement before adjudication
- Annotation prompts: adjudicate conflicts and version the rule
- Annotation prompts: protect holdouts and sample drift
- Annotation quality prompt decisions
- Evaluation sets: measure the failure cases that matter
- Report prompts: reconcile every claim with the packet
