Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Active-learning annotator disagreement and adjudication

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Selection often concentrates ambiguous records. Disagreement may expose unclear policy or bad evidence, not merely an annotator who needs correction.

Route hard cases intentionally

A tax-code exception can depend on contract context absent from an invoice scan. If the query strategy repeatedly picks these cases, a reviewer cannot produce a reliable binary label by guessing. Allow “insufficient evidence” and capture what document was missing. Later evidence can change the target state; that is not the same as changing the analyst’s opinion.

Separate model uncertainty from label uncertainty

Two models may disagree while trained analysts agree. Conversely, models may agree on a shortcut while analysts dispute whether the invoice is exempt. Record initial independent votes and a final adjudicated outcome. Do not train on the final result without storing the disagreement and policy version. Selection effects can amplify disputed cases.

Budget second review by impact

Double annotation of every record may exhaust the budget. Route policy-sensitive classes, contradictory evidence and novel templates for independent second review; sample some easy cases to measure background error. The code marks records for adjudication when votes differ or either vote is unknown.

Avoid majority-vote theater

Three analysts repeating an unclear rule do not resolve a broken definition. Escalate recurrent disagreements to a policy owner and version the instruction. After a policy change, revisit labels made under the old rule and keep test labels under the same target contract as training.

Measure downstream value

Compare model quality with raw first votes and adjudicated labels at the same staff-minute cost. If most queries are unanswerable, revise the acquisition policy or data collection rather than adding more labels. Stopping rules need this signal.

Implementation

python
review_votes = [
    {"invoice": "bill-47", "first": "review", "second": "review"},
    {"invoice": "bill-62", "first": "clear", "second": "review"},
    {"invoice": "bill-83", "first": "unknown", "second": "unknown"},
]

def adjudication_queue(votes):
    return [record["invoice"] for record in votes
            if record["first"] != record["second"]
            or "unknown" in (record["first"], record["second"])]

assert adjudication_queue(review_votes) == ["bill-62", "bill-83"]

Performance and operating cost

Routing N vote pairs costs O(N) time and O(N) output storage. A second review nearly doubles direct labeling work on selected records; policy adjudication may take longer than annotation. Measure total resolution time and label reversals, not just the number of completed tasks.

Common Mistakes

  • Do not force an answer when evidence is missing.
  • Do not erase disagreement after adjudication.
  • Do not blame annotators for an undefined label policy.

Read next

ai-data
machine-learning
Storage details