Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Adjudicate text labels before they become training truth

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Reviewer disagreement can expose a vague taxonomy, a missing context window or a genuinely ambiguous message. Keep those causes distinct.

Preserve independent judgments

For a sample of messages, collect at least two independent labels with policy version and reviewer ID. Do not show one reviewer the other’s answer before the initial decision. Compare intent, span boundary and relation labels separately. A high agreement score on the dominant “other” label can hide poor agreement on a rare critical category. Store disagreement reasons and the exact original text revision.

Fix policy before relabeling

If reviewers disagree about whether “refund pending” belongs to billing or returns, revise the taxonomy with explicit inclusion and exclusion cases. If they disagree about a code-switched phrase, provide context or mark it unresolved. Do not force a majority vote when the policy lacks a defensible answer. Record the adjudicator’s decision, rationale and effective policy version. Corpus contracts define the upstream identity of each label.

Measure label debt

Track unresolved count, disagreement rate by category and source channel, turnaround time and the fraction of training rows created under retired policies. A model trained across incompatible taxonomies can appear to have noisy examples when the real problem is mixed definitions. Before retraining, migrate or exclude old labels intentionally. Keep the final evaluation labels reviewed under one frozen policy, otherwise model comparisons are not meaningful.

Close the acquisition loop

Use disagreements to refine guidelines and choose targeted review, but retain a random sample to estimate true traffic. After a policy update, recheck a fixed set of hard examples and measure whether agreement improved. The sampling lesson chooses cases; the acquisition project keeps their provenance and adjudication state.

Implementation

python
def adjudication_state(first_label, second_label, policy_version):
    if not policy_version:
        raise ValueError("label policy version is required")
    if first_label is None or second_label is None:
        return "awaiting-independent-review"
    if first_label == second_label:
        return "agreed"
    return "needs-adjudication"

assert adjudication_state("refund", "billing", "policy-v7") == "needs-adjudication"
assert adjudication_state("refund", "refund", "policy-v7") == "agreed"

Performance and operating cost

This state check is O(1) time and space. Double-labeling every example roughly doubles first-pass human effort, so use a planned overlap sample and targeted high-risk categories. Adjudication cost rises when taxonomy language is vague. Measure review minutes per accepted label and downstream error reduction rather than optimizing agreement by collapsing useful classes.

Common Mistakes

  • Calling majority vote ground truth when policy is ambiguous.
  • Mixing labels from incompatible policy versions.
  • Reporting overall agreement while rare critical labels disagree.
  • Showing reviewers the model suggestion without measuring anchoring effects.

Read next

Continue the workflow: Dialect-aware NLP audits: slices and label disagreement.

ai-data
natural-language-processing
Storage details