Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Reviewer agreement: inspect confusion cells before one kappa number

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Agreement between two reviewers depends on observed matches, their label marginals and the prevalence of each label.

Audit the labeling setup

Two claims reviewers independently classify the same sampled cases as duplicate or legitimate. Independence means neither sees the other’s answer before submitting; it does not mean their judgments are statistically independent. Define the label policy, ambiguous-case route, random audit sample and allowed evidence. An agreement rate from handpicked easy claims does not represent the queue. Diagnostic estimates inherit any error in the reference labels.

Read the complete two-by-two table

Record cases both call duplicate, both call legitimate, and the two disagreement directions. Observed agreement is the share on the matching diagonal. Cohen’s kappa subtracts an expected agreement derived from each reviewer’s label proportions and rescales the remainder. When true duplicates are rare, both reviewers can agree on many legitimate cases yet kappa can be modest. The code demonstrates this pattern; it does not identify which reviewer is right. Rare-label counts should accompany the statistic.

Interpret disagreement as work

Inspect the actual disagreement cases, especially missed suspected duplicates and policy edge cases. A prevalence shift or different review threshold changes the label marginals and therefore kappa, even when observed agreement remains high. Neither kappa nor percent agreement proves validity against an independent truth. Use a third adjudicator or evidence-based resolution process for disputed cases and preserve the original pair of judgments. Annotation workflows can connect policy revision to a stable label ledger.

Guard downstream evaluation

If only flagged claims are double-reviewed, the agreement sample is selected and cannot describe all claims without an audit design. Report sample construction, each confusion cell, observed agreement, kappa, positive-label prevalence and adjudication outcome. The project holds a model-quality claim until unflagged cases and reviewer disagreements are examined; the frame states which labels the result covers.

Implementation

python
def binary_reviewer_agreement(both_duplicate, both_legitimate,
                              first_only, second_only):
    cells = (both_duplicate, both_legitimate, first_only, second_only)
    if any(cell < 0 or int(cell) != cell for cell in cells) or sum(cells) == 0:
        raise ValueError("nonnegative agreement counts required")
    total = sum(cells)
    observed = (both_duplicate + both_legitimate) / total
    first_positive = (both_duplicate + first_only) / total
    second_positive = (both_duplicate + second_only) / total
    expected = (first_positive * second_positive +
                (1 - first_positive) * (1 - second_positive))
    kappa = None if expected == 1 else (observed - expected) / (1 - expected)
    return {"observed_agreement": observed, "kappa": kappa}

report = binary_reviewer_agreement(2, 88, 5, 5)
assert report["observed_agreement"] == 0.90
assert 0 < report["kappa"] < 0.30

Performance and operating cost

The four-cell calculation is O(1) time and space. Double review and adjudication cost far more than the statistic, but those cases reveal policy ambiguity that a single agreement number hides. A selected easy-case sample can make both cheap agreement and kappa irrelevant to the production queue.

Common Mistakes

  • Calling a high agreement rate proof that labels are correct.
  • Treating a modest kappa under rare positives as a diagnosis by itself.
  • Dropping disagreement cases after adjudication instead of preserving both original labels.
  • Generalizing flagged-only double review to all claims.

Read next

Continue the workflow: Verification sampling: recover label error when flags get unequal review.

ai-data
applied-statistics
Storage details