Agreement between two reviewers depends on observed matches, their label marginals and the prevalence of each label.
Reviewer agreement: inspect confusion cells before one kappa number
Audit the labeling setup
Two claims reviewers independently classify the same sampled cases as duplicate or legitimate. Independence means neither sees the other’s answer before submitting; it does not mean their judgments are statistically independent. Define the label policy, ambiguous-case route, random audit sample and allowed evidence. An agreement rate from handpicked easy claims does not represent the queue. Diagnostic estimates inherit any error in the reference labels.
Read the complete two-by-two table
Record cases both call duplicate, both call legitimate, and the two disagreement directions. Observed agreement is the share on the matching diagonal. Cohen’s kappa subtracts an expected agreement derived from each reviewer’s label proportions and rescales the remainder. When true duplicates are rare, both reviewers can agree on many legitimate cases yet kappa can be modest. The code demonstrates this pattern; it does not identify which reviewer is right. Rare-label counts should accompany the statistic.
Interpret disagreement as work
Inspect the actual disagreement cases, especially missed suspected duplicates and policy edge cases. A prevalence shift or different review threshold changes the label marginals and therefore kappa, even when observed agreement remains high. Neither kappa nor percent agreement proves validity against an independent truth. Use a third adjudicator or evidence-based resolution process for disputed cases and preserve the original pair of judgments. Annotation workflows can connect policy revision to a stable label ledger.
Guard downstream evaluation
If only flagged claims are double-reviewed, the agreement sample is selected and cannot describe all claims without an audit design. Report sample construction, each confusion cell, observed agreement, kappa, positive-label prevalence and adjudication outcome. The project holds a model-quality claim until unflagged cases and reviewer disagreements are examined; the frame states which labels the result covers.
Implementation
def binary_reviewer_agreement(both_duplicate, both_legitimate,
first_only, second_only):
cells = (both_duplicate, both_legitimate, first_only, second_only)
if any(cell < 0 or int(cell) != cell for cell in cells) or sum(cells) == 0:
raise ValueError("nonnegative agreement counts required")
total = sum(cells)
observed = (both_duplicate + both_legitimate) / total
first_positive = (both_duplicate + first_only) / total
second_positive = (both_duplicate + second_only) / total
expected = (first_positive * second_positive +
(1 - first_positive) * (1 - second_positive))
kappa = None if expected == 1 else (observed - expected) / (1 - expected)
return {"observed_agreement": observed, "kappa": kappa}
report = binary_reviewer_agreement(2, 88, 5, 5)
assert report["observed_agreement"] == 0.90
assert 0 < report["kappa"] < 0.30
Performance and operating cost
The four-cell calculation is O(1) time and space. Double review and adjudication cost far more than the statistic, but those cases reveal policy ambiguity that a single agreement number hides. A selected easy-case sample can make both cheap agreement and kappa irrelevant to the production queue.
Common Mistakes
- Calling a high agreement rate proof that labels are correct.
- Treating a modest kappa under rare positives as a diagnosis by itself.
- Dropping disagreement cases after adjudication instead of preserving both original labels.
- Generalizing flagged-only double review to all claims.
Read next
- Diagnostic performance: separate sensitivity, specificity and predictive value
- Project: separate claims-flagger errors from reviewer disagreement
- Rare proportions: keep interval uncertainty visible at zero and one
- Population, estimand and sampling frame: name the quantity before calculating
- Data Annotation & Label Quality Tutorial
Continue the workflow: Verification sampling: recover label error when flags get unequal review.
