A stratified adjudication sample needs its review probabilities in the analysis, especially when flagged and unflagged cases are sampled differently.
Verification sampling: recover label error when flags get unequal review
Name the verification mechanism
An invoice team sends nearly every automated flag to a reviewer but audits only a fraction of unflagged invoices. The adjudicated set is therefore not a simple random sample of all eligible invoices. Computing sensitivity, specificity, or prevalence from that reviewed set without its sampling fractions overrepresents flagged items. Build the full flag-by-review ledger and retain the probability that each invoice entered manual verification. The survey-weight lesson uses the same population accounting idea.
Weight the validated cells
For a sampled invoice, use inverse verification probability to estimate how much of the full eligible population that record represents. Summing weights within the true-label by automated-flag cells gives estimated population cell totals. Divide the weighted true-positive total by all weighted true positives and false negatives for sensitivity; analogous weighted cells yield specificity. The code illustrates the weighted table, not its uncertainty interval. Positivity matters: if no unflagged invoice can be reviewed, missed duplicates cannot be estimated from this design.
Avoid designing a collider
If review probability also depends on branch, invoice size, or other recorded risk cues, the weighting model must include those cues or the design must stratify on them. If reviewers see the original flag, their adjudication can be influenced by it; consider blinded or independently checked reviews for a sample. The reviewer-agreement lesson handles disagreement, which is separate from unequal selection into review.
Report design and uncertainty
Show eligible counts by flag status, verified counts, sampling fractions, effective weight concentration, adjudication procedure, and uncertainty that reflects the stratified design. A point estimate from a handful of unflagged audits can be fragile even when thousands of flagged invoices were reviewed. The correction lesson can then use the validation estimates under its own assumptions. The project gates both steps before publication.
Implementation
def weighted_verification_cells(reviewed_invoices):
cells = {(truth, flag): 0.0 for truth in (False, True)
for flag in (False, True)}
for confirmed_duplicate, auto_flagged, review_probability in reviewed_invoices:
if not 0 < review_probability <= 1:
raise ValueError("positive verification probability required")
cells[(bool(confirmed_duplicate), bool(auto_flagged))] += 1 / review_probability
return cells
cells = weighted_verification_cells([(True, True, 1.0),
(False, True, 1.0),
(True, False, 0.25)])
assert cells[(True, False)] == 4.0
assert cells[(True, True)] == 1.0
Performance and operating cost
Accumulating n reviewed invoices takes O(n) time and O(1) table space for binary truth and flag states. Checking frame completeness and adjudication quality dominates. Extremely small review probabilities create unstable weights rather than free information.
Common Mistakes
- Estimating specificity from a flagged-only review file.
- Assigning zero to a cell that could never enter verification.
- Ignoring branch-dependent verification probabilities.
- Treating reviewer consensus as infallible ground truth.
Read next
- Outcome misclassification: correct an observed rate only under stated label-error assumptions
- Project: estimate duplicate-invoice prevalence with audited label error
- Stratified survey estimates: weight toward the named population
- Reviewer agreement: inspect confusion cells before one kappa number
- Diagnostic performance: separate sensitivity, specificity and predictive value
