Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Verification sampling: recover label error when flags get unequal review

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A stratified adjudication sample needs its review probabilities in the analysis, especially when flagged and unflagged cases are sampled differently.

Name the verification mechanism

An invoice team sends nearly every automated flag to a reviewer but audits only a fraction of unflagged invoices. The adjudicated set is therefore not a simple random sample of all eligible invoices. Computing sensitivity, specificity, or prevalence from that reviewed set without its sampling fractions overrepresents flagged items. Build the full flag-by-review ledger and retain the probability that each invoice entered manual verification. The survey-weight lesson uses the same population accounting idea.

Weight the validated cells

For a sampled invoice, use inverse verification probability to estimate how much of the full eligible population that record represents. Summing weights within the true-label by automated-flag cells gives estimated population cell totals. Divide the weighted true-positive total by all weighted true positives and false negatives for sensitivity; analogous weighted cells yield specificity. The code illustrates the weighted table, not its uncertainty interval. Positivity matters: if no unflagged invoice can be reviewed, missed duplicates cannot be estimated from this design.

Avoid designing a collider

If review probability also depends on branch, invoice size, or other recorded risk cues, the weighting model must include those cues or the design must stratify on them. If reviewers see the original flag, their adjudication can be influenced by it; consider blinded or independently checked reviews for a sample. The reviewer-agreement lesson handles disagreement, which is separate from unequal selection into review.

Report design and uncertainty

Show eligible counts by flag status, verified counts, sampling fractions, effective weight concentration, adjudication procedure, and uncertainty that reflects the stratified design. A point estimate from a handful of unflagged audits can be fragile even when thousands of flagged invoices were reviewed. The correction lesson can then use the validation estimates under its own assumptions. The project gates both steps before publication.

Implementation

python
def weighted_verification_cells(reviewed_invoices):
    cells = {(truth, flag): 0.0 for truth in (False, True)
             for flag in (False, True)}
    for confirmed_duplicate, auto_flagged, review_probability in reviewed_invoices:
        if not 0 < review_probability <= 1:
            raise ValueError("positive verification probability required")
        cells[(bool(confirmed_duplicate), bool(auto_flagged))] +=             1 / review_probability
    return cells

cells = weighted_verification_cells([(True, True, 1.0),
                                     (False, True, 1.0),
                                     (True, False, 0.25)])
assert cells[(True, False)] == 4.0
assert cells[(True, True)] == 1.0

Performance and operating cost

Accumulating n reviewed invoices takes O(n) time and O(1) table space for binary truth and flag states. Checking frame completeness and adjudication quality dominates. Extremely small review probabilities create unstable weights rather than free information.

Common Mistakes

  • Estimating specificity from a flagged-only review file.
  • Assigning zero to a cell that could never enter verification.
  • Ignoring branch-dependent verification probabilities.
  • Treating reviewer consensus as infallible ground truth.

Read next

ai-data
applied-statistics
Storage details