Build a verification sample across flagged and unflagged invoices, then decide whether a corrected duplicate rate is defensible.
Project: estimate duplicate-invoice prevalence with audited label error
Freeze a population and a truth rule
An accounts team reports a low duplicate-invoice rate from an automated flag. Define eligible invoices by approval date, one invoice as the counting unit, the rule version, and a consistent independent adjudication standard. Preserve flagged and unflagged invoices in the frame. Decide whether the target is confirmed duplicates, disputed duplicates, or automated flags; these are not the same event. The frame lesson gives the denominator.
Design the verification sample
Sample invoices from both flag groups and record each inclusion probability. Oversample flags if review resources are tight, but retain some unflagged records so missed duplicates and specificity can be learned. Record branch and invoice-value bands if selection varies across them. Keep reviewers from simply copying the automated flag, and sample disagreements for a second adjudication. The verification lesson turns unequal review rates into weighted cell estimates.
Calculate with uncertainty and sensitivity
Estimate flag sensitivity and specificity from the designed validation sample and compare the raw flag rate with a corrected duplicate prevalence under the stable-error assumption. Propagate uncertainty from both the observed flag rate and validation estimates. If a change in rule version alters error rates, produce version-specific results rather than applying one correction to all months. A correction that leaves the zero-to-one range signals an incompatible point model or sampling noise that deserves review. The correction lesson makes the assumption visible.
Release an audit trail
The packet includes eligible counts, review draws and fractions, adjudication protocol, weighted confusion cells, estimated error rates, corrected rate and interval, rule version, and unresolved disagreement. The gate below blocks a headline if one flag stratum was never verified or if adjudicators were not assigned a documented truth rule. Passing it begins statistical review, not a guarantee that manual review is perfect. A recurring audit should check whether the labeling process drifts.
Implementation
def invoice_label_gate(audit):
if not audit["truth_rule_documented"]:
return "hold:adjudication-rule"
if audit["verified_flagged"] == 0 or audit["verified_unflagged"] == 0:
return "hold:verification-support"
if not audit["selection_probabilities_recorded"]:
return "hold:sample-weights"
if audit["unresolved_disagreements"]:
return "hold:adjudication"
return "review:corrected-prevalence"
audit = {"truth_rule_documented": True, "verified_flagged": 37,
"verified_unflagged": 0, "selection_probabilities_recorded": True,
"unresolved_disagreements": 0}
assert invoice_label_gate(audit) == "hold:verification-support"
assert invoice_label_gate({**audit, "verified_unflagged": 29}) == "review:corrected-prevalence"
Performance and operating cost
The gate is O(1), and weighted table construction is O(n) in reviewed invoices. Human verification is the limiting resource. Increasing only the flagged review count cannot identify missed duplicates when the unflagged stratum receives no review.
Common Mistakes
- Using automatic flags as the truth variable.
- Verifying only flags and reporting a corrected population rate.
- Dropping selection probabilities after oversampling.
- Reusing an old rule-version validation estimate without testing stability.
Read next
- Outcome misclassification: correct an observed rate only under stated label-error assumptions
- Verification sampling: recover label error when flags get unequal review
- Population, estimand and sampling frame: name the quantity before calculating
- Reviewer agreement: inspect confusion cells before one kappa number
- Rare proportions: keep interval uncertainty visible at zero and one
