Compare a receipt-routing change using a protected audit sample, mature outcomes and a visible unobserved population.
Project: audit blind spots in receipt-risk feedback
Freeze the rollout identities
A candidate receipt model reduces manual review volume. Record its model digest, threshold revision, stable assignment cohort and every served decision. The existing review queue supplies labels mainly for routed receipts, so reviewed-only error is not a whole-population quality measure. Create a policy-approved audit lane from otherwise released receipts before outcomes are known; write its inclusion probability with each assignment. Selective-label coverage defines the counts to report.
Wait for comparable labels
Hold back enough time for customer corrections and independent adjudication to mature. Distinguish missing, pending and negative outcomes. Keep audit outcomes separate from labels created by routine risk review. The candidate’s reviewed-only error falls, yet the audit sample finds more missed high-value blurry receipts. Report both observations and their denominators, plus uncertainty for the small affected slice. Outcome joins prevent late corrections from being assigned to the wrong decision revision.
Challenge the apparent win
Replay both models on a shared frozen evaluation set and compare scores and routes, but do not pretend that it reveals outcomes for never-reviewed live receipts. Use the audit lane to estimate released-route error under recorded selection probabilities. If one cohort had no audit support, mark that comparison unresolved and expand coverage rather than extrapolating. The policy-shift lesson explains the support requirement.
Deliver a release decision
Hold expansion until the missed-case signal is resolved, or restore the prior model-and-policy pair if the quality gate fails. Preserve customer exposure, reviewer workload, audit cost and the unresolved population in the packet. A later retrain may use routine labels under the normal lineage process, but the protected evaluation audit cannot quietly become both training data and final approval data. Retraining governance should explicitly record that boundary.
Implementation
def feedback_release_gate(scored, matured, audit_supported,
baseline_missed, candidate_missed):
if scored <= 0 or not 0 <= matured <= scored:
raise ValueError("invalid population")
if not audit_supported or matured / scored < 0.18:
return "hold:insufficient-observation"
if candidate_missed > baseline_missed:
return "hold:missed-cases"
return "review:expansion"
assert feedback_release_gate(4_700, 914, True, 7, 13) == "hold:missed-cases"
assert feedback_release_gate(4_700, 500, True, 7, 5) == "hold:insufficient-observation"
assert feedback_release_gate(4_700, 914, True, 7, 5) == "review:expansion"
Performance and operating cost
The gate is O(1) time and space after population and audit aggregation. Auditing 235 of 4,700 receipts adds reviewer work; delaying rollout until outcomes mature costs time but avoids treating absent labels as success. The count threshold is only a coverage screen, not a statistical significance test; include interval estimates and slice sample sizes in the final packet.
Common Mistakes
- Declaring victory from a lower reviewed-only error rate.
- Assigning audit cases after a customer correction arrives.
- Reusing the protected audit sample to tune and approve the same candidate.
- Treating an unsupported cohort as if its quality matched the observed cohort.
Read next
- Selective labels: measure what the model never lets reviewers see
- Feedback policy shift: compare models when labels depend on routing
- Prediction-outcome joins: evaluate only mature, matched decisions
- Retraining decisions: require a reason and a challenger comparison
- Model experiment guardrails: stop harm without misreading the sample
