A new review threshold changes the labeled population; compare candidates with an observation design that survives that change.
Feedback policy shift: compare models when labels depend on routing
Separate model and policy changes
If a new model and review threshold launch together, observed reviewer error can improve simply because different cases reached review. Pin model digest, score policy, assignment cohort and label channel for each decision. Evaluate the candidate on a shared frozen labeled frame before changing the policy, then run a staged online comparison with an independent audit lane. Threshold control ensures a policy change has its own revision.
Know what offline replay can answer
A frozen dataset can compare scores and routes for examples already labeled, but it cannot reveal outcomes for cases never observed under the old policy. Do not fill missing outcomes with the model prediction and call that ground truth. If logged assignment probabilities and support exist, weighted estimators can be considered; when a route had zero chance of observation, no weighting can recover its missing labels. Audit sampling creates support where permitted.
Gate the online comparison
Use sticky assignment to keep a receipt on one model-and-policy pair. Record exposure only when the decision actually served, then wait for outcome maturity. Compare quality, review load, override rate and unresolved share by cohort. A candidate that reduces review volume may make its observed label rate look better while missing hard cases; audit outcomes are the check. Experiment guardrails address interference and operational limits.
Publish limits explicitly
Report the population covered by labels, the audit inclusion probabilities, effective sample size, uncertainty and any route with no observation support. Mark off-policy estimates as unsupported when their assumptions fail. The right disposition may be “need more audit coverage,” even if the reviewed-only metric looks favorable. The project tests that decision on a receipt triage rollout.
Implementation
def weighted_audit_rate(audits):
if not audits:
raise ValueError("audit rows required")
weighted_errors = 0.0
total_weight = 0.0
for audit in audits:
probability = audit["selection_probability"]
if not 0 < probability <= 1:
raise ValueError("unsupported selection probability")
weight = 1 / probability
weighted_errors += weight * bool(audit["error"])
total_weight += weight
return weighted_errors / total_weight
audits = [{"selection_probability": 0.25, "error": True},
{"selection_probability": 0.50, "error": False},
{"selection_probability": 0.50, "error": False}]
assert weighted_audit_rate(audits) == 0.5
Performance and operating cost
The weighted pass is O(a) time and O(1) extra space for a audited rows. It is a self-normalized descriptive estimate, not a causal proof of a policy improvement; large inverse probabilities increase variance. Independent audits cost reviewer time, and delayed outcomes extend the comparison window. Report uncertainty before using the estimate for promotion.
Common Mistakes
- Treating missing labels as correct predictions.
- Comparing reviewed-only error after changing the review threshold.
- Using inverse weights when some route had zero observation probability.
- Ignoring audit variance because the point estimate improved.
Read next
- Selective labels: measure what the model never lets reviewers see
- Project: audit blind spots in receipt-risk feedback
- Score thresholds are release policy, not model metadata
- Model experiment guardrails: stop harm without misreading the sample
- Prediction-outcome joins: evaluate only mature, matched decisions
Continue the workflow: Bandit offline evaluation: support, variance and release limits.
