Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Feedback policy shift: compare models when labels depend on routing

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A new review threshold changes the labeled population; compare candidates with an observation design that survives that change.

Separate model and policy changes

If a new model and review threshold launch together, observed reviewer error can improve simply because different cases reached review. Pin model digest, score policy, assignment cohort and label channel for each decision. Evaluate the candidate on a shared frozen labeled frame before changing the policy, then run a staged online comparison with an independent audit lane. Threshold control ensures a policy change has its own revision.

Know what offline replay can answer

A frozen dataset can compare scores and routes for examples already labeled, but it cannot reveal outcomes for cases never observed under the old policy. Do not fill missing outcomes with the model prediction and call that ground truth. If logged assignment probabilities and support exist, weighted estimators can be considered; when a route had zero chance of observation, no weighting can recover its missing labels. Audit sampling creates support where permitted.

Gate the online comparison

Use sticky assignment to keep a receipt on one model-and-policy pair. Record exposure only when the decision actually served, then wait for outcome maturity. Compare quality, review load, override rate and unresolved share by cohort. A candidate that reduces review volume may make its observed label rate look better while missing hard cases; audit outcomes are the check. Experiment guardrails address interference and operational limits.

Publish limits explicitly

Report the population covered by labels, the audit inclusion probabilities, effective sample size, uncertainty and any route with no observation support. Mark off-policy estimates as unsupported when their assumptions fail. The right disposition may be “need more audit coverage,” even if the reviewed-only metric looks favorable. The project tests that decision on a receipt triage rollout.

Implementation

python
def weighted_audit_rate(audits):
    if not audits:
        raise ValueError("audit rows required")
    weighted_errors = 0.0
    total_weight = 0.0
    for audit in audits:
        probability = audit["selection_probability"]
        if not 0 < probability <= 1:
            raise ValueError("unsupported selection probability")
        weight = 1 / probability
        weighted_errors += weight * bool(audit["error"])
        total_weight += weight
    return weighted_errors / total_weight

audits = [{"selection_probability": 0.25, "error": True},
          {"selection_probability": 0.50, "error": False},
          {"selection_probability": 0.50, "error": False}]
assert weighted_audit_rate(audits) == 0.5

Performance and operating cost

The weighted pass is O(a) time and O(1) extra space for a audited rows. It is a self-normalized descriptive estimate, not a causal proof of a policy improvement; large inverse probabilities increase variance. Independent audits cost reviewer time, and delayed outcomes extend the comparison window. Report uncertainty before using the estimate for promotion.

Common Mistakes

  • Treating missing labels as correct predictions.
  • Comparing reviewed-only error after changing the review threshold.
  • Using inverse weights when some route had zero observation probability.
  • Ignoring audit variance because the point estimate improved.

Read next

Continue the workflow: Bandit offline evaluation: support, variance and release limits.

ai-data
mlops
Storage details