Audit action logs and mature rewards, reject unsupported policy actions and run a bounded canary.
Project: evaluate a pickup-notification bandit before rollout
Freeze the decision frame
A delivery app sends one of three pickup reminders. Each decision captures eligible timings, chosen timing, logging probability, policy revision, order cohort and a decision ID. Join a pickup-completion reward only after its observation window closes. A duplicate client retry keeps its original decision rather than drawing another action. The decision ledger prevents accidental duplication and missing probabilities.
Find the unsupported candidate
The candidate heavily favors late reminders for remote districts. Historic policy logs show zero late allocations there. Do not compute an offline reward for that candidate by pretending early rewards apply to late timing. Restrict the candidate to supported actions for the first comparison or allocate a small, separately approved exploration cohort. Another candidate has nonzero support but effective sample size of only 47 despite 820 recorded decisions; hold it while rare high weights are investigated. The estimate gate exposes both problems.
Protect the live experiment
For a supported candidate with adequate effective sample size, reserve a limited user cohort and a fixed maximum notification rate. Measure completion, cancellations, opt-outs and messages sent, with no overlapping notification experiments in the same orders. Keep the incumbent and a reversible allocation switch. Do not declare a win before the completion window matures. Experiment guardrails keep interactions and exposure accounting honest.
Publish the decision record
Deliver action-support tables, logged probability checks, reward maturity, weight distribution, effective sample size, cohort outcomes and the canary stop rule. Record whether an offline gain survived live measurement. The project is complete only when a poor or unsupported policy has an explicit hold decision, and a safe candidate has an owner for rollback. Link the case to mature rewards and promotion evidence.
Implementation
def pickup_policy_gate(report, limits):
if report["unsupported_contexts"]:
return "hold:unsupported"
if report["effective_n"] < limits["minimum_effective_n"]:
return "hold:low-effective-sample"
if report["notification_rate"] > limits["maximum_rate"]:
return "hold:burden"
if not report["rewards_mature"]:
return "hold:immature-reward"
return "canary:bounded-cohort"
limits = {"minimum_effective_n": 82, "maximum_rate": 0.47}
report = {"unsupported_contexts": 0, "effective_n": 47,
"notification_rate": 0.39, "rewards_mature": True}
assert pickup_policy_gate(report, limits) == "hold:low-effective-sample"
assert pickup_policy_gate({**report, "effective_n": 120}, limits) == "canary:bounded-cohort"
Performance and operating cost
The aggregate gate is O(1) time and space; collecting valid exploration data and mature outcomes is the real cost. Extra notifications can disturb users, so the experiment has a burden cap. A more conservative policy may need longer to gather enough supported examples for a reliable comparison.
Common Mistakes
- Assigning a counterfactual reward to an unchosen reminder.
- Reusing the same order twice because a client retried.
- Ignoring effective sample size when a few decisions dominate the estimate.
- Calling an experiment complete before pickup outcomes mature.
Read next
- Bandit decision logs: preserve action probabilities and delayed rewards
- Bandit offline evaluation: support, variance and release limits
- Model experiment guardrails: stop harm without misreading the sample
- Prediction-outcome joins: evaluate only mature, matched decisions
- Promotion evidence: bind evaluation, contract and rollback to one digest
