Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: evaluate a pickup-notification bandit before rollout

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Audit action logs and mature rewards, reject unsupported policy actions and run a bounded canary.

Freeze the decision frame

A delivery app sends one of three pickup reminders. Each decision captures eligible timings, chosen timing, logging probability, policy revision, order cohort and a decision ID. Join a pickup-completion reward only after its observation window closes. A duplicate client retry keeps its original decision rather than drawing another action. The decision ledger prevents accidental duplication and missing probabilities.

Find the unsupported candidate

The candidate heavily favors late reminders for remote districts. Historic policy logs show zero late allocations there. Do not compute an offline reward for that candidate by pretending early rewards apply to late timing. Restrict the candidate to supported actions for the first comparison or allocate a small, separately approved exploration cohort. Another candidate has nonzero support but effective sample size of only 47 despite 820 recorded decisions; hold it while rare high weights are investigated. The estimate gate exposes both problems.

Protect the live experiment

For a supported candidate with adequate effective sample size, reserve a limited user cohort and a fixed maximum notification rate. Measure completion, cancellations, opt-outs and messages sent, with no overlapping notification experiments in the same orders. Keep the incumbent and a reversible allocation switch. Do not declare a win before the completion window matures. Experiment guardrails keep interactions and exposure accounting honest.

Publish the decision record

Deliver action-support tables, logged probability checks, reward maturity, weight distribution, effective sample size, cohort outcomes and the canary stop rule. Record whether an offline gain survived live measurement. The project is complete only when a poor or unsupported policy has an explicit hold decision, and a safe candidate has an owner for rollback. Link the case to mature rewards and promotion evidence.

Implementation

python
def pickup_policy_gate(report, limits):
    if report["unsupported_contexts"]:
        return "hold:unsupported"
    if report["effective_n"] < limits["minimum_effective_n"]:
        return "hold:low-effective-sample"
    if report["notification_rate"] > limits["maximum_rate"]:
        return "hold:burden"
    if not report["rewards_mature"]:
        return "hold:immature-reward"
    return "canary:bounded-cohort"

limits = {"minimum_effective_n": 82, "maximum_rate": 0.47}
report = {"unsupported_contexts": 0, "effective_n": 47,
          "notification_rate": 0.39, "rewards_mature": True}
assert pickup_policy_gate(report, limits) == "hold:low-effective-sample"
assert pickup_policy_gate({**report, "effective_n": 120}, limits)        == "canary:bounded-cohort"

Performance and operating cost

The aggregate gate is O(1) time and space; collecting valid exploration data and mature outcomes is the real cost. Extra notifications can disturb users, so the experiment has a burden cap. A more conservative policy may need longer to gather enough supported examples for a reliable comparison.

Common Mistakes

  • Assigning a counterfactual reward to an unchosen reminder.
  • Reusing the same order twice because a client retried.
  • Ignoring effective sample size when a few decisions dominate the estimate.
  • Calling an experiment complete before pickup outcomes mature.

Read next

ai-data
mlops
Storage details