Bandit rewards should be evaluated only after a declared observation window, while pending outcomes and late corrections remain visible in the denominator.
Delayed bandit rewards and outcome maturity
Define when a reward exists
For a rapid-lane routing decision, the reward might be an on-time handoff minus a staffing penalty. An on-time status cannot be finalized until the handoff window closes; some cases may remain disputed. Record decision time, outcome deadline, actual label-ready time and current status separately. The code partitions a log into mature and pending records at a fixed audit time. The action log links those records by immutable ID.
Do not call pending cases failures
If rapid-lane successes tend to be recorded quickly while standard-lane failures appear after several days, scoring only currently labeled cases biases a comparison. The opposite pattern is also possible. Report how many decisions have matured by action, depot and time cohort. A row with a missing reward is not automatically a zero reward. Delayed-label monitoring applies the same principle to supervised models.
Handle censoring as a design issue
For some shipments, the final outcome may never be recorded. A fixed maturity window reduces ordinary delay but does not solve outcome-dependent missingness. Trace missing labels, compare data-collection routes and consider sensitivity ranges. Inverse-propensity weights correct action selection under their assumptions; they do not automatically correct selective reward observation. IPS therefore needs a mature, sufficiently observed cohort.
Separate interim and final rewards
A quick scan confirmation is a proxy for the eventual service outcome. If used for fast learning, label it as an interim reward and compare it with the final reward on matured cases. Do not change a production policy on a proxy gain without checking downstream harm. Proxy measurement explains how apparent improvement can reflect the wrong target.
Keep the evaluation clock fixed
Specify the audit time and maturity age before viewing outcome rates. A later backfill can update an analysis but should not rewrite what the model knew at a historical decision. For policy comparison, use the same reward window and case eligibility for every candidate. The release review refuses a value estimate built from incompatible clocks.
Implementation
# Decision ID, chosen lane, decision day, reward-ready day, final reward.
outcomes = [
("R901", "rapid", 1, 4, 1.0),
("R902", "standard", 2, 7, 0.0),
("R903", "rapid", 4, 9, 1.0),
("R904", "standard", 5, None, None),
]
def reward_maturity(rows, audit_day):
mature, pending = [], []
for decision_id, lane, decision_day, ready_day, reward in rows:
if ready_day is not None and ready_day <= audit_day:
if reward is None or ready_day < decision_day:
raise ValueError("invalid mature reward")
mature.append((decision_id, lane, reward))
else:
pending.append((decision_id, lane))
return mature, pending
mature, pending = reward_maturity(outcomes, audit_day=7)
assert [decision_id for decision_id, _, _ in mature] == ["R901", "R902"]
assert len(pending) == 2Performance and operating cost
Partitioning N decisions costs O(N) time and O(N) memory when retaining case lists, or O(A) counters for A actions when streamed. The difficult costs are durable joins, label-quality review and waiting long enough for the outcome to represent the declared reward.
Common Mistakes
- Do not encode pending or unknown rewards as zero.
- Do not compare actions with different maturity windows.
- Do not assume action-selection weighting fixes missing-outcome bias.
Read next
- Contextual bandit action logging and support
- Inverse-propensity policy value
- Exploration budget and action guardrails
- Contextual bandit policy release review project
Continue the workflow: Episode returns, temporal difference and terminal states.
