Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Delayed bandit rewards and outcome maturity

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Bandit rewards should be evaluated only after a declared observation window, while pending outcomes and late corrections remain visible in the denominator.

Define when a reward exists

For a rapid-lane routing decision, the reward might be an on-time handoff minus a staffing penalty. An on-time status cannot be finalized until the handoff window closes; some cases may remain disputed. Record decision time, outcome deadline, actual label-ready time and current status separately. The code partitions a log into mature and pending records at a fixed audit time. The action log links those records by immutable ID.

Do not call pending cases failures

If rapid-lane successes tend to be recorded quickly while standard-lane failures appear after several days, scoring only currently labeled cases biases a comparison. The opposite pattern is also possible. Report how many decisions have matured by action, depot and time cohort. A row with a missing reward is not automatically a zero reward. Delayed-label monitoring applies the same principle to supervised models.

Handle censoring as a design issue

For some shipments, the final outcome may never be recorded. A fixed maturity window reduces ordinary delay but does not solve outcome-dependent missingness. Trace missing labels, compare data-collection routes and consider sensitivity ranges. Inverse-propensity weights correct action selection under their assumptions; they do not automatically correct selective reward observation. IPS therefore needs a mature, sufficiently observed cohort.

Separate interim and final rewards

A quick scan confirmation is a proxy for the eventual service outcome. If used for fast learning, label it as an interim reward and compare it with the final reward on matured cases. Do not change a production policy on a proxy gain without checking downstream harm. Proxy measurement explains how apparent improvement can reflect the wrong target.

Keep the evaluation clock fixed

Specify the audit time and maturity age before viewing outcome rates. A later backfill can update an analysis but should not rewrite what the model knew at a historical decision. For policy comparison, use the same reward window and case eligibility for every candidate. The release review refuses a value estimate built from incompatible clocks.

Implementation

python
# Decision ID, chosen lane, decision day, reward-ready day, final reward.
outcomes = [
    ("R901", "rapid", 1, 4, 1.0),
    ("R902", "standard", 2, 7, 0.0),
    ("R903", "rapid", 4, 9, 1.0),
    ("R904", "standard", 5, None, None),
]

def reward_maturity(rows, audit_day):
    mature, pending = [], []
    for decision_id, lane, decision_day, ready_day, reward in rows:
        if ready_day is not None and ready_day <= audit_day:
            if reward is None or ready_day < decision_day:
                raise ValueError("invalid mature reward")
            mature.append((decision_id, lane, reward))
        else:
            pending.append((decision_id, lane))
    return mature, pending

mature, pending = reward_maturity(outcomes, audit_day=7)
assert [decision_id for decision_id, _, _ in mature] == ["R901", "R902"]
assert len(pending) == 2

Performance and operating cost

Partitioning N decisions costs O(N) time and O(N) memory when retaining case lists, or O(A) counters for A actions when streamed. The difficult costs are durable joins, label-quality review and waiting long enough for the outcome to represent the declared reward.

Common Mistakes

  • Do not encode pending or unknown rewards as zero.
  • Do not compare actions with different maturity windows.
  • Do not assume action-selection weighting fixes missing-outcome bias.

Read next

Continue the workflow: Episode returns, temporal difference and terminal states.

ai-data
machine-learning
Storage details