Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Offline trajectory support and simulator risk

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An offline sequential policy can enter state-action paths missing from the logged behavior, so its estimated value needs path support and a separately checked environment model.

Look beyond one-step overlap

A historical stock policy may try both replenishment and waiting in low-stock states, yet never dispatch after a particular replenishment delay. A candidate that repeatedly chooses a new sequence changes the state mix. One-step action probabilities are necessary for some estimators but do not by themselves establish coverage of the resulting trajectories. Bandit support is a simpler, single-decision boundary.

Count observed state-action pairs

The code compares a proposed state-action plan with counts in a fixed transition log. Missing pairs are blockers for direct evidence. Counts above zero are only a coarse screen: context may be too broad, and rare pairs produce unstable estimates. Include state features, time, policy version and next-state distributions in a serious audit. The state definition determines what the counts mean.

Do not trust a learned simulator by construction

Value iteration can find an optimal policy for any supplied transition model, including a wrong one. Test a simulator’s one-step transitions and reward predictions on held-out episodes, then check longer rollouts for compounding error. A model that matches common dispatches can still fail on rare shortage paths that matter most. Value iteration solves the stated model, not the real warehouse.

Keep conservative alternatives

Where a candidate enters unsupported states, stay with a reviewed behavior policy, restrict the action set, or collect controlled evidence under operational approval. A synthetic rollout is a stress test, not observed proof. Report coverage gaps and sensitivity to transition assumptions rather than a single projected return. Q-learning cannot repair absent transitions merely by repeated updates.

Plan prospective review

Before a live pilot, name capacity and safety guardrails, assignment logging, monitoring windows and rollback criteria. If actions affect future customer or staff outcomes, evaluate those outcomes at their real maturity time. The release project keeps unsupported claims visible.

Implementation

python
# Historical transitions are summarized by observed state and action.
logged_visits = {
    ("low-stock", "replenish"): 47,
    ("low-stock", "wait"): 13,
    ("ready-stock", "dispatch"): 52,
    ("ready-stock", "hold"): 0,
}
candidate_plan = [("low-stock", "replenish"),
                  ("ready-stock", "hold"),
                  ("ready-stock", "dispatch")]

def support_screen(plan, counts, minimum_visits):
    return [{"state": state, "action": action,
             "visits": counts.get((state, action), 0)}
            for state, action in plan
            if counts.get((state, action), 0) < minimum_visits]

gaps = support_screen(candidate_plan, logged_visits, minimum_visits=8)
assert gaps == [{"state": "ready-stock", "action": "hold", "visits": 0}]

Performance and operating cost

Screening P proposed state-action entries against a hash map costs O(P) expected time and O(G) output space for G gaps. Reliable sequential evaluation is much more costly: coverage deteriorates with path length, and simulator error can compound across transitions. A low-cost count is a blocker check, not a value estimator.

Common Mistakes

  • Do not treat a one-step supported action as proof that a new trajectory is supported.
  • Do not present simulated long-run reward as observed policy value.
  • Do not use a nonzero visit count as evidence of adequate precision without inspecting its context.

Read next

ai-data
machine-learning
Storage details