Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Sequential decision state and reward contract

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A sequential decision process records the state before an action, the action taken, immediate reward, next state and whether the episode has ended.

Distinguish a sequence from a one-step choice

A rapid-lane assignment has an immediate result but may also change the next queue length and the choices available to later shipments. When actions alter future state, evaluating each intake as an isolated bandit can miss delayed capacity cost. Define a decision interval, state snapshot, eligible actions, transition and terminal condition. The bandit log remains useful for one-step action probabilities; sequential evaluation also needs next-state records.

Put enough history into the state

A stock-control state containing only current bin level may be inadequate if an outstanding replenishment order will arrive tomorrow. The Markov assumption says that, conditional on recorded state and action, earlier history adds no information to the next-state and reward distribution. It is an assumption to test, not a label attached to any convenient feature vector. The code validates transition shape and legal actions for a small warehouse process.

Define reward from actual operations

Dispatching a parcel can earn service value, replenishing costs labor, and a shortage can impose a later penalty. Choose units and timing before optimizing; otherwise an agent may improve the logged immediate score while increasing real delays. Include capacity, safety and prohibited-action rules outside the reward as hard constraints where necessary. Cost-based decisions show why the business action contract matters.

Track episode boundaries

A terminal state has no future value to bootstrap. A time-limit cutoff can be a truncation rather than a genuine terminal business outcome; those cases may need different handling. Record both the stop reason and reward maturity. Temporal-difference updates use this distinction; delayed rewards affect the input log.

Check who generated the trajectory

A historical policy determines which state-action pairs appear in data. A candidate that repeatedly chooses a rare action can move into states the log never visited. A single-step propensity check does not cover that entire path. Offline trajectory support audits the sequence-level gap before release.

Implementation

python
eligible_actions = {
    "low-stock": {"replenish", "wait"},
    "ready-stock": {"dispatch", "hold"},
    "closed": set(),
}
transitions = [
    ("low-stock", "replenish", -2.0, "ready-stock", False),
    ("ready-stock", "dispatch", 5.0, "low-stock", False),
    ("low-stock", "wait", -7.0, "closed", True),
]

def validate_trajectory(rows, action_map):
    for position, (state, action, reward, next_state, terminated) in enumerate(rows):
        if action not in action_map[state] or next_state not in action_map:
            raise ValueError("invalid state-action transition")
        if terminated != (next_state == "closed"):
            raise ValueError("terminal state mismatch")
        if position and state != rows[position - 1][3]:
            raise ValueError("broken trajectory chain")
    return len(rows)

assert validate_trajectory(transitions, eligible_actions) == 3

Performance and operating cost

Validating T recorded transitions costs O(T) time and O(SA) action-map storage for S states and their eligible actions. Real logs also need timestamps, policy versions, state history, reward maturity and constraints. A larger state space increases collection and coverage demands faster than simple validation cost.

Common Mistakes

  • Do not call a one-step score the value of a policy that changes future state.
  • Do not mark every time limit as a true terminal business outcome.
  • Do not encode a safety constraint only as a small negative reward.

Read next

ai-data
machine-learning
Storage details