Build a support conversation flow that remembers corrected order references, asks when uncertain and never issues an action without verification and confirmation.
Project: track a refund conversation with correction-safe handoff
Define the conversation
A customer asks about a refund, first names order ZX-47, then corrects it to ZX-82. The system must retain the correction event, make ZX-82 current and treat both as unverified until an authorized lookup. A customer-visible response may explain the next step but cannot claim a refund is complete. The output includes current state, verification status, pending question and handoff reason.
Build reviewed turns
Annotate full multi-turn cases with slot updates, explicit corrections, pronoun references and confirmation events. Include copied agent messages, mixed languages and multiple orders. Split by conversation and customer before training or evaluation; a single turn in both sets would leak the answer. The state contract defines set, clear and unchanged; entity extraction supplies candidate identifiers.
Rehearse action gates
Replay each reviewed case turn by turn. Compare joint state, correction success, unnecessary clarification rate and any wrong-action attempt. Stub the verification tool so tests can simulate ownership failure and ineligible refunds without touching customer accounts. A case passes only if no external action occurs until the current order is verified and the customer confirms. The handoff gate makes that condition explicit.
Release with review
Deploy the state model, schema and policy together. Keep an audit trail of state changes and tool decisions under support access controls. Route unresolved references and unsupported language mixes to an agent with a concise evidence-backed summary. Review a sample of successful automated cases as well as handoffs. An increase in completed conversations is not a win if the wrong-action rate rises.
Implementation
def replay_order_correction(events):
state = {}
for event in events:
if event["operation"] == "set-order":
state = {**state, "order_id": event["value"], "verified": False}
elif event["operation"] == "verify-order":
if state.get("order_id") != event["value"]:
raise ValueError("verification is for an old order")
state["verified"] = True
return state
turns = [{"operation": "set-order", "value": "ZX-47"},
{"operation": "set-order", "value": "ZX-82"},
{"operation": "verify-order", "value": "ZX-82"}]
assert replay_order_correction(turns) == {"order_id": "ZX-82", "verified": True}
Performance and operating cost
Replay is O(t) time for t turn events and O(s) active state for s slots; retained audit history is O(t). The expensive steps are verification calls and human resolution, so measure p95 tool latency and handoff completion time. Keep validation strict even if it adds a turn: a wrong refund has a different cost from a delayed but correct reply.
Common Mistakes
- Verifying an earlier order after the customer corrected it.
- Treating a model answer as confirmation from the customer.
- Evaluating each turn independently instead of replaying a whole case.
- Logging the entire conversation into an unrestricted metrics stream.
Read next
- Dialogue state: apply slot updates, corrections and deletions
- Dialogue handoff: gate actions on verified state and uncertainty
- Project: ship an auditable support-entity extractor
- Project: route multilingual tickets with measured abstention
- Project: enforce a privacy-safe support-text pipeline
Continue the workflow: Project: audit support commitments and conversation repairs.
