A conversation state is a hypothesis about user intent. External actions need verified fields, authorization and an explicit confirmation boundary.
Dialogue handoff: gate actions on verified state and uncertainty
Separate understanding from execution
The assistant may infer that a customer wants a refund, but intent classification is not permission to issue one. Mark slots as asserted, extracted, verified or disputed. Verify order ownership and refund eligibility through the authorized system, then ask for confirmation of the exact action. A high model score cannot replace the account check. The slot reducer tracks which value is current; this lesson decides whether it is actionable.
Design a clear handoff
Escalation should carry a concise state snapshot, source turn IDs, open questions, verification results and why automation stopped. Do not dump an unfiltered conversation into a broad queue. If a human changes a slot, record the correction as a new event rather than overwriting history. The next model turn must see that correction or abstain; otherwise the assistant may repeat a rejected guess. Preserve the original user wording behind role-based access.
Test the failure paths
Exercise account mismatch, corrected order ID, multiple orders, ambiguous pronouns, unsupported language mix, expired eligibility and interrupted confirmation. Evaluate wrong-action rate, missed useful handoff, repeated clarification loops and time to human resolution. A state tracker can score well on slots while still acting too early. Abstention and multilingual fallback provide related guardrails.
Operate with versioned policy
Store dialogue model, slot schema, action rules and tool permissions as one release bundle. Changes to an API authorization rule should invalidate cached verification even if the text model is unchanged. Audit every external action with the confirmed state and policy version, not raw model text alone. The refund project provides an end-to-end rehearsal before enabling any live action.
Implementation
def refund_action_gate(state, order_verified, eligible, confirmed):
if state.get("intent") != "refund":
return "handoff", "intent-not-confirmed"
if not state.get("order_id") or not order_verified:
return "handoff", "order-not-verified"
if not eligible:
return "handoff", "eligibility-unconfirmed"
if not confirmed:
return "ask-confirmation", "action-not-confirmed"
return "ready-for-authorized-tool", "all-gates-passed"
assert refund_action_gate({"intent": "refund", "order_id": "ZX-82"},
True, True, False)[0] == "ask-confirmation"
Performance and operating cost
This policy gate is O(1) time and space. Verification calls, human handoff and waiting for user confirmation dominate latency. Measure wrong-action prevention and resolution time, not only turn-level state accuracy. Cache verification cautiously: ownership or eligibility can change, and a stale positive response should not permit an external action.
Common Mistakes
- Calling an external refund tool after intent detection alone.
- Passing a handoff without source turn IDs or unresolved questions.
- Reusing verification after the customer changes the order ID.
- Allowing a model update to change action policy without a separate review.
Read next
- Dialogue state: apply slot updates, corrections and deletions
- Project: track a refund conversation with correction-safe handoff
- Text classification evaluation: inspect slices and allow abstention
- Calibrate multilingual text decisions and fallback routes
- Text inference: package tokenizer, labels and reject paths
Continue the workflow: Speaker turns: attribution, overlap and unresolved voices.
Continue the workflow: Conversation repair: corrections, clarification and unresolved targets.
