Track requests, promises and corrections in support conversations without turning proposed language into an authorized action.
Project: audit support commitments and conversation repairs
Set the workflow boundary
A customer reports a failed gateway and asks for a retry-policy change. An agent promises to investigate, then corrects the affected region. The system may stage a task and a proposed correction; it may not modify production policy from text alone. Store speaker, turn, source revision, act, target and authorization state. Speech-act policy separates request, report, promise and approval.
Construct conversations that break shortcuts
Include “yes” without a clear target, quoted promises, agent reports of completed work, a customer correction after several similar turns and an ASR transcript with overlapping speakers. Group each full conversation in one evaluation split. Reviewers mark the act, target, repaired field and whether the speaker had authority. Add a case where a generated assistant promise is never carried out; the task must remain open until separate completion evidence arrives.
Stage commitments and repairs
Classify turns with history, but keep unresolved targets in a clarification queue. Apply a repair as a proposed overlay that preserves the old claim. Create a commitment record only when an identified speaker makes a definite promise, and keep it distinct from the operational action. Before any external change, check account authorization through the real action service. Repair state ensures a corrected region does not rewrite unrelated history.
Release on workflow safety
Measure false commitments, missed promises, wrong repair targets, unsupported completion and unauthorized action attempts. Review the final task queue and conversation transcript side by side. Shadow the workflow before enabling task creation. If an ambiguous short acknowledgment begins closing tasks, roll back the act policy and inspect affected conversations. A good label score cannot excuse an incorrect account or production change.
Implementation
def stage_commitment(act, speaker_role, target_action, authorized):
if act != "commitment":
return {"state": "no-commitment"}
if not target_action:
return {"state": "clarify", "reason": "action-target-missing"}
if speaker_role != "agent" or not authorized:
return {"state": "review", "reason": "authority-unverified"}
return {"state": "task-proposed", "action": target_action,
"completed": False}
task = stage_commitment("commitment", "agent", "inspect gateway-west logs", True)
assert task == {"state": "task-proposed",
"action": "inspect gateway-west logs", "completed": False}
assert stage_commitment("report", "agent", "inspect logs", True)["state"] == "no-commitment"
Performance and operating cost
This staging gate is O(1) per turn. Linking targets and checking authority against external systems add latency; do not replace them with a language-model score. Human review cost should be concentrated on ambiguous commitments and corrections that affect actions. Track cost per correctly staged task and false-completion rate, not just classification throughput.
Common Mistakes
- Executing a production change because text was classified as a commitment.
- Marking a task done when only a promise was recorded.
- Accepting approval from a speaker without checking account authority.
- Erasing the original region after a correction.
Read next
- Speech acts: requests, reports and commitments in support dialogue
- Conversation repair: corrections, clarification and unresolved targets
- Dialogue handoff: gate actions on verified state and uncertainty
- Project: track a refund conversation with correction-safe handoff
- Project: review support-call transcripts before handoff
