A voice-agent evaluation needs synchronized input audio, recognition events, output audio, tool events, and a final account state. A text transcript alone hides early tool calls, unheard speech, and false interruption. Build cases for self-correction, pause inside a sentence, digit confusion, background speech, genuine barge-in, incidental noise, slow lookup, and ambiguous effect. Score whether the agent used a final turn, selected the right entity, confirmed action-changing fields, stopped stale playback, respected tool authority, and described the confirmed outcome. Set hard failures for wrong-account actions and duplicate effects. Review recordings with privacy controls because raw calls may contain sensitive speech.
Real-time voice evaluations: test timing and state, not transcript fluency
Operational case
A test suite for the library renewal agent contains 38 calls: 7 code corrections, 6 mid-sentence pauses, 5 noisy digit pairs, 5 true interruptions, 4 incidental sounds, 6 slow-tool cases, and 5 lost-response cases. A new prompt reduces repeated confirmations but renews BK-412 in one correction call before the final transcript says BK-421. That is a critical failure despite better average latency. The team records the audio event sequence, action time, receipt, and loan state, then blocks release. A separate case catches a false barge-in caused by a dropped metal cup; it counts as a usability defect, not a wrong-record effect.
Case VC-17: interim BK-412; final BK-421; no call before final.
Case VC-24: cup noise; playback should continue.
Case VC-31: submit timeout; one receipt or ambiguous state.
Release: zero wrong-record renewals and zero duplicate effects.Performance and operating cost
For C recorded calls, V prompt variants, and M review modes, exhaustive comparison takes O(CVM) runs plus audio playback time proportional to total duration. Automated event assertions can catch early calls and duplicate receipts, while human listening is needed for intelligibility and false interrupts. Keep a fixed critical set on every candidate release and sample less risky pronunciation variants separately. Store model, recognizer, endpoint, voice, and tool versions with each case; a prompt comparison is misleading when the audio pipeline changed at the same time. Minimize recording retention and access to raw speech.
Common Mistakes
- Do not grade voice behavior from the final transcript alone.
- Do not average a wrong-record renewal away with lower latency.
- Do not compare prompt variants across different audio-pipeline versions without recording the difference.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Evaluation sets: measure the failure cases that matter
- Paired prompt evaluation: count changes, then inspect uncertainty
- Audio prompts: mark overlapping speech before assigning speakers
- Voice prompts: do not commit an action from interim speech
- Voice interruptions: stop stale speech before handling a new turn
- Spoken confirmations: repeat the risky field, not the whole form
- Voice tool waits: report progress without inventing an outcome
- Project: a voice call that renews the right library item
- Real-time voice prompt decisions
