Voice AI lives or dies on the call: what it said, what it understood, and whether the task got done. Voice AI evaluation grades transcripts of real calls plus task completion, from recordings you provide.
We grade outputs. We evaluate what your voice system said and did, from call recordings and transcripts you provide. We do not train models, run fine-tuning, or sell experts we cannot verify.
Voice AI evaluation is the grading of a voice agent's calls on what the system said and whether the caller's task got completed. Voice product teams need it because voice has failure modes text never sees: misheard names, talked-over callers, wrong transfers, and confirmations the caller never actually gave. Good looks like accurate understanding of what the caller said, correct task completion, natural turn-taking, and clean handoffs to humans when needed. We grade from your recordings and transcripts: what the system said, and what it did.
Voice AI fails in the space between hearing and understanding. We grade for the patterns that lose callers.
Speech recognition gets the words slightly wrong and the agent acts on the wrong version with total confidence. The caller said "cancel," the system heard "confirm." We grade transcripts against the call outcome to catch understanding failures, not just transcription typos.
The voice agent says it booked, scheduled, or updated, but the task never completed, or completed wrong. In voice there is no screen to double-check. We grade task completion separately from what the agent claimed on the call.
Interrupting, answering before the caller finishes, or going silent at the wrong moment. It reads as rudeness even when the information is right. We grade conversational behavior from the transcript: interruptions, dead air, and recovery.
When the voice agent gives up, the handoff to a human has to carry the caller's story with it. Agents that transfer without context, or refuse to transfer at all, fail the caller's actual goal. We grade whether escalation happened at the right moment with the right context.
You send 20 to 50 call recordings with transcripts, plus your task definitions and quality rules, or we help you write a rubric from them. Everything stays confidential.
Trained reviewers grade each call on what the system said and whether the task got done, backed by automated checks. Fabricated confirmations are P0s.
Call-level grades, understanding vs completion splits, the failure patterns across the batch, recommended fixes, and a walkthrough call.
A call can sound perfect while the task never completes, and a call can complete the task while sounding awkward. We grade both separately so the report tells you whether to fix comprehension, execution, or conversational behavior. Lumping them into one score hides exactly the trade-off you need to see. Our human evaluation process handles the judgment calls that automated scoring misses.
Interruptions, dead air, talking over the caller, and recovery after a misunderstanding all show up in the transcript. Reviewers mark them as behavior notes with timestamps, so your team can find exactly where the call went wrong. Audio is checked where it matters: tone, pauses, and anything the transcript marks unclear.
We also grade recovery, because calls go wrong and what matters is what the system does next. A misheard name that gets corrected gracefully is a minor note. A misheard name that gets acted on without checking is a failure. The report separates the two, because recovery behavior is what separates a usable voice agent from a frustrating one. Good recovery follows a pattern: notice the confusion, confirm before acting, and keep the caller's goal intact. We grade each recovery attempt against that pattern, so your team knows whether the agent fails gracefully or fails twice.
Two things, separately: what the system said, graded from the transcript, and what it did, graded from the call outcome. A call can sound perfect and still fail if the booking never happened. It can also complete the task while sounding robotic, which loses points on behavior but passes on completion.
Reviewers work from transcripts for grading accuracy, and check audio where it matters: interruptions, long pauses, tone problems, and anything the transcript marks as unclear. If you can only provide transcripts, we can still grade, but audio makes behavior grading much stronger.
Recordings stay confidential and are used only for your evaluation. We recommend redacting payment details and other sensitive identifiers before sending, and we will confirm handling requirements with you before the pilot starts.
Yes. Outbound appointment setters, reminders, and follow-up calls get graded on the same structure: what was said, whether the goal was reached, and how the conversation behaved. Compliance with your calling rules becomes part of the rubric.
Yes, in the same pilot. Text sessions get graded like a support AI evaluation, voice calls get the full voice treatment, and the report shows both. Many voice products are really both, and grading them together catches inconsistencies between the two.
A standard pilot uses 20 to 50 calls covering your main call types: the happy path, the edge cases, and the calls that usually end in escalation. If you have distinct call flows, include samples of each so no flow goes ungraded.
Send 20 to 50 recordings. We will grade what was said and what got done.
Not sure where to start? Talk to us and we will point you at the right evaluation.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.