Product
Test every call before a customer makes one
Connect your voice agent, run thousands of simulated calls across accents and failure modes, score every turn on what actually matters, and gate your deploys on regressions. Before a single real caller is affected.
From agent endpoint to regression alert
Four steps from connection to confidence. No agents to instrument, no changes to your production stack.
Connect the agent
Point Vexa at your SIP endpoint or WebSocket URI. No SDK to install, no changes to your production stack. Vexa dials in as if it were a real caller, over the same telephony path your customers use.
Simulate thousands of calls
Vexa runs your suite across the full scenario library: 47 voices and accents, background noise conditions, fast talkers, code-switching, barge-in sequences, wrong numbers and warm transfers. Each call is a distinct persona and edge case.
Score every turn
Every agent response is scored in real time on task success, tone, latency and barge-in handling. Word error rate is estimated by comparing your STT output against what the simulated caller said.
Alert on regressions
Turn scores roll up to a suite score. If a new deploy drops task success below your threshold, Vexa blocks the merge and posts the failing turns directly to your Slack channel or PagerDuty routing key.
Works with any provider
Vexa connects via SIP and WebRTC. It does not care whether your agent runs on Twilio, Deepgram, ElevenLabs, Cartesia or a private model. If it can take a call, Vexa can test it.
Block a bad deploy before it ships
Set thresholds on task success, WER and latency. Vexa runs as a CI step and blocks the deploy if your suite score drops below threshold.
See the configWhat Vexa sees on every call
Every turn in every simulated call is scored in real time. This is what the output looks like for a booking-intent call that encounters a barge-in on turn three, and fails to recover.
“I’d like to book a table for two this Saturday, around seven in the evening.”
“Great, I can help with that. Can I confirm the date, this coming Saturday the 26th?”
Yes, Satur, wait, does the 6:45 slot work? Or, I mean, seven is fine too.
“I’m sorry, I didn’t catch that. Could you repeat your request, please?”
Five dimensions, every turn
Vexa scores every agent response on what actually determines whether a call goes well, not just whether it completed.
Did the agent actually complete the caller's goal?
91.4%
suite avg
The most important signal. A model judge evaluates whether the caller's stated intent, booking a table, disputing a charge, completing a transfer, was satisfied by the end of the turn or call. Scored 0 to 100 per turn, averaged across the suite. Threshold-gated in CI.
Professional, warm and on-script
The agent's language is rated for professionalism, warmth and adherence to persona guidelines. Catches agents that turn clipped or robotic under edge-case pressure.
88.2 / 100 suite avg
p50 and p95, per turn
Measured from the end of the caller's utterance to the first audio byte from the agent. Reported as p50 and p95 across all turns. Latency above roughly 700 ms drives abandonment in real caller populations.
342 ms
p50
614 ms
p95
Word error rate on your STT stack
Estimated by comparing what the simulated caller said against what your speech-to-text layer transcribed. Surfaces STT errors before they propagate to your NLU and cause task failures.
5.8% suite avg
Interruption handling, scored
Vexa injects barge-in events mid-utterance, as real callers do. The score measures whether the agent handled the interruption gracefully: stopping cleanly, not repeating itself, and recovering the task thread.
74% recovery rate
Not just which call failed. Which turn.
Suite-level scores tell you something regressed. Turn-level scores tell you exactly where, and which dimension broke. Vexa drills to the utterance.
The regression lives in the turn, not the call
A call that fails on task success failed for a specific reason at a specific turn. An agent that handles the opening question well but collapses on a barge-in at turn four has a different bug than one that fails immediately. Vexa pinpoints the turn, the score drop and the dimension that broke.
Each failing turn shows the caller utterance, the agent response, and the per-dimension scores. You can filter the suite to show only turns below threshold and sort by the dimension that matters most.
Search by scenario or voice
Failures often cluster. A barge-in recovery bug will appear across every scenario that includes an interruption. A latency regression will show under callers with fast speech, because your STT takes longer to settle. Vexa groups failures by scenario type, voice, accent and edge-case flag so patterns surface immediately.
Every failing turn links to a full call replay: the audio waveform, the transcript, the STT output and the per-dimension scores side by side. No context switching.
Failure drill-down · booking-intent · v1.4.3 vs v1.4.2
barge-in-mid-utterance
en-IN-wavenet-F
14 affected turns
background-noise-coffeeshop
en-US-standard
9 affected turns
fast-talker-high-wpm
es-US-bilingual
6 affected turns
The AI that runs the test
Honest about how the AI works
A language model role-plays each caller persona. A separate judge model scores every turn. A routing layer picks the right model for each task. Read the technical write-up, including known limits of the approach.
Every voice, tested.
Connect one agent and run your first suite of simulated calls in a minute, free.