Skip to content

Product

Test every call before a customer makes one

Connect your voice agent, run thousands of simulated calls across accents and failure modes, score every turn on what actually matters, and gate your deploys on regressions. Before a single real caller is affected.

How it works

From agent endpoint to regression alert

Four steps from connection to confidence. No agents to instrument, no changes to your production stack.

The pipeline
01

Connect the agent

Point Vexa at your SIP endpoint or WebSocket URI. No SDK to install, no changes to your production stack. Vexa dials in as if it were a real caller, over the same telephony path your customers use.

02

Simulate thousands of calls

Vexa runs your suite across the full scenario library: 47 voices and accents, background noise conditions, fast talkers, code-switching, barge-in sequences, wrong numbers and warm transfers. Each call is a distinct persona and edge case.

03

Score every turn

Every agent response is scored in real time on task success, tone, latency and barge-in handling. Word error rate is estimated by comparing your STT output against what the simulated caller said.

04

Alert on regressions

Turn scores roll up to a suite score. If a new deploy drops task success below your threshold, Vexa blocks the merge and posts the failing turns directly to your Slack channel or PagerDuty routing key.

Works with any provider

Vexa connects via SIP and WebRTC. It does not care whether your agent runs on Twilio, Deepgram, ElevenLabs, Cartesia or a private model. If it can take a call, Vexa can test it.

Block a bad deploy before it ships

Set thresholds on task success, WER and latency. Vexa runs as a CI step and blocks the deploy if your suite score drops below threshold.

See the config
Scored call

What Vexa sees on every call

Every turn in every simulated call is scored in real time. This is what the output looks like for a booking-intent call that encounters a barge-in on turn three, and fails to recover.

Suite: booking-intent · Run #4,218 · gpt-4.1-mini · Twilio
2 failures2 / 4 turns passing
Turn 01Caller
Pass

“I’d like to book a table for two this Saturday, around seven in the evening.”

Task success94%
Tone91
Latency302 ms
Turn 02Agent
Pass

“Great, I can help with that. Can I confirm the date, this coming Saturday the 26th?”

Task success89%
Tone94
Latency268 ms
Turn 03Caller
barge-inFail

Yes, Satur, wait, does the 6:45 slot work? Or, I mean, seven is fine too.

Task success61%
Tone78
Latency730 ms
Turn 04Agent
recovery failedFail

“I’m sorry, I didn’t catch that. Could you repeat your request, please?”

Task success44%
WER28%
RecoveryFailed
Suite WER avg6.2%p50 latency342 msBarge-in recovery62% ↓Regression detected · 2 turns failed vs v1.4.2
What Vexa scores

Five dimensions, every turn

Vexa scores every agent response on what actually determines whether a call goes well, not just whether it completed.

Did the agent actually complete the caller's goal?

91.4%

suite avg

The most important signal. A model judge evaluates whether the caller's stated intent, booking a table, disputing a charge, completing a transfer, was satisfied by the end of the turn or call. Scored 0 to 100 per turn, averaged across the suite. Threshold-gated in CI.

0Threshold-gated in CI100

Professional, warm and on-script

The agent's language is rated for professionalism, warmth and adherence to persona guidelines. Catches agents that turn clipped or robotic under edge-case pressure.

88.2 / 100 suite avg

p50 and p95, per turn

Measured from the end of the caller's utterance to the first audio byte from the agent. Reported as p50 and p95 across all turns. Latency above roughly 700 ms drives abandonment in real caller populations.

342 ms

p50

614 ms

p95

Word error rate on your STT stack

Estimated by comparing what the simulated caller said against what your speech-to-text layer transcribed. Surfaces STT errors before they propagate to your NLU and cause task failures.

5.8% suite avg

Interruption handling, scored

Vexa injects barge-in events mid-utterance, as real callers do. The score measures whether the agent handled the interruption gracefully: stopping cleanly, not repeating itself, and recovering the task thread.

74% recovery rate

Failure pinpointing

Not just which call failed. Which turn.

Suite-level scores tell you something regressed. Turn-level scores tell you exactly where, and which dimension broke. Vexa drills to the utterance.

The regression lives in the turn, not the call

A call that fails on task success failed for a specific reason at a specific turn. An agent that handles the opening question well but collapses on a barge-in at turn four has a different bug than one that fails immediately. Vexa pinpoints the turn, the score drop and the dimension that broke.

Each failing turn shows the caller utterance, the agent response, and the per-dimension scores. You can filter the suite to show only turns below threshold and sort by the dimension that matters most.

Search by scenario or voice

Failures often cluster. A barge-in recovery bug will appear across every scenario that includes an interruption. A latency regression will show under callers with fast speech, because your STT takes longer to settle. Vexa groups failures by scenario type, voice, accent and edge-case flag so patterns surface immediately.

Every failing turn links to a full call replay: the audio waveform, the transcript, the STT output and the per-dimension scores side by side. No context switching.

Failure drill-down · booking-intent · v1.4.3 vs v1.4.2

barge-in-mid-utterance

en-IN-wavenet-F

−18 pp task success

14 affected turns

background-noise-coffeeshop

en-US-standard

−11 pp task success

9 affected turns

fast-talker-high-wpm

es-US-bilingual

+2.1 pp WER

6 affected turns

3 regression clusters · deploy blocked
Under the hood

The AI that runs the test

Honest about how the AI works

A language model role-plays each caller persona. A separate judge model scores every turn. A routing layer picks the right model for each task. Read the technical write-up, including known limits of the approach.

How the AI works

Every voice, tested.

Connect one agent and run your first suite of simulated calls in a minute, free.