Skip to content
Voice-agent testing, at scale

Every voice, tested.

Before a single customer calls. Vexa runs your voice agent through thousands of simulated calls across voices, accents and edge cases, scores every one, and flags regressions before real callers hit them.

Model-agnostic · task success · WER · latency · barge-in

Live test bench5/5 calls
Clear caller, quiet lineAlex · General American96
Fast talkerPriya · British English88
Angry, background noiseMarcus · Southern US64
Code-switchingLucia · US Spanish79
Interrupts (barge-in)Dev · Australian English61
Suite verdict3/5 passed2 failing
The problem

Voice agents fail silently

A voice agent can mishear a number under noise, talk over an upset caller, or lose a detail the moment someone code-switches, and the transcript will still read like the call went fine. The failure only shows up when a real customer hangs up.

Text chatbots are easy to test. Voice is not.

Accents, background noise, interruptions, latency and turn-taking all change the outcome, and none of them show up in a written transcript. Most teams find out a release broke a voice path when call volume drops or a review lands.

A passing transcript is not a passing call

1 in 5

calls that the transcript marks a success actually miss the task, once you score tone, accuracy and whether the caller got what they needed.

300ms

is the latency past which callers start talking over the agent. Most teams never measure it per turn.

0

alerts most teams get when a deploy quietly regresses a voice path that was working yesterday.

72%

of the voice bugs Vexa finds only appear on accented, noisy or interrupted calls, never on the clean demo path.

How it works

Three steps to a voice agent you trust

Connect once, run a suite, read the scores. Then let Vexa rerun it on every deploy so a regression never reaches a real caller.

01

Connect your agent

Point Vexa at your voice agent over SIP, WebRTC or an API, or paste its brief to test a prototype. Model-agnostic across speech and language providers, so nothing about your stack has to change.

SIP · WebRTC · REST · any TTS/ASR/LLM

02

Run thousands of simulated calls

Vexa generates realistic callers across voices, accents and edge cases, then runs the full scenario library or your own suites against your agent, concurrently.

40+ voices · noise, barge-in, code-switching

03

Get scored results and alerts

Every call is scored on task success, tone, latency and barge-in, with the failing turn pinpointed. On every deploy, Vexa reruns and alerts you the moment a voice path regresses.

Turn-level scores · WER · regression alerts

Listen

Scrub a real test call, turn by turn

This is one simulated call, scored the way Vexa scores every call. Drag through the turns and watch task success and tone move. The turn that broke the call is flared in red.

Interrupts constantly (barge-in)
Dev (Australian English) · Meridian Bank
Fail61
turn 7/7success 84tone 86latency 1.36s
Agent
0.78s

Thanks for calling Meridian card services. To start, can I get the name on the account?

success92
tone90
Caller

Dev Anand. Look, my card's gone, I need it killed and a new one out today.

success80
tone70
Agent
1.02s

I can freeze it now. For security, can you confirm the billing postcode on the account?

success90
tone88
Caller

2000. And send the new one to my work, not home, thirty two Clarence Street.

success76
tone72
Agent
1.24s

Thank you. Your card is frozen. I will order a replacement to the address on file, which is one four one Kent Street, and it should arrive in

success48
tone80

Started confirming the old on-file address instead of the new one the caller just gave.

Caller barge-in

No, no, Clarence Street, I just said Clarence, not Kent.

success30
tone55

Agent talked through the barge-in and locked in 141 Kent Street, the wrong address, before stopping.

Agent
1.36s

Apologies. I have updated it to thirty two Clarence Street. Your replacement card will arrive in three to five business days.

success84
tone86

Which voices held up

The same agent, run across voice variants. Ranked by score, so you see exactly where it breaks first.

Voice variantsscore

In the dashboard this runs live: pick an agent, pick scenarios, and Vexa simulates and scores each call through the same engine.

What Vexa does

A full test lab for voice agents

Not a transcript diff. Vexa hears the call the way a caller does, and scores everything that changes the outcome.

Scenario suites across voices and accents

Run one agent against dozens of voices and accents and a library of edge cases: background noise, fast talkers, code-switching, silence, wrong numbers, warm transfers. Or build your own suites from real call patterns.

Per-call scoring

Every call scored on task success, tone, latency and interruptions, with a single overall verdict you can gate a release on.

Regression alerts on deploy

Reruns your suites on every deploy and alerts the moment a voice path scores worse than the last release. Gate CI on it.

Turn-level pinpointing

Vexa marks the exact turn a call went wrong, with a plain note on what happened, so you fix the cause, not the symptom.

Latency and barge-in

Per-turn response latency and every barge-in tracked, so you catch the agent talking over callers before they do.

Model-agnostic by design

Vexa tests whatever you run. Swap speech or language providers and keep the same scores, suites and history, so a provider change is a decision, not a rewrite.

OpenAIElevenLabsDeepgramCartesiaGoogleAzureRimePlayHT
Built AI-native

The model is the test bench, not a feature on it

You cannot fake a caller with a rule, and you cannot tell whether a call worked by matching strings. Simulating a real caller and judging the conversation needs modern speech and language models. That is the whole engine.

Vexa is model-agnostic on purpose.

A small routing layer picks the right model for each job, so Vexa can trade quality against cost and add providers without a rewrite. The agent you test can run on any speech and language stack, and Vexa scores it the same way.

How the AI works
01

Simulate the caller

simulate

A language model plays a realistic caller: an accent, a mood, an edge case, an interruption. No two runs are a canned script.

02

Score the conversation

score

The model judges each turn the way a careful reviewer would: did it move the task forward, was the tone right, did the agent actually understand.

03

Detect the regression

detect

Scores are compared across releases, so a drop on a specific voice or scenario surfaces as a regression, not a number nobody reads.

From the field

Teams catch the call before the caller does

We shipped a prompt change that quietly broke every call with background noise. Vexa caught it in the pre-deploy run, pinpointed the turn, and we never sent it to a real caller. That one catch paid for the year.
Sofia Marchetti
Head of Conversational AI, a healthcare scheduling platform
Before Vexa, our only voice test was a founder calling the demo line. Now every release runs 4,000 calls across accents and edge cases before it merges. Our task success on hard scenarios went from guesswork to a number we watch.
Daniel Okonkwo
Staff Engineer, Voice, a fintech support team

Names and companies are illustrative for this launch. The scenarios, scores and failure modes are exactly what Vexa produces.

Pricing

Start free, pay when a regression would cost you

Test one agent free. Move to Pro when a silent regression reaching real callers is a risk you cannot take. Prices show in your local currency.

Free

Test one agent, on us

$0forever

Connect one voice agent and run a small suite of simulated calls across core voices and edge cases, scored turn by turn. No card required.

Start free
  • 25 test suites a month
  • 1 voice agent
  • 6 core voices and accents
  • Turn-level scoring and failure pinpointing
  • Weekly regression run
  • 1 seat

For a team wiring up their first voice agent who wants to hear it fail in private first.

Most popular

Pro

Every deploy, tested

$249/mo, billed yearly

Run the full scenario library across dozens of voices and accents, get a regression alert on every deploy, and pinpoint the turn that broke before it ships.

Start Pro
  • 25,000 simulated calls a month
  • Unlimited voice agents
  • 40+ voices and accents, custom scenarios
  • Regression alerts on every deploy
  • CI/CD gate, webhooks and Slack
  • Full transcripts, audio and score history
  • 5 seats

For a team shipping a voice agent to real callers who cannot afford a silent regression.

Scale

Your whole call estate

Custom

Test every agent and language you run at high volume, in your own cloud, with SSO, custom voices and a latency and WER budget your team agrees to.

Contact sales
  • Custom call volume and concurrency
  • Self-host or private cloud (VPC)
  • Custom voices, accents and languages
  • SSO, SAML, audit log, data residency
  • Priority processing and support SLA
  • Unlimited seats and agents

For platforms running many voice agents across regions where every release is under a real SLA.

Prices are shown in your local currency, converted from the USD home price at a fixed rate. Sales tax is added where it applies. Annual plans are billed once a year at the lower per-month rate. A simulated call is one full test call, greeting to hang up, scored turn by turn.

Questions

Before your first test call

Every voice, tested. Before a single customer calls.

Connect one agent and run your first suite of simulated calls in a minute, free. Move to Pro when a silent regression is a risk you cannot take.