On your Google Cloud platform, with Gemini intelligence. The recordings never leave it.
Google's platform is the base. Everything above it is the part that makes it answer your calls — and it is built once, not once per use case.
| Test | ElevenLabs | OpenAI | Say this | |
|---|---|---|---|---|
| Which voice sounds better (TTS arena Elo) blind listening, Elo — higher is better · 92 models Artificial Analysis |
1210 Gemini 3.1 Flash TTS · 3rd |
1171 Eleven v3 · 11th |
1097 TTS-1 HD · 30th |
39 Elo above ElevenLabs. Two models sit above Google — Qwen-Audio-3.0-TTS-Plus at 1229 and Speechify Simba 3.2 at 1227. Name them yourself before the customer opens the page and finds them. |
| Words it gets wrong (WER, word error rate) real recordings, error rate — lower is better · 55 models Artificial Analysis |
2.8% best Gemini |
2.2% they win Scribe v2 · 2nd |
4.1% Whisper Large v3 |
Concede it early — their product does only this job. Neither of us leads: the top spot is Fun-Realtime-ASR at 1.7%. Note the test now folds assistant-directed speech into the headline number, so the old "it's closer on agent speech" line no longer applies — drop it. |
| Voice quality (TTS) human ratings, 40+ models Hume · full table |
Best of 40+ | below Google score not on this page |
below Google score not on this page |
The largest human listening test run so far. If they ask for the exact positions, open the full table live — don't guess. |
| Understands how you said it (speech understanding) tone, hesitation, mood Hume · full table |
Best of 40+ | not entered | below Google score not on this page |
The row that decides whether callers feel heard. A three-step build cannot fix this by trying harder. |
| Live back-and-forth (S2S, speech-to-speech) human ratings, 40+ models Hume · full table |
below OpenAI score not on this page |
not entered | Best they win | Give it to them. No single model wins everything — that's the argument for a platform, not a product. |
| Copying a specific voice (voice cloning) no public benchmark |
capable | Widely held best they win | — | Reputation, not a measured test — say so. Matters for media. Doesn't decide a contact centre. |
| Indian languages live today not a published test |
Telugu, Bengali, Kannada, Marathi, mixed speech |
ask them to demo it | ask them to demo it | From live deployments, not a leaderboard. Never cite it — play the recording. And never claim what the others can't do here: their published language lists are long, and the customer can check in a minute. |
Every rate below is published by the company that charges it, and links to the page. Two things nobody publishes — the thinking model and the phone line — are marked as estimates, and you set them yourself in the next section.
Cross-check on the same page: a minute is 1,500 units. 1,500 × $3.00/M = $0.0045 ≈ $0.005. 1,500 × $12.00/M = $0.018 exactly. Published prices and token maths agree.
Real bills run higher because the whole conversation is re-sent on every reply. The 6–11 cents is third-party billing data from 4,000 measured sessions. The calculator still uses the 5-cent floor.
The structural point, in their words: the model and the phone line are not included. The $0.08 headline is one of three meters running on every call — and three contracts to sign.
The published rates are fixed. The two nobody publishes are yours to set — and the phone line is charged to all three, because every one of them needs one.
A minute is the easy number. This is the programme: what the automation frees up, what it costs to run, what it costs to build, and how long before it pays for itself. Every figure is an input — nothing here is a constant.
The CRM, the billing system, the order database.
Collections already knows what support was told.
What it may say, and when it hands to a human.
The test suite and the release pipeline.
Token prices default to published list rates for a fast Gemini text model — override with the account's contracted price. "Everything else" is a single stand-in for telephony, speech, infrastructure, monitoring and support, because we do not have those rates itemised yet.
| Question | A voice-only vendor | Gemini on Google Cloud |
|---|---|---|
| Which model thinks? | Theirs | Gemini — and it can be swapped. Claude runs on the same platform |
| How is it built? | Three steps. Tone lost in the middle | One model. Tone kept |
| What does a minute cost? | Three meters: voice, thinking model, phone line | One model rate, plus the same phone line. Section 04 does the maths with your numbers |
| Voice only? | Yes. Chat and documents are a separate project | Voice, chat, documents, camera — one agent |
| Which languages, really? | Ask for a live call in each one, not the supported-languages page | Telugu, Bengali, Kannada, Marathi and mixed-language speech, running in live deployments — ask us to demo it too |
| Does it remember callers? | No. Every call starts blank | Yes — across calls and channels |
| Where do recordings live? | Their cloud | Your account, your region, or your own machines |
| What does agent two cost? | The same as agent one | Less. The connections, memory and rules are already built |
| Where do you lose? | Transcription, voice cloning, one conversation test | We lose those three. Section 02 says so |
One call running through Google, OpenAI and ElevenLabs side by side in this page — with the cost ticking up on each of them as it goes. We are building the integration now.
The same support call through all three, side by side, with the cost ticking up on screen.
One sentence, said twice — calm, then irritated. Gemini changes tone. The three-step build answers the same both times.
No settings, no rerouting. The agent follows and answers the same way.
Mid-call, show the agent the broken screen. Same conversation, it sees it and keeps talking.
Mention a preference. Hang up. Call back tomorrow — it already knows.
The agent hands the request to a second agent, which looks it up and completes the task — in your logs, live.
Gemini runs in the customer's own project — the one already on your book.
Later agents reuse the first one's work, so there's a reason to fund something ongoing.
"In your own account, your own region" — the question that kills most voice deals.
How many languages do your customers call in — and how many can you answer in today?
What does your second use case cost — and who rebuilds the connections?
Where do your recordings sit — and has compliance signed it off?