Answers calls. Remembers people. Runs in your own cloud.

On your Google Cloud platform, with Gemini intelligence. The recordings never leave it.

Hears tone, not just words Remembers who called Your account. Your data
Checked 7 August 2026 · prices and rankings move — open the links and re-check before you present
01 · The stack

What you are actually buying.

Google's platform is the base. Everything above it is the part that makes it answer your calls — and it is built once, not once per use case.

The solution A codified implementation layer
Codified, repeatable use cases
Collections Inbound servicing Outbound lead gen
Orchestration layer
Telephony, integrations & compliance workflowsWhatsApp · IVR · CRM · core systems
Indic language coverageHindi · Tamil · Telugu · Bengali · Kannada · Marathi · code-mixed
Telemetry, state management & routingTurn-taking · barge-in · noise suppression · fallbacks
Google platform
Agent platform on Vertex AIGemini Live · Gemini Flash · Chirp · TTS · STT
The red band is Google's, and it is the part with published prices and published benchmarks — sections 02 and 03. The layers above it are the build, and they are what section 04 prices. This is the slide that turns a model conversation into a platform conversation: the customer can see that swapping voice vendors changes only the bottom band, while everything that took the time sits above it.
02 · Quality

Tested by outsiders. Including where we lose.

Test Google ElevenLabs OpenAI Say this
Which voice sounds better (TTS arena Elo)
blind listening, Elo — higher is better · 92 models
Artificial Analysis
1210
Gemini 3.1 Flash TTS · 3rd
1171
Eleven v3 · 11th
1097
TTS-1 HD · 30th
39 Elo above ElevenLabs. Two models sit above Google — Qwen-Audio-3.0-TTS-Plus at 1229 and Speechify Simba 3.2 at 1227. Name them yourself before the customer opens the page and finds them.
Words it gets wrong (WER, word error rate)
real recordings, error rate — lower is better · 55 models
Artificial Analysis
2.8%
best Gemini
2.2% they win
Scribe v2 · 2nd
4.1%
Whisper Large v3
Concede it early — their product does only this job. Neither of us leads: the top spot is Fun-Realtime-ASR at 1.7%. Note the test now folds assistant-directed speech into the headline number, so the old "it's closer on agent speech" line no longer applies — drop it.
Voice quality (TTS)
human ratings, 40+ models
Hume · full table
Best of 40+ below Google
score not on this page
below Google
score not on this page
The largest human listening test run so far. If they ask for the exact positions, open the full table live — don't guess.
Understands how you said it (speech understanding)
tone, hesitation, mood
Hume · full table
Best of 40+ not entered below Google
score not on this page
The row that decides whether callers feel heard. A three-step build cannot fix this by trying harder.
Live back-and-forth (S2S, speech-to-speech)
human ratings, 40+ models
Hume · full table
below OpenAI
score not on this page
not entered Best they win Give it to them. No single model wins everything — that's the argument for a platform, not a product.
Copying a specific voice (voice cloning)
no public benchmark
capable Widely held best they win Reputation, not a measured test — say so. Matters for media. Doesn't decide a contact centre.
Indian languages live today
not a published test
Telugu, Bengali,
Kannada, Marathi,
mixed speech
ask them to demo it ask them to demo it From live deployments, not a leaderboard. Never cite it — play the recording. And never claim what the others can't do here: their published language lists are long, and the customer can check in a minute.
Every row links to the test it came from. Where a cell says score not on this page, the ranking is public but we haven't copied the number here — open the table and read it live rather than trust a figure retyped onto a slide. Two rows carry no published test at all and are flagged as such.
03 · Prices

Where every price comes from.

Every rate below is published by the company that charges it, and links to the page. Two things nobody publishes — the thinking model and the phone line — are marked as estimates, and you set them yourself in the next section.

Google — Gemini Live Google's price page
Listening
$0.005 / min
Google publishes this per-minute price itself
Speaking
$0.018 / min
Same page
Camera, if used
$0.002 / min
Live camera or screen, mid-call
Phone line
estimate
Not Google's. A phone call needs a carrier whichever model answers it — you set the rate in section 04
listening + (how much it talks × speaking) + your phone line

Cross-check on the same page: a minute is 1,500 units. 1,500 × $3.00/M = $0.0045 ≈ $0.005. 1,500 × $12.00/M = $0.018 exactly. Published prices and token maths agree.

OpenAI — gpt-realtime their price page, checked against one, two, three
Listening
≈ $0.019 / min
$32 per million units, converted
Speaking
≈ $0.077 / min
$64 per million. Where two conversions exist, we used the one cheaper for them
listening + (how much it talks × speaking) + your phone line This is the floor. Teams who measured their own bills report well above it.

Real bills run higher because the whole conversation is re-sent on every reply. The 6–11 cents is third-party billing data from 4,000 measured sessions. The calculator still uses the 5-cent floor.

ElevenLabs + a separate model their price page · cost review
Voice
$0.08 / min
Doubles to $0.16 over your call limit. Silence and hold time are charged too
Thinking model
estimate
Not included in any plan — their page says so. You pay Claude, or another model, separately
Phone line
estimate
Also separate, also stated on their page
their voice rate + your thinking model + your phone line Three meters, running on the same minute.

The structural point, in their words: the model and the phone line are not included. The $0.08 headline is one of three meters running on every call — and three contracts to sign.

Three companies bill every minute. One bills one. That is a structure, not a discount.
04 · The cost

Set your own numbers. Watch the three bills.

The published rates are fixed. The two nobody publishes are yours to set — and the phone line is charged to all three, because every one of them needs one.

$0
Saved a year against ElevenLabs plus a model and a phone line
Google published rate + your phone line$—
OpenAI published rate + your phone line$—
ElevenLabs published voice rate + your model + your phone line$—
Running costs only. No staff, build work or seats, on any of the three — and at most volumes those cost more than everything shown here. The ElevenLabs bar uses their normal rate, not their over-limit rate or the higher figures independent reviews of full setups report. The OpenAI bar uses the token floor, not the higher bills teams actually measured. Set the phone line to zero and the gap widens — leaving it in is the honest version, so leave it in.

And the whole business case, not just the minutes.

A minute is the easy number. This is the programme: what the automation frees up, what it costs to run, what it costs to build, and how long before it pays for itself. Every figure is an input — nothing here is a constant.

What the later agents reuse
Built once

System connections

The CRM, the billing system, the order database.

Built once

Memory

Collections already knows what support was told.

Built once

Safety rules

What it may say, and when it hands to a human.

Built once

Testing

The test suite and the release pipeline.

Demand & value
Build & replication
Model & running cost

Token prices default to published list rates for a fast Gemini text model — override with the account's contracted price. "Everything else" is a single stand-in for telephony, speech, infrastructure, monitoring and support, because we do not have those rates itemised yet.

Capacity released / year
Running cost / year
First build
Payback
05 · Nine questions

Ask every vendor. Including us.

QuestionA voice-only vendorGemini on Google Cloud
Which model thinks?TheirsGemini — and it can be swapped. Claude runs on the same platform
How is it built?Three steps. Tone lost in the middleOne model. Tone kept
What does a minute cost?Three meters: voice, thinking model, phone lineOne model rate, plus the same phone line. Section 04 does the maths with your numbers
Voice only?Yes. Chat and documents are a separate projectVoice, chat, documents, camera — one agent
Which languages, really?Ask for a live call in each one, not the supported-languages pageTelugu, Bengali, Kannada, Marathi and mixed-language speech, running in live deployments — ask us to demo it too
Does it remember callers?No. Every call starts blankYes — across calls and channels
Where do recordings live?Their cloudYour account, your region, or your own machines
What does agent two cost?The same as agent oneLess. The connections, memory and rules are already built
Where do you lose?Transcription, voice cloning, one conversation testWe lose those three. Section 02 says so
06 · See it live

Six things we can show you.

Coming soon

Try all three, on the same call, right here.

One call running through Google, OpenAI and ElevenLabs side by side in this page — with the cost ticking up on each of them as it goes. We are building the integration now.

D1

One call, three stacks

The same support call through all three, side by side, with the cost ticking up on screen.

Makes the price gap visible in one call.
D2

The annoyed caller

One sentence, said twice — calm, then irritated. Gemini changes tone. The three-step build answers the same both times.

Their architecture cannot fix this. Lead with it.
D3

Telugu and English, one sentence

No settings, no rerouting. The agent follows and answers the same way.

Running in live customer calls today.
D4

Point your camera at it

Mid-call, show the agent the broken screen. Same conversation, it sees it and keeps talking.

Neither competitor has an answer to this one.
D5

The returning caller

Mention a preference. Hang up. Call back tomorrow — it already knows.

A demo no-memory vendors cannot run on stage.
D6

A call that finishes the work

The agent hands the request to a second agent, which looks it up and completes the task — in your logs, live.

Turns a voice demo into a platform conversation.
07 · Seller notes

What this means for your account.

Usage

It lands in your account

Gemini runs in the customer's own project — the one already on your book.

Deal shape

A platform, not a pilot

Later agents reuse the first one's work, so there's a reason to fund something ongoing.

Regulated

The data question has an answer

"In your own account, your own region" — the question that kills most voice deals.

Three questions to open with.

01

How many languages do your customers call in — and how many can you answer in today?

02

What does your second use case cost — and who rebuilds the connections?

03

Where do your recordings sit — and has compliance signed it off?

The close
Anyone can win a demo.
These numbers hold up when you check them.