memujo
AI6 min read

Gemini 3.8 Live: Google's Voice AI Cost Breakdown

Google's Gemini 3.8 Live costs $0.005 per input minute and $0.018 per output minute, well under OpenAI's $0.05 per minute voice rate. Here is the per-hour math.

By Alice

In this article
  1. 01The facts at list price
  2. 02The math
  3. 03What the pitch leaves out
  4. 04Our Read

The real question about Google's Gemini 3.8 Live is not whether it is smart. Google already published the benchmarks, and the numbers are strong. The question that actually matters to anyone building a voice product is what a conversation costs, and whether the price gap against OpenAI's GPT-Live-1 is big enough to change a decision. It is, and the gap is wider than the headline suggests, but only in one direction of the conversation.

The facts at list price

On September 15, Google DeepMind released two live dialogue models in the Gemini API and Google AI Studio: Gemini 3.8 Live, built for scale and cost efficiency, and Gemini 3.8 Live Extended Thinking, built for high-complexity multi-step reasoning. Google's announcement describes them as its most advanced live dialogue models yet. Both are native speech-to-speech models, meaning they take audio in and produce audio out in a single model rather than chaining a speech-to-text step, a text model, and a text-to-speech step. The company positions them as a streamlined alternative to those cascaded pipelines.

The pricing, straight from Google's developer announcement, is $0.005 per minute for audio input and $0.018 per minute for audio output on the standard model. The Extended Thinking variant adds charges for reasoning tokens and for extra inputs such as video and documents.

That is the number to hold in your head, because the comparison is against a different pricing shape. OpenAI released GPT-Live-1 in its API on September 10, five days earlier, at $0.05 per minute for the front-end voice layer, with the backend model and tool usage billed separately, per OpenAI's launch post. So Google splits the bill into an input rate and an output rate, while OpenAI charges a flat per-minute rate for the voice layer on top of whatever the backend model costs.

The functional difference between the two is also worth stating plainly. Google's models run tool calls and API calls in the background while the conversation keeps flowing, and they support 97 or more languages with automatic language detection and mid-conversation switching. OpenAI's GPT-Live-1 is full duplex, which means it can listen and speak at the same time, and it delegates reasoning to a backend model such as GPT-6 Astra. Both architectures let the agent work while it talks. They get there by different routes.

The math

Let us work the numbers for a one hour session, because per-minute rates are easy to misread.

On input alone, the gap is stark. Google charges $0.005 per minute of audio in, OpenAI charges $0.05 per minute for the voice layer. That is a ten to one difference in the input direction. If a user speaks for 30 minutes in an hour, Google bills 30 times $0.005, which is $0.15, while OpenAI's voice layer bills the full hour at $0.05, which is $3.00.

But a real conversation has two directions. Assume a balanced hour where the user speaks 30 minutes and the model speaks 30 minutes. Google bills $0.15 for the input and 30 times $0.018, or $0.54, for the output, for a total of $0.69. OpenAI's voice layer bills the full 60 minutes at $0.05, or $3.00. So on a balanced hour, Google costs about $0.69 and OpenAI's voice layer costs $3.00, roughly a four to one difference.

The worst case for Google is a long, talkative model. If the model speaks for the full 60 minutes and the user also speaks 60 minutes, Google bills $0.30 plus $1.08, or $1.38, while OpenAI still bills $3.00. Even then, Google is about half the price. The output rate of $0.018 per minute is where Google's advantage narrows, because it is 3.6 times the input rate, and voice agents that talk a lot spend more time in the output direction.

Now scale it to a call center, which is where voice AI economics actually get decided. Take 10,000 calls a month, four minutes each, with the user speaking two minutes and the agent speaking two minutes. Per call, Google costs 2 times $0.005 plus 2 times $0.018, or $0.046. OpenAI's voice layer costs 4 times $0.05, or $0.20. Over the month, Google bills $460 and OpenAI's voice layer bills $2,000, a difference of $1,540 per month before a single backend token is counted. At that volume the choice of voice layer is a real line item, not a rounding error.

What the pitch leaves out

There are three caveats that the headline price does not show.

First, the comparison is not perfectly even. OpenAI's $0.05 per minute covers only the front-end voice layer. The backend model, the one doing the reasoning and tool calls, is billed separately on its own token rates. Google's $0.005 and $0.018 rates are for the standard model, and the Extended Thinking variant adds reasoning tokens and extra input charges on top. So a fair total cost of ownership has to add a backend to both sides, and the gap narrows once you do. Google's price is genuinely lower, but it is lower on the voice layer, not necessarily on the whole agent.

Second, the benchmark story is close, not one sided. Google reports that Gemini 3.8 Live Extended Thinking takes the top spot on the Artificial Analysis Speech to Speech Quality Index at 82.6, ahead of GPT-Live-1 Astra at 81.5 and Grok Voice Think Fast 2.0 at 81.3. On agentic task completion it scores 68.6 percent on tau-Voice and 35.1 percent on Sierra's tau-Voice-banking, and 97.7 percent on Big Bench Audio. Those are Google's reported figures. OpenAI's own numbers for GPT-Live-1 are also strong, with 86.2 percent on Tau3 Voice when paired with GPT-6 Astra at medium effort, and 97.3 percent on conversational dynamics. On that last metric in particular, OpenAI's model leads. The two companies are measuring different things, and each wins on its own scoreboard.

Third, there is a quality tradeoff that reviewers have flagged. The Decoder, which covered the launch, noted that OpenAI's model should still deliver more natural conversations thanks to full duplex, and that judging from the demos it also sounds better, suggesting Google once again optimized for price over quality. That is a fair characterization of the tradeoff, and it is the kind of thing a demo cannot fully settle. For a customer support line where the user just needs an answer, the price may dominate. For a product where the voice is the brand, the feel of the conversation may matter more than the invoice.

Our Read

Google's play here is the same one it has run across the Gemini line: match the frontier on capability, then undercut on price. The 3.8 Flash release earlier this month set the tone, and 3.8 Live extends it into the voice layer, where the unit economics of a conversation are unusually sensitive to the per-minute rate.

The honest read is that Google has won the cost war on the voice layer, and by a wide margin. Ten to one on input, roughly four to one on a balanced hour, and half the price even in the worst case. For high volume, low margin voice work, call triage, booking, spoken search, field service, the math is decisive.

But the voice layer is not the whole agent. OpenAI's flat rate sits on top of a backend that you pay for separately, and Google's Extended Thinking model adds its own reasoning charges. The total cost of a production voice agent is the sum of the voice layer plus the reasoning model plus the tools, and on that total the gap is real but smaller than the $0.005 versus $0.05 headline implies.

So the decision is not Google or OpenAI. It is which layer you are optimizing. If you are building a high volume voice product and the conversation quality is good enough, Google's lower voice layer price is a genuine cost advantage that compounds with scale. If the voice is the product, and the feel of the conversation is the differentiator, OpenAI's full duplex model may still be worth the premium. The numbers in this article are the starting point for that call, and they are all traceable to the two companies' own announcements.

See also: Local AI vs API: The Real Cost Math for 2026 and Google's Gemini 3.8 Flash and Flash Cyber Release.

  • #gemini
  • #voice-ai
  • #api-pricing
  • #google-deepmind
  • #speech-to-speech

Sources

Share this story