AERIOXFLUX
Create & Earn
Create & Earn · voice tts

Google Priced Voice at $1.38 an Hour and Took the Top Spot

Gemini 3.8 Live runs $0.005 a minute in and $0.018 out, against roughly $3 an hour for GPT-Live-1. The Extended Thinking variant leads the Artificial Analysis speech-to-speech index at 82.6.

Flux Desk·2026-09-17·5 min read

On September 15, Google released two speech-to-speech models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Pricing is $0.005 per minute of audio input and $0.018 per minute of output — roughly $1.38 an hour for a typical conversational mix.

OpenAI's GPT-Live-1 runs about $0.05 per minute, roughly $3.00 an hour at minimum.

On the Artificial Analysis Speech-to-Speech Index, the Extended Thinking variant debuted at #1 with 82.6, and took the top slot on the Tau Voice benchmark at 68.6%. Standard Gemini 3.8 Live placed fifth at 76.0. The models support over 97 languages, can make API calls in the background, and process visual input while talking.

The number that changes the business case

Voice agents have been technically viable and economically marginal for two years. The math is simple and unforgiving: a support call that runs eight minutes costs about $0.40 in model inference at GPT-Live-1 pricing. Add telephony, orchestration, and observability, and you are near the point where a human in a low-cost market is competitive.

At $1.38 an hour, that same eight-minute call costs about $0.18. That is not an incremental improvement to a margin. It is the difference between a voice agent being a premium deployment and being the default for any volume support queue.

The categories that open up are the ones where call volume is high and value per call is low: appointment confirmation, order status, delivery windows, outbound reminders, qualification calls. Those were never worth $3 an hour of inference. They are worth $1.38.

What Google traded away to get there

Price is not free, and the reviewers who spent time with both models are consistent about where Google gave ground.

OpenAI's GPT-Live-1 uses full duplex — it listens and speaks simultaneously, the way people actually talk, which is what makes interruptions and overlaps feel natural rather than transactional. Gemini 3.8 Live is reported to sound less natural in exactly that dimension.

That is a real gap, and it maps cleanly onto use cases. If your voice agent is the primary interface for a consumer product, the quality of turn-taking is the product. If your voice agent confirms a dental appointment, nobody cares whether it handles overlapping speech gracefully.

The Extended Thinking variant complicates this further. It is the one that tops the leaderboard, and it costs meaningfully more — reported at around $3.50 an hour for input against $0.84 for the standard model. A voice agent that thinks before speaking scores better on benchmarks and introduces latency into a medium where latency is the thing users notice most. Whether you want a model that pauses to reason mid-conversation depends entirely on whether the conversation is a task or a chat.

The pattern this fits

This is the third time this cycle Google has entered a category at roughly half the incumbent's price and taken the benchmark lead in the same announcement.

The reason it can is TPUs. Google runs inference on silicon it designed, in data centers it owns, without paying Nvidia's margin. That advantage compounds in exactly the workloads where inference is continuous rather than bursty — and a voice session is the most continuous workload there is. Every second of a call is tokens.

OpenAI's structural answer has been its Broadcom custom inference program, which is real but not shipping at the volume that would close this gap today. In the interim, OpenAI competes on quality of experience, which is a defensible position for consumer voice and a weak one for the long tail of business automation.

The benchmark caveat worth stating

A #1 on the Artificial Analysis Speech-to-Speech Index at 82.6 is a meaningful result, and it is also an index score. Speech-to-speech quality is unusually hard to benchmark because the thing users judge — does this feel like talking to someone — decomposes badly into measurable components.

Tau Voice at 68.6% is arguably the more useful number, because it measures task completion in voice-mediated agentic scenarios rather than conversational quality in the abstract. A model that finishes the job it was called about is what a business is buying.

Both numbers came from the same week Yale researchers re-graded a physics benchmark and found that many frontier-model "errors" were benchmark errors — bad answer keys and underspecified questions. That is worth holding in mind whenever a leaderboard position is the headline.

What to watch

Whether OpenAI cuts. GPT-Live-1 at $3 an hour against a #1-ranked competitor at $1.38 is not a stable configuration. A price move from OpenAI within a quarter would be the normal outcome, and it would confirm that voice pricing is now set by whoever has the cheapest inference rather than the best model.

Whether full duplex arrives on Gemini. If Google closes the turn-taking gap at current pricing, the quality argument for GPT-Live-1 narrows to preference.

Volume shifting in the telephony layer. The honest indicator is not benchmarks — it is which model the voice-agent platforms default to. Those companies pay the inference bill and will move on price the moment quality is acceptable.

#gemini-live#voice-agents#speech-to-speech#openai#pricing

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.