AERIOXFLUX
← Create & Earn
Create & Earn · voice tts

Google and Alibaba Just Started a Voice Price War

On the same day, Google shipped two Gemini 3.8 speech models at promotional prices and Alibaba cut its Qwen audio APIs by up to 95%. Synthetic voice is turning into a commodity line item.

Flux Desk·2026-09-24·5 min read

For most of the last two years, the expensive part of a voice product was the voice. Transcription got cheap early. Language models got cheap fast. Good speech output, the part a listener actually hears, stayed stubbornly priced.

On September 23, two of the largest model makers moved on that line in the same news cycle. Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Alibaba's Qwen team released Qwen-Audio-3.1, a five-model audio stack, and cut its speech prices by between roughly 70% and 95%. Neither announcement is a breakthrough on its own. Together they say something plainer: speech is becoming a commodity input, and the vendors are now competing on price per token.

What Google shipped

Google's post, by Leland Rechis and Alan Cowen of the Gemini audio team, splits the job in two. Flash TTS is pitched at creative work such as audiobooks, podcasts and game characters, with line-by-line direction for acting cues, pacing, dialect and non-verbal sounds like laughs and sighs. Flash-Lite TTS is the bulk option, aimed at dubbing, read-aloud features and voice agents where volume matters more than performance.

The headline feature is voice design by prompt. Instead of choosing from a fixed menu, a developer describes the voice in plain language. Google says the preset library has grown from 30 baseline voices to more than 2,000 production-ready ones, across more than 100 languages and dialects, including Mexican Spanish, Quebec French and Scots English. Both models also support two-speaker scenes natively.

Voice replication needs a 30-second sample plus a consent recording. All output carries a SynthID watermark, and Google says it supports C2PA credentials. Replication is not offered in Illinois, Texas, the EEA, the UK, Switzerland or India, which tells you where Google expects legal trouble over cloned voices.

On quality, Google cites Hume AI's Voice Design Benchmark, where Flash TTS scored 71.4, first overall, and 60.8 on accent modeling, also first. That is a vendor-reported ranking on a third party's benchmark. It has not been independently reproduced here.

The price, and the catch

The number that matters is on Google's pricing page. Flash TTS costs $0.50 per million input tokens and $9 per million audio output tokens. Flash-Lite TTS costs $0.50 in and $6 out. Batch mode halves both.

Those are promotional rates. The page says they hold through December 31, 2026. From January 1, 2027, both double: Flash TTS goes to $18 per million output tokens and Flash-Lite to $12. For comparison, Gemini 3.1 Flash TTS Preview lists at $20 per million output tokens, and Gemini 2.5 Flash Preview TTS at $10.

So the durable price cut is smaller than the launch price suggests. After January, Flash TTS is about 10% cheaper than the 3.1 preview on output, and Flash-Lite is 40% cheaper. The launch rate is a three-month incentive to move workloads now. Anyone budgeting a 2027 product should use the January numbers.

What Alibaba shipped

Qwen's release is broader. Qwen-Audio-3.1 upgrades the existing ASR, TTS and Realtime models and adds two new ones: TTS-Next, which generates speech, sound effects and background audio from one script in a single pass, and ASR-Next, which adds speaker labels, timestamps and detection of emotion and ambient sound. ASR-Next's API was not yet available at launch, according to 36Kr.

The price cuts are the story. Alibaba says speech recognition is down by up to 95%, TTS by about 70% and Realtime by about 85%. On Qwen Cloud, the international model page lists Qwen-Audio-3.1-Realtime-Plus at $6.40 per million audio input tokens, $0.80 per million text input tokens and $24 per million audio output tokens. It has a 262K-token context window, full-duplex audio, function calling and web search. 36Kr reports the domestic rates in yuan, including TTS-Flash at 1.5 yuan per million input tokens and 12 yuan per million output tokens.

Treat the percentage cuts with care. They are Alibaba's figures, and one independent pricing analysis from OrcaRouter notes that rates differ by region and tier, so the headline 70% will not map cleanly onto every bill. Third-party resellers also list different numbers. EmpirioLabs, for instance, lists Realtime-Plus audio input at $12.80 per million tokens, double the Qwen Cloud price.

Qwen's own benchmark claims are also vendor-reported. The most interesting one is practical: on Full-Duplex-Bench v1.5, Qwen says the Realtime model's rate of responding to irrelevant background speech fell from 73.0% to 13.0%. For anyone who has had a voice agent answer the television, that number matters more than a leaderboard score.

Why this matters for builders

The two launches aim at different layers. Google is selling expressive, directable output for content, with a large voice library and strong safety labeling. Alibaba is selling the whole loop, from hearing to talking, and pricing the realtime layer aggressively.

For a creator running a faceless channel or an audiobook pipeline, the practical question is no longer whether synthetic narration is affordable. At $6 to $9 per million output tokens, cost is rarely what stops a project. The questions now are control, consistency across long sessions and licensing. Google's per-line direction and prompt-built voices target exactly those problems.

For teams building voice agents, the calculation is about the full turn. Realtime models that listen and speak at once cost more than a cascade of separate transcription, language and speech models, and Qwen's cut narrows that gap. The Full-Duplex-Bench result also points to where competition is heading: less about how natural the voice sounds and more about whether the agent knows when to stay quiet.

The pressure on specialist voice companies is obvious. When two platform vendors that also sell the language model underneath cut speech prices in the same week, a standalone TTS vendor has to win on quality, voice rights or workflow, not price. The watermark and consent features Google shipped suggest the next fight is over trust.

Two things to watch. First, whether Google's January price doubling holds or quietly slips once usage is locked in. Second, whether independent evaluations confirm the benchmark claims both companies made. Until then, the only numbers you can plan around are the ones on the pricing pages, and one of those pages has an expiry date.

#text-to-speech#gemini#qwen#voice-agents#pricing

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.