
Two years ago, about 20 people a month typed "gemini tts" into Google. As of August 2026 it is 1,300, a 55x lift, and that reading predates the thing everyone is now searching for. On 23 September 2026 Google shipped Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models that let you write a voice into existence with a sentence of plain English and then direct it line by line.
Search the term today and every result on page one belongs to Google: AI Studio, the Cloud docs, DeepMind, the Gemini API reference. Nothing editorial ranks. So nobody has put the one number next to it that decides whether any of this matters to a business: what a minute of generated speech costs. Google's own pricing page has the answer buried in a footnote about token rates. Converted, Gemini 3.8 Flash TTS runs $0.81 per hour of audio, and the cheaper of the two models runs $0.54. We did the conversion, and we have the demand curve nobody else is holding.
Key takeaways:
- "gemini tts" grew from 20 to 1,300 US monthly searches in two years, up 120% year over year, classified as EXPONENTIAL in the Rising Trends database (data as of August 2026).
- Google shipped two models on 23 September 2026: Flash TTS for studio work and Flash-Lite TTS for volume. Both do voice design from a text prompt and voice replication from a 30-second sample.
- At Google's promotional rate, generated speech costs $0.81 an hour on Flash TTS and $0.54 on Flash-Lite. Both prices double on 1 January 2027.
- That undercuts the model it replaces. Gemini 3.1 Flash TTS Preview was $1.80 an hour, and Google's own Chirp 3: HD voices are still billed per character at $30 per million.
- The demand has moved anyway. "ai voice agents" is up 10,479% in a year to 201,000 searches a month, while "elevenlabs" grew 22% (data as of August 2026). Google is pricing the layer underneath the part of the market that is growing.
- Every open-source TTS model term we track is down year over year. The only other model name matching Gemini's volume is qwen tts, at 1,300 and up 306%.
Let's get into it.
The search spike, in numbers
Here is the monthly search volume for "gemini tts" over the last two years, straight from our database.

The shape matters more than the height. This is not a launch spike. The breakout was May 2025, when the term first crossed a quarter of its eventual peak, and it has been a staircase since: 590 a month in September 2025, 1,300 by December, a peak of 1,600 in April 2026, and 1,300 again in August. A term that holds a four-figure base for nine months before the flagship product arrives is a term with real intent behind it. People were already using the 2.5 and 3.1 preview models in production.
Now put it next to the other model names in the category.

Two things jump out. Gemini and Alibaba's Qwen are tied at the top, and everything else is a rounding error. And the open-source side is shrinking: index tts is down 33% in a year, orpheus tts down 46%, kitten tts down 89%. Two years ago the self-hosted models were where the curiosity was. Now the searches follow the labs.
What Gemini TTS actually is
It is an API, not an app. You send text, you get audio back, and the interesting part is everything you can attach to the text on the way.

The API guide splits control into two layers. Sustained direction goes into a speech_metadata block on each turn: a style string like "whispering" or "out of breath", plus the speaker label. Momentary events stay inline in the transcript as angle-bracket tags, so <laugh>, <sigh> and <short pause> land where you put them. That split is new in 3.8. In the preview models you stuffed stage directions into the text and hoped.
Three capabilities sit on top of that. Voice design builds a voice from a written description, which Google demonstrates with a high-energy Melbourne DJ and a Japanese dragon. Voice replication copies a real voice from a 30-second sample, gated behind a recorded consent clip from the voice owner and watermarked with SynthID plus C2PA credentials. And saved voices persist a designed voice so it does not drift between sessions, capped at 200 per project with a one-year expiry.
The limits are worth knowing before you plan around them. A single request handles at most two speakers, and only with prebuilt voices, so a three-character scene means synthesising each turn separately and stitching. Flash TTS covers over 130 languages and Flash-Lite over 100. Voice replication is switched off in Illinois, Texas, the EEA, the UK, Switzerland and India.
Why people are searching for it
Because the thing that used to be a paid feature is now a prompt. Until this release, giving a synthetic voice a specific character meant either picking from a fixed menu or paying for a cloned one. Google's pitch is that you can now write the character. Its announcement claims the top spot on Hume AI's Real World VoiceEQ benchmark for voice design, at 71.4 overall and 60.8 on accent modelling, with the two 3.8 models taking first and second on Hume's overall quality index.
Treat that carefully. The benchmark is Hume's, and the paper behind it argues the opposite of a single-number verdict: voice AI, the authors write, "should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score." Google is quoting one row of a table its own model tops.
The developer reaction started within days. This review, published 25 September, walks through voice design, cloning and emotion control two days after launch.
What an hour of speech costs
This is the part nobody ranking for the term has done. Google prices TTS output in audio tokens, and a footnote on every table in the Gemini API pricing page says audio tokens run at 25 per second. That makes one million audio tokens equal to 40,000 seconds, or 11.11 hours of speech. Divide through and the per-million figures turn into something a finance team can read.
| Model | Per 1M audio tokens | Per hour of speech | From 1 Jan 2027 |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $9.00 | $0.81 | $1.62 |
| Gemini 3.8 Flash-Lite TTS | $6.00 | $0.54 | $1.08 |
| Gemini 3.8 Flash-Lite TTS, batch | $3.00 | $0.27 | $0.54 |
| Gemini 3.1 Flash TTS Preview | $20.00 | $1.80 | $1.80 |
Read the last column first. The $9.00 and $6.00 rates are promotional and expire on 31 December 2026, which Google states plainly on the same page. Anyone modelling a voice product on these numbers is modelling a price that has a date on it.
Even so, the comparison holds. Gemini 3.8 Flash TTS is 55% cheaper than the preview model it replaces, and still 10% cheaper after the doubling. Against Google's own alternatives the gap is wider: Cloud Text-to-Speech still bills Chirp 3: HD voices at $30 per million characters and instant custom voice at $60, and its price list has no 3.8 row at all. The new models landed in the Gemini API and AI Studio first, with Gemini Enterprise marked coming soon.
Against ElevenLabs, the incumbent, the cheapest per-minute number on its pricing page is "low-latency TTS as low as 5c/minute", which is $3.00 an hour, and it sits behind the $990-a-month Business plan. Google's $0.54 has no seat count and no monthly commitment. That is not a small difference in a category where the unit is hours of audio.
What the early reaction says
The complaints on the preview models tell you what Google was aiming at. A developer on r/TextToSpeech building with the 3.1 preview wrote, "Hey, I am playing around with the new flash TTS preview and it seems very slow" (thread). Another thread on r/GeminiAI is titled, flatly, "Gemini TTS Preview: Great quality, terrible latency". A third asked whether anyone had "achieved consistent voice identity" while researching the preview model "as the primary TTS engine for a long-form audiobook" application (thread).
Latency and drift. Those are exactly the two things the 3.8 release notes address: Flash-Lite exists for speed, and saved custom voices plus long-form generation exist for stability. Whether the fix holds at scale is the open question, and it will show up in those same subreddits before it shows up in a benchmark.
Google's own partner list is the other tell. Agora, LiveKit, Pipecat and Vercel are named as the platforms shipping Gemini TTS integrations. None of them is a creative tool. All four are voice-agent infrastructure.
What it means for the voice market
The most useful thing our data says about this launch is that Google is not entering the market people think it is entering.

ElevenLabs is still the biggest name in synthetic voice by a distance, at 368,000 searches a month. But it grew 22% in a year. AI voice agents grew 10,479% to 201,000, and realtime voice ai grew 1,660%. The money and the attention have moved from "make me a voice file" to "answer my phone", and a TTS model is a component of the second thing, not a competitor to it.
That reframes the pricing. Google is not trying to win the creator who needs one narrator. It is trying to be the default speech layer inside every agent someone builds on the Gemini API, which is the same play we described in our 2026 AI agent report. Cheap components win platform fights. The people who should be nervous are not the voice studios but the middleware vendors whose margin was the gap between raw TTS and a working agent.
For everyone buying rather than building, three things change. Per-seat voice pricing gets harder to defend when the underlying model is under a dollar an hour. The customer experience teams running contact-centre pilots can now model voice as a variable cost instead of a licence. And the brands treating a distinctive voice as an asset just got a reason to lock it down, because replication now takes 30 seconds of audio and a consent recording.
What to watch next
The January price. Both promotional rates expire on 31 December 2026. If Google lets the doubling happen, Flash-Lite lands at $1.08 an hour and the gap to ElevenLabs narrows from roughly six times to under three. If it quietly extends the promotion, that tells you Google cares more about the agent platform than the audio margin.
Gemini Enterprise and Cloud. As of today the 3.8 models are Gemini API and AI Studio only. Cloud Text-to-Speech still lists the 2.5 and 3.1 models on its price sheet. Enterprise availability is the signal that Google thinks the safety tooling is ready for regulated buyers.
The search terms. Watch gemini tts against qwen tts. Alibaba's model is the only other name at the same volume, and it is growing faster off a smaller base at 306% a year. If Gemini's curve steps up through the autumn while Qwen's flattens, the launch worked. If ai voice agents keeps compounding while both model terms stay flat, the interesting layer was never the model, which is the pattern we keep seeing across generative AI and the advertising businesses built on top of it.
The honest summary is that Gemini 3.8 TTS is a very good component sold at a price designed to make you stop shopping. The voice is not the product. The agent that uses it is, and Google has just made the raw material cheap enough that the agent is where the argument moves next.
Want to spot the next breakout product before the launch post lands? Read our guide on how to identify market trends, follow the live gemini tts trend page, or browse what is breaking out right now on the Rising Trends dashboard.



