Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact

Gemini 3.8 Flash TTS: Prices, Voices, Arena Rank

At a glanceQuick answers
What did Google release on September 23, 2026?
Two text-to-speech models in the Gemini API and Google AI Studio: Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) for voice design, acting and long-form audio, and Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) for high-volume audio and voice agents. Both take text and return audio.
What do they cost?
Through December 31, 2026: $0.50 per million text tokens in for both, and $9 (Flash TTS) or $6 (Flash-Lite TTS) per million audio tokens out, which is 1.35 or 0.9 cents per minute of speech. From January 1, 2027 every price doubles.
How good are they?
On Artificial Analysis’ blind provider voice arena, read September 23, Flash TTS ranks second (1260 Elo) and Flash-Lite TTS sixth (1235); the leader, Cartesia Sonic 3.6, is at 1273, within overlapping intervals. Google also reports the top spot on Hume AI’s Voice Design Benchmark.
Editorial data illustration on a near-white ground: a script page with laugh and mhm tags feeds a navy ribbon into a coral speaking head that emits four colored sound waves; headline ONE SCRIPT, ANY VOICE and three numbers, 130 languages, 1.35 cents a minute, a 30-second sample to replicate a voice
Fig 0One script in, many voices out, at a fraction of a cent a second. Figures from Google's September 23 pages.

Google’s new speech models are priced to be the default for anyone generating voice at volume: 1.35 cents a minute for the flagship and 0.9 for the fast one, until January. Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026 (the post’s own timestamp is 15:15 UTC; @GoogleAI’s thread followed at 15:26 UTC), pitching the pair as a move “from static presets into a dynamic creative studio”. The Gemini API changelog lists both as generally available under September 22.

This page is the record: what shipped, what a minute of speech costs, the four ways to get a voice, the consent rules, and where the two models land in an independent blind listening test. Every figure is from Google’s announcement, its speech generation guide, pricing page and changelog, or the Artificial Analysis voice arena, all read on release day.

On this page · 10 sectionsOpen
  1. What shipped
  2. What a minute of speech costs
  3. Four ways to get a voice
  4. Consent and watermarking
  5. Where the two models rank
  6. Where you can use them
  7. Where we sit
  8. What we are watching for
  9. Update log
  10. Sources
Key points6 · 9 min full read
  1. A sheet of script paper with two small bracket marks in the margin: the text a speech model reads.
    Google released two text-to-speech models on September 23, 2026: gemini-3.8-flash-tts, its flagship for voice design and long-form acting, and gemini-3.8-flash-lite-tts, built for high volume and voice agents. The Gemini API changelog marks both generally available under September 22.
  2. A price tag with a coin beside a short sound wave: the price of a minute of speech.
    Paid tier through December 31, 2026, per million tokens: $0.50 text in for both, $9 audio out for Flash TTS and $6 for Flash-Lite TTS. At Google’s 25 audio tokens per second that is 1.35 and 0.9 cents a minute of speech. Every price doubles on January 1, 2027; batch is half.
  3. A globe with three speech bubbles around it: many languages.
    Flash TTS speaks 130 languages and Flash-Lite TTS 101, and both detect the input language on their own. Google’s post highlights regional varieties such as Mexican Spanish, Quebec French and Scots English.
  4. A sound wave passing through a padlock with a check mark: a voice copied only with consent.
    Voice replication works from a 30-second sample, gated by a verbal consent recording from the voice owner that has to match the reference speaker. Every generated clip carries a SynthID watermark.
  5. A small podium with headphones resting on the second step: second place in a listening test.
    On Artificial Analysis’ blind provider voice arena, read September 23, Flash TTS is second at 1260 Elo behind Cartesia Sonic 3.6 at 1273, inside overlapping 95% intervals, and Flash-Lite TTS is sixth at 1235. The highest ElevenLabs entry is twelfth.
  6. Two speech bubbles facing each other: a two-person dialogue.
    Both models stage two-speaker dialogue from one script and take inline cues such as a laugh, a sigh or an mhm backchannel. The single two-speaker request works with prebuilt voices; designed or replicated voices are generated turn by turn and joined.

§ 01What shipped

Item Gemini 3.8 Flash TTS Gemini 3.8 Flash-Lite TTS
Model code gemini-3.8-flash-tts gemini-3.8-flash-lite-tts
Status Generally available, changelog dated September 22, 2026 Generally available, changelog dated September 22, 2026
Google’s positioning Flagship: voice design, acting, dialects, long-form stability High volume and low latency; replaces gemini-3.1-flash-tts-preview
Languages 130, detected automatically 101, detected automatically
Input and output Text in, audio out Text in, audio out
Two-speaker scripts Yes, with prebuilt voices Yes, with prebuilt voices
Consumer product Gemini Notebook Google Vids
Table 1The two models at release, per Google’s docs and changelog

The changelog’s line on the smaller model is the one that matters for anyone already on Google’s older speech preview: “Fast, cost-efficient TTS model built to replace gemini-3.1-flash-tts-preview for high-throughput production and real-time voice agent cascades”. The same release added a Voices endpoint (/v1beta/voices) for listing, designing and storing voices.

§ 02What a minute of speech costs

Line Flash TTS Flash-Lite TTS
Text in, through Dec 31, 2026 $0.50 $0.50
Audio out, through Dec 31, 2026 $9.00 $6.00
Text in, from Jan 1, 2027 $1.00 $1.00
Audio out, from Jan 1, 2027 $18.00 $12.00
Batch audio out, through Dec 31 $4.50 $3.00
Priority audio out, through Dec 31 $16.20 $10.80
One minute of speech, through Dec 31 1.35 cents 0.9 cents
Table 2Paid tier, per million tokens, Google’s pricing page, September 23, 2026

Google bills audio at 25 tokens per second, so a minute of speech is 1,500 output tokens. At the launch rates that is 1.35 cents a minute on Flash TTS and 0.9 cents on Flash-Lite TTS; from January 1, 2027 it is 2.7 and 1.8 cents. The preview Flash-Lite replaces, gemini-3.1-flash-tts-preview, is listed at $20 per million audio tokens, 3 cents a minute, so the replacement is 70% cheaper until January and 40% cheaper after. The free tier covers both models for testing.

Cents per minute of generated speech, Gemini API paid tierBar chart of cents per minute of speech: 3.1 Flash TTS preview at 3.0, 3.8 Flash TTS highlighted at 1.35 through December and 2.7 from January, 3.8 Flash-Lite TTS at 0.9 and 1.83.1 Flash TTS preview3.03.8 Flash TTS, to Dec 311.353.8 Flash TTS, from Jan 12.73.8 Flash-Lite TTS, to Dec 310.93.8 Flash-Lite TTS, from Jan 11.8Cents per minute of generated speech, Gemini API paid tierBar chart of cents per minute of speech: 3.1 Flash TTS preview at 3.0, 3.8 Flash TTS highlighted at 1.35 through December and 2.7 from January, 3.8 Flash-Lite TTS at 0.9 and 1.83.1 Flash TTS preview3.03.8 Flash TTS, to Dec 311.353.8 Flash TTS, from Jan 12.73.8 Flash-Lite TTS, to Dec 310.93.8 Flash-Lite TTS, from Jan 11.8
Fig 1Cents per minute of generated speech, Gemini API paid tier
Horizontal bar chart of Elo scores: Cartesia Sonic 3.6 1273, Gemini 3.8 Flash TTS 1260 highlighted in coral, Qwen-Audio-3.0-TTS-Plus 1259, Gemini 3.8 Flash-Lite TTS 1235, ElevenLabs Eleven v3 1167, axis starting at 1100

01The blind listening arena: second and sixth, with overlapping intervals at the top

Three price bars for a minute of speech: Gemini 3.1 Flash TTS preview 3.0 cents in grey, 3.8 Flash TTS 1.35 cents with a 2.7 cents from January tag, 3.8 Flash-Lite TTS 0.9 cents with a 1.8 cents tag

02A minute of speech: 1.35 and 0.9 cents until January, then double

Four panels: a row of microphones labeled prebuilt, 30 curated studio voices; a card drawer labeled extended library; a text box turning into a wave labeled voice design; a waveform with a padlock labeled replication, 30-second sample, consent check, watermark

03Four ways to get a voice, from 30 prebuilt to a 30-second replica

1 / 3
Fig 2The launch in three pictures

§ 03Four ways to get a voice

Option What it is Limits
Prebuilt 30 curated studio voices, such as Kore, Puck and Zephyr None stated
Extended Voice Library Hundreds of additional voices, listed through the Voices endpoint None stated
Voice design A new persona generated from a text description Stored voices: 200 per project, kept one year
Voice replication A copy of a real voice from a 30-second sample, with consent Stateless keys expire after seven days
Table 3Voice options in the Gemini API, per Google’s speech generation guide

Google’s own pages do not agree on the size of the catalog. The guide says 30 curated voices plus hundreds more, the changelog says 150+ prebuilt and custom voices, and the launch post says “Scale up from 30 original voices to an infinite library.” and counts 2,000+ production-ready voices. The honest reading: 30 curated voices you can name today, a larger library behind an API call, and voice design for anything else. Remixing a library voice is announced, not shipped: “Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent.”

Direction happens in the script. Both models take turn-level style notes and inline cues such as <laughs>, <sigh> and <gasp>, plus backchannels written in pipes such as |mhm| and |yeah|. A single request can stage two speakers with prebuilt voices; the guide says designed or replicated voices in a multi-character scene are generated one turn at a time and joined.

Replication is the feature most likely to be misused, and Google put the gate in the product: “users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created”. The sample itself is short, 30 seconds of your voice or one you have the rights to use. Separately, “every audio clip generated by our Gemini Audio models is watermarked with SynthID”, and replicated voices carry C2PA credentials.

§ 05Where the two models rank

Google’s launch numbers come from third-party benchmarks it reports on:

Claim Figure
Hume AI Voice Design Benchmark Flash TTS first overall, 71.4
Hume AI accent modeling Flash TTS leading, 60.8
Hume AI Overall Quality Index Flash TTS first, Flash-Lite TTS second
Voice Arena, blind human preference Top positions in Japanese, Brazilian Portuguese, Vietnamese, Arabic (MSA), Mexican Spanish and Hindi
Table 4Launch-day claims, Google’s announcement

The independent read we could check directly is Artificial Analysis’ provider voice arena, where listeners pick between two unlabeled clips and the votes become Elo scores:

Rank Model Elo 95% interval Listed price per 1M characters
1 Cartesia Sonic 3.6 1273 plus or minus 17 $49.0
2 Gemini 3.8 Flash TTS 1260 plus or minus 17 $16.5
3 Qwen-Audio-3.0-TTS-Plus 1259 plus or minus 17 $27.6
4 Inworld Realtime TTS-2 1245 plus or minus 18 $20.8
5 Speechify Simba 3.2 1237 plus or minus 14 $6.6
6 Gemini 3.8 Flash-Lite TTS 1235 plus or minus 16 $11.0
10 Gemini 3.1 Flash TTS 1199 plus or minus 12 $18.3
12 ElevenLabs v3 Conversational 1196 plus or minus 15 $50.0
17 ElevenLabs Eleven v3 1167 plus or minus 11 $100.0
Table 5Artificial Analysis provider voice arena, read September 23, 2026
Blind listening arena Elo, Artificial Analysis, September 23, 2026Bar chart of Elo scores: Cartesia Sonic 3.6 at 1273, Gemini 3.8 Flash TTS highlighted at 1260, Qwen-Audio-3.0-TTS-Plus 1259, Gemini 3.8 Flash-Lite TTS 1235, ElevenLabs v3 Conversational 1196, Eleven v3 1167Cartesia Sonic 3.61273Gemini 3.8 Flash TTS1260Qwen-Audio-3.0-TTS-Plus1259Gemini 3.8 Flash-Lite TTS1235ElevenLabs v3 Conversational1196ElevenLabs Eleven v31167Blind listening arena Elo, Artificial Analysis, September 23, 2026Bar chart of Elo scores: Cartesia Sonic 3.6 at 1273, Gemini 3.8 Flash TTS highlighted at 1260, Qwen-Audio-3.0-TTS-Plus 1259, Gemini 3.8 Flash-Lite TTS 1235, ElevenLabs v3 Conversational 1196, Eleven v3 1167Cartesia Sonic 3.61273Gemini 3.8 Flash TTS1260Qwen-Audio-3.0-TTS-Plus1259Gemini 3.8 Flash-Lite TTS1235ElevenLabs v3 Conversational1196ElevenLabs Eleven v31167
Fig 3Blind listening arena Elo, Artificial Analysis, September 23, 2026

Two cautions. The top three sit within each other’s intervals, so second place is a statistical tie with first and third, not a clear loss. And the arena scores a fixed set of voices and prompts; it does not measure voice design, replication or long-form stability, which is where Google says Flash TTS is strongest. What it does show cleanly is the price side: on the arena’s own per-character column, Flash TTS is a third of Cartesia’s listed price and a sixth of Eleven v3’s.

§ 06Where you can use them

Surface Flash TTS Flash-Lite TTS
Gemini API and Google AI Studio From September 23 From September 23
Gemini Enterprise Coming soon, via API Coming soon, via API
Consumer Gemini Notebook Google Vids
Table 6Availability, per Google’s announcement

Google names Agora, LiveKit, Pipecat and Vercel as developer platforms building on the API, and Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as companies integrating the models for dubbing, localization and voice agents. Vertex AI is not named in the announcement.

§ 07Where we sit

CellCog does not use these models as of this page. Our speech generation runs on OpenAI, ElevenLabs and MiniMax voices, and the multi-voice dialogue behind our podcast runs on ElevenLabs. A two-to-one price gap at arena-level quality is the kind of number that moves a routing decision; if ours moves, it ships as a product update on this blog, the way our Gemini 3.8 Flash switch did.

§ 08What we are watching for

  • Vertex AI and Gemini Enterprise availability for both models.
  • Voice remixing, announced as coming soon.
  • The January 1, 2027 price step, and whether it moves.
  • The arena with more votes: whether second place holds against Cartesia and Qwen.
  • A response from ElevenLabs on price or quality.

§ 09Update log

As of September 23, 2026, 21:30 ET: page opened.

§ 10Sources

Frequently asked5 questions

Q1Which one should a developer pick?

Google’s split: Flash TTS for designed voices, line-by-line acting, audiobooks, games and podcasts; Flash-Lite TTS for dubbing, bulk audio and voice agents where cost and latency matter. Flash-Lite is also the named replacement for gemini-3.1-flash-tts-preview, and at 0.9 cents a minute it is 70% cheaper than that preview until January.

Q2How many voices are there?

It depends which Google page you read. The docs list 30 curated prebuilt voices plus hundreds more in the Extended Voice Library; the changelog says 150+ prebuilt and custom voices; the launch post says 2,000+ production-ready voices. You can also design new ones from a text description or replicate one from a sample.

Q3Can I clone any voice?

No. Google says replication needs your own voice or one you have rights to use, and a verbal consent recording from the voice owner that matches the sample. Stored custom voices are capped at 200 per project and kept for a year; stateless replication keys expire after seven days.

Q4Where can I use them without writing code?

Flash TTS is in Gemini Notebook and Flash-Lite TTS in Google Vids, per Google’s post. Both are in Google AI Studio’s audio playground for developers, and Gemini Enterprise access is listed as coming soon.

Q5Does CellCog use Gemini TTS?

Not as of this page. CellCog speech generation runs on OpenAI, ElevenLabs and MiniMax voices, and our multi-voice podcast dialogue runs on ElevenLabs. If that changes, it ships as a product update on this blog.

Published 23 September 2026 All Choosing a platform →