Cover: AI-generated editorial composition by TMRW. The Hume connection and the independent arena ranking were first reported by OrcaRouter; product details are from Google's announcement and API documentation.
Google released two text-to-speech models on 23 September, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The headline feature is voice design: you describe a voice in a sentence and get a reusable one back. It works, and it will change how small teams cast audiobooks, games and support lines. Before you repeat Google's claim that it is the best on the market, though, look at where that number comes from.
A voice from a sentence
Until now, most text-to-speech meant picking from a list or cloning someone's recording. With voice design, you write something like "a crisp, energetic sports announcer in her 30s with a slight Midwestern accent," and the API returns a stored voice ID plus a sample clip to audition. Google's demos include a high-energy DJ from Melbourne and a Japanese dragon.
Once the voice exists, you direct it line by line: "whispered urgently," "cheerful and energetic." The models also stage two-speaker scenes from one script and take cues for laughs, sighs and listener noises like "mhm." Google says both hold a voice steady across hours of audio, which is the usual failure point for audiobooks.
The rest of the kit:
- A library of 2,000+ ready voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.
- Voice replication from a 30-second sample. The voice's owner must record a spoken consent that matches the sample. Replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland or India.
- Watermarks on every clip. Output carries Google's SynthID watermark, and replicated voices also get C2PA content credentials.
- Voice remixing, which adjusts an existing voice's pitch or accent, is announced as coming soon. It is not available yet.
Check whose benchmark it is
Google says Flash TTS took first place on Hume AI's Voice Design Benchmark with 71.4, and that the two models rank first and second on Hume's Overall Quality Index. Hume is a respected voice lab, and the benchmark is real.
But one of the two named authors of Google's announcement is Alan Cowen, Hume's founder. Wired reported in January that Google DeepMind hired Cowen and about seven Hume engineers as part of a licensing deal. Hume still operates independently. That does not make the score wrong. It does mean the score is not a neutral third party grading Google, and a reader should know that.
The independent figure is less flattering and more useful. On Artificial Analysis's Speech Arena, where listeners vote blind between clips, OrcaRouter's 24 September snapshot put Gemini 3.8 Flash TTS second, behind Cartesia Sonic 3.6, and Flash-Lite sixth. Second place in a blind vote is a strong result. It is not "number one."
Where you can use it
Developers can try both models now in the Gemini API and in Google AI Studio's new audio playground, which includes a two-speaker script editor. Flash-Lite TTS is also rolling out to everyone in Google Vids, Google's video editor. Gemini Enterprise access is listed as coming soon.
Two limits in the documentation matter if you build a product on it. A project can store 200 custom voices, shared between designed and replicated ones, and stored voices expire after a year. That is enough for a fixed cast of characters, but not for an app that creates a voice for every customer.
Who should switch
If you produce audiobooks, narrated courses or games with many characters, voice design removes a real cost: casting. Write the voice descriptions down and keep them in version control, so you can recreate a voice if it expires or drifts.
If you run a support line that reads order numbers aloud, you probably don't need the flagship. Run your own script through both tiers, since the API schema is identical and switching is one parameter. Then decide by ear. Voice quality claims are easy to publish and hard to hear, and your own script will tell you more than any leaderboard.



