Skip to content
Google DeepMind·· 2026-04-16

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

AI summary

Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.

Selection record

AdmittedSum of both 142 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:Google发布新一代AI语音合成模型
Why it was chosen
Audio-tag controls, Elo scores and access across platforms help assess controllability in speech generation.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Google DeepMind · deepmind.google