Google DeepMind·· 2026-04-16
Gemini 3.1 Flash TTS: the next generation of expressive AI speech
Gemini 3.1 Flash TTS: the next generation of expressive AI speech
AI summary
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Selection record
Threshold 60Official, first-handFirst 71Second 71
AdmittedSum of both 142 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:Google发布新一代AI语音合成模型
- Why it was chosen
- Audio-tag controls, Elo scores and access across platforms help assess controllability in speech generation.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google DeepMind · deepmind.google