Google DeepMind·· 2026-08-27
Intelligent transcription with Gemini 3.5 Transcribe
Intelligent transcription with Gemini 3.5 Transcribe
AI summary
Google DeepMind released Gemini 3.5 Transcribe, calling it the most accurate speech-to-text model available, converting raw audio directly into accurate, formatted text.
Selection record
Threshold 60Official, first-handFirst 74Second 71
AdmittedSum of both 145 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:Google发布Gemini语音转文字模型
- Why it was chosen
- WER, latency and multilingual metrics, plus integration details for two APIs, support assessment of speech pipeline options.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google DeepMind · deepmind.google