Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to a native live conversation model to create a dynamic visual avatar that can hear, see and speak. It is available on Gemini Enterprise from today.
Meta positioned Muse as a personal agent at Connect 2026 and announced hardware and feature updates around it. Muse supports voice and live video, runs tasks in the background and will come to all Meta glasses. Each Muse has its own email address, the Mac version supports computer use, and its connector platform has more than 1,500 apps.
Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The former targets character design and line-by-line performance direction, while the latter targets high-concurrency voiceovers and voice agents.
Google DeepMind released two real-time conversation models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The former targets scale and cost efficiency, while the latter targets complex tasks and multistep reasoning.
OpenAI launched GPT-Live-1 in the API, supporting natural full-duplex voice conversations with stronger instruction following, custom voices and telephony support.
Voice Arena and Hugging Face added Monsoon en-IN and Monsoon hi-IN evaluation datasets to Open ASR Leaderboard, making Hindi its first Indian language.
Google DeepMind released Gemini 3.5 Transcribe, calling it the most accurate speech-to-text model available, converting raw audio directly into accurate, formatted text.
Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.
Google Research released AMIE (Video), built on Gemini and Project Astra, for real-time video clinical consultations. It can perceive non-verbal cues and guide virtual physical examinations.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Google DeepMind released its next-generation music model Lyria 3.5 on Google Flow Music. It claims improvements in musicality, lyrics, vocals and creative control, including more complex and natural melodic structures, lyrics with better prompt adherence and structural awareness, more expressive vocals with clearer pronunciation, and easier control of rhythm and duration.
Hume released Real World VoiceEQ, a speech evaluation benchmark covering over 40 proprietary and open-source voice models, more than 15 evaluation dimensions and over 60 metrics across ASR, TTS, S2S and speech understanding.
Hugging Face and Cerebras jointly demonstrated a real-time speech-to-speech pipeline, accelerating Gemma 4 31B inference with Cerebras and combining Nvidia Parakeet speech recognition with Alibaba Qwen3TTS synthesis.
Google DeepMind released Gemini 3.5 Live Translate, a near-real-time speech-to-speech translation model supporting more than 70 languages, with automatic language detection and preservation of speakers' intonation, rhythm and pitch.
Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Mistral AI released its first text-to-speech model, Voxtral TTS, with 4B parameters and support for nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It is available through the API and Mistral Studio at $0.016 per 1k characters.
Google Research released WAXAL, a large open speech dataset initially covering 27 sub-Saharan African languages spoken by more than 100 million people, under CC-BY-4.0.
Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.