Google DeepMind released Gemini 3.5 Live Translate, a near-real-time speech-to-speech translation model supporting more than 70 languages, with automatic language detection and preservation of speakers' intonation, rhythm and pitch.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Mistral released Mistral Medium 3.5, its first 128B dense model combining instruction following, reasoning and coding. Its weights are available under a modified MIT licence, with a 256k context window and self-hosting possible on a minimum of four GPUs.
Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family, generating high-quality video from combined image, audio, video and text inputs and supporting iterative video editing through natural-language conversations.
Google DeepMind launched the Gemini 3.5 family with 3.5 Flash, focusing on agents and coding. It is available from today in the Gemini app, Google Search AI Mode, Google Antigravity, Gemini API and Gemini Enterprise.
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Google DeepMind released Gemini Robotics-ER 1.6, an upgraded reasoning-first robotics model with stronger spatial reasoning and multi-view understanding. New instrument reading capabilities cover circular pressure gauges, level indicators and digital readouts.
Mistral AI released its first text-to-speech model, Voxtral TTS, with 4B parameters and support for nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It is available through the API and Mistral Studio at $0.016 per 1k characters.
Mistral AI released Mistral Small 4, the next major version in the series and its first model to combine Magistral reasoning, Pixtral multimodal capabilities and Devstral agentic coding in one model, under Apache 2.0.
Mistral AI released Leanstral, the first open-source coding agent for Lean 4, using a sparse architecture with 6B active parameters. Its weights are available under Apache 2.0, and it is integrated into Mistral vibe and the free labs-leanstral-2603 API endpoint.
Mistral AI released Mistral OCR 3 with an overall 74% win rate against Mistral OCR 2 on forms, scans, complex tables and handwriting. The company says its accuracy exceeds enterprise document-processing and AI-native OCR solutions.
Mistral AI released the next-generation Devstral 2 coding family, including 123B Devstral 2 under a modified MIT licence and 24B Devstral Small 2 under Apache 2.0, both open-source.
Mistral released the Mistral 3 family, comprising 14B, 8B and 3B small dense models and its strongest yet Mistral Large 3, a sparse MoE with 41B active and 675B total parameters. All are open-source under Apache 2.0.
Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.
Mistral AI released its first code-focused embedding model, Codestral Embed, outperforming Voyage Code 3, Cohere Embed v4.0 and OpenAI's large embedding model on real-world code retrieval.