Intelligent transcription with Gemini 3.5 Transcribe
Google DeepMind released Gemini 3.5 Transcribe, calling it the most accurate speech-to-text model available, converting raw audio directly into accurate, formatted text.
Google DeepMind released Gemini 3.5 Transcribe, calling it the most accurate speech-to-text model available, converting raw audio directly into accurate, formatted text.
The IBM Granite team released Granite 4.2, its first dense, decoder-only reasoning model family, in 3B, 8B and 30B sizes, all open-source under Apache 2.0.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Google DeepMind released Gemini 3.7 Flash, positioning it as its strongest workhorse model for coding and agents, just three weeks after Gemini 3.6 Flash.
Google DeepMind released SL2T, a multilingual sign-language-to-text model, bringing the capability to consumer products for the first time. Gboard and Live Transcribe on Pixel 11 support ASL-to-English dictation, with more devices and languages to follow.
NVIDIA released Magpie TTS Multilingual, a 364M-parameter open-weight speech synthesis model supporting English, Spanish, French, German, Italian, Vietnamese, Chinese, Hindi and Japanese, plus newly added Modern Standard Arabic, Korean and Brazilian Portuguese, for 12 languages in total.
Meta released Muse Glimmer, a multimodal model distilled from Muse to 30B parameters under Apache 2.0, targeting local agent use cases such as coding, document analysis and personal assistants.
Google DeepMind released the WeatherNext AI model, achieving state-of-the-art cyclone track, intensity and wind-field structure forecasts and adding an average extra day of warning, equivalent to roughly a decade of meteorological progress.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Google DeepMind released Gemini Robotics ER 2 as a high-level brain for robots, supporting video understanding, multi-step task orchestration and multi-robot collaboration, while delegating action execution to lower-level VLA models.
Google DeepMind released its next-generation music model Lyria 3.5 on Google Flow Music. It claims improvements in musicality, lyrics, vocals and creative control, including more complex and natural melodic structures, lyrics with better prompt adherence and structural awareness, more expressive vocals with clearer pronunciation, and easier control of rhythm and duration.
Google DeepMind released Gemini Robotics 2, a next-generation robot intelligence layer achieving whole-body control of a complete humanoid robot for the first time, with fine manipulation using both hands and grippers.
NVIDIA introduced Cosmos-H-Dreams, a real-time action-conditioned generative simulator for surgical robots. It runs through FlashDreams on one RTX PRO 6000 GPU and supports closed-loop interaction by humans and policies.
Google DeepMind released three models: Gemini 3.6 Flash, 3.5 Flash-Lite and cybersecurity-focused 3.5 Flash Cyber.
Google DeepMind released Gemini 3.5 Flash Cyber, a lightweight cybersecurity model fine-tuned from 3.5 Flash for rapid vulnerability discovery, validation and patching. Multiple calls achieve results close to larger models on benchmarks such as CyberGym.
DharmaOCR scored 0.925 on a Portuguese OCR benchmark, ahead of Mistral OCR4's 0.798 and Unlimited-OCR's 0.7587. It specialised through two-stage training: supervised fine-tuning on Portuguese corpora, then DPO to stabilise inference. The author argues that concentrating parameters on one language remains a structural advantage despite emerging architectures.
Thinking Machines released Inkling on Hugging Face, a multimodal MoE model with around one trillion parameters, a one-million-token context and native image, text and audio inputs, alongside Inkling-Small with 276 billion total and 12 billion active parameters.
Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.
Mistral released Leanstral 1.5 under Apache-2.0, with 119B total and 6B active parameters and major formal verification improvements. It achieves 100% on miniF2F, solves 587 of 672 PutnamBench problems, and sets current best results of 87% on FATE-H and 34% on FATE-X.
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image) and Gemini Omni Flash (gemini-omni-flash-preview).
Google Research released TabFM for tabular classification and regression. It reframes table prediction as in-context learning, producing predictions in one forward pass without manual training, hyperparameter tuning or feature engineering.
Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.
Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.
Google DeepMind released experimental open-source model DiffusionGemma, generating text blocks in parallel through text diffusion and achieving up to 4 times faster inference on dedicated GPUs.
Google DeepMind released Gemini 3.5 Live Translate, a near-real-time speech-to-speech translation model supporting more than 70 languages, with automatic language detection and preservation of speakers' intonation, rhythm and pitch.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Mistral released Mistral Medium 3.5, its first 128B dense model combining instruction following, reasoning and coding. Its weights are available under a modified MIT licence, with a 256k context window and self-hosting possible on a minimum of four GPUs.
Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family, generating high-quality video from combined image, audio, video and text inputs and supporting iterative video editing through natural-language conversations.
Google DeepMind launched the Gemini 3.5 family with 3.5 Flash, focusing on agents and coding. It is available from today in the Gemini app, Google Search AI Mode, Google Antigravity, Gemini API and Gemini Enterprise.
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Google DeepMind released Gemini Robotics-ER 1.6, an upgraded reasoning-first robotics model with stronger spatial reasoning and multi-view understanding. New instrument reading capabilities cover circular pressure gauges, level indicators and digital readouts.
Google DeepMind released the open-source Gemma 4 family in four sizes—E2B, E4B, 26B MoE and 31B Dense—under Apache 2.0.
Mistral AI released its first text-to-speech model, Voxtral TTS, with 4B parameters and support for nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It is available through the API and Mistral Studio at $0.016 per 1k characters.
Mistral AI released Mistral Small 4, the next major version in the series and its first model to combine Magistral reasoning, Pixtral multimodal capabilities and Devstral agentic coding in one model, under Apache 2.0.
Mistral AI released Leanstral, the first open-source coding agent for Lean 4, using a sparse architecture with 6B active parameters. Its weights are available under Apache 2.0, and it is integrated into Mistral vibe and the free labs-leanstral-2603 API endpoint.
Mistral released Voxtral Transcribe 2, comprising Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for real-time use.
Mistral AI released Mistral OCR 3 with an overall 74% win rate against Mistral OCR 2 on forms, scans, complex tables and handwriting. The company says its accuracy exceeds enterprise document-processing and AI-native OCR solutions.
Mistral AI released the next-generation Devstral 2 coding family, including 123B Devstral 2 under a modified MIT licence and 24B Devstral Small 2 under Apache 2.0, both open-source.
Mistral released the Mistral 3 family, comprising 14B, 8B and 3B small dense models and its strongest yet Mistral Large 3, a sparse MoE with 41B active and 675B total parameters. All are open-source under Apache 2.0.
Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.