Google DeepMind released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The former targets long-horizon coding and autonomous agents, priced like 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens.
Google DeepMind released Gemini 3.5 Transcribe, calling it the most accurate speech-to-text model available, converting raw audio directly into accurate, formatted text.
Google DeepMind released Gemini 3.7 Flash, positioning it as its strongest workhorse model for coding and agents, just three weeks after Gemini 3.6 Flash.
Google DeepMind released SL2T, a multilingual sign-language-to-text model, bringing the capability to consumer products for the first time. Gboard and Live Transcribe on Pixel 11 support ASL-to-English dictation, with more devices and languages to follow.
Meta released Muse Glimmer, a multimodal model distilled from Muse to 30B parameters under Apache 2.0, targeting local agent use cases such as coding, document analysis and personal assistants.
Google DeepMind released the WeatherNext AI model, achieving state-of-the-art cyclone track, intensity and wind-field structure forecasts and adding an average extra day of warning, equivalent to roughly a decade of meteorological progress.
Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.
Google DeepMind released Gemini Robotics ER 2 as a high-level brain for robots, supporting video understanding, multi-step task orchestration and multi-robot collaboration, while delegating action execution to lower-level VLA models.
Google DeepMind released its next-generation music model Lyria 3.5 on Google Flow Music. It claims improvements in musicality, lyrics, vocals and creative control, including more complex and natural melodic structures, lyrics with better prompt adherence and structural awareness, more expressive vocals with clearer pronunciation, and easier control of rhythm and duration.
Google DeepMind released Gemini Robotics 2, a next-generation robot intelligence layer achieving whole-body control of a complete humanoid robot for the first time, with fine manipulation using both hands and grippers.
Google DeepMind released Gemini 3.5 Flash Cyber, a lightweight cybersecurity model fine-tuned from 3.5 Flash for rapid vulnerability discovery, validation and patching. Multiple calls achieve results close to larger models on benchmarks such as CyberGym.
Thinking Machines released Inkling on Hugging Face, a multimodal MoE model with around one trillion parameters, a one-million-token context and native image, text and audio inputs, alongside Inkling-Small with 276 billion total and 12 billion active parameters.
Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.
Mistral released Leanstral 1.5 under Apache-2.0, with 119B total and 6B active parameters and major formal verification improvements. It achieves 100% on miniF2F, solves 587 of 672 PutnamBench problems, and sets current best results of 87% on FATE-H and 34% on FATE-X.
Google Research released TabFM for tabular classification and regression. It reframes table prediction as in-context learning, producing predictions in one forward pass without manual training, hyperparameter tuning or feature engineering.
Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.
Google DeepMind released experimental open-source model DiffusionGemma, generating text blocks in parallel through text diffusion and achieving up to 4 times faster inference on dedicated GPUs.