GDM last shipped a larger-than-Flash model in February ( 3.1 Pro ), and after successive incremental 3.x Flash versions and the big GDM management shakeup last month, the largest question for GDM was when they would catch up to peers who have in the meantime launched Fable and Astra class models.
Google unveiled Gemini 4 Argon, its new frontier model that closes the gap with rivals from OpenAI and Anthropic, beating some of them on key benchmarks. While it may not clearly lead the pack, it is relatively cheap for a frontier model, at least at the introductory price.
Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
TypeSafe AI founder and CEO Diogo Almeida introduced Jev on the Latent Space podcast, describing it as a System One large model designed for consumption by software.
NVIDIA submitted preview results for Vera Rubin NVL72 to MLPerf Inference v6.1 for the first time, achieving throughput up to 3.7 times that of GB300 NVL72 on Qwen3-VL and up to 2.5 times on DeepSeek-R1.
NVIDIA announced several advances for Vera Rubin and the DSX platform at AI Infra Summit, focusing on energy-efficiency improvements in token throughput per megawatt.
Multiverse Computing published a paper proposing Quantization-Aware Healing (QAH). After compressing GPT-OSS 120B to 60B parameters and quantising it to MXFP4, the method distils directly from the original uncompressed model rather than a reconstructed bfloat16 checkpoint.
Liquid AI released DSpark draft model checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. Speculative decoding accelerates decoding without changing output quality, increasing throughput by up to 3.18 times on GPUs and 2.87 times on devices.
Google Research released the Chain-of-Evidence (CoE) verifiability framework, implemented in a Science One Framework prototype, with automated CoE Audit metrics.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Thinking Machines released Inkling on Hugging Face, a multimodal MoE model with around one trillion parameters, a one-million-token context and native image, text and audio inputs, alongside Inkling-Small with 276 billion total and 12 billion active parameters.
Mistral released Leanstral 1.5 under Apache-2.0, with 119B total and 6B active parameters and major formal verification improvements. It achieves 100% on miniF2F, solves 587 of 672 PutnamBench problems, and sets current best results of 87% on FATE-H and 34% on FATE-X.
Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.
Google DeepMind released experimental open-source model DiffusionGemma, generating text blocks in parallel through text diffusion and achieving up to 4 times faster inference on dedicated GPUs.
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
Mistral AI released Mistral Small 4, the next major version in the series and its first model to combine Magistral reasoning, Pixtral multimodal capabilities and Devstral agentic coding in one model, under Apache 2.0.
Mistral AI released Leanstral, the first open-source coding agent for Lean 4, using a sparse architecture with 6B active parameters. Its weights are available under Apache 2.0, and it is integrated into Mistral vibe and the free labs-leanstral-2603 API endpoint.
Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.