TurboQuant: Redefining AI efficiency with extreme compression
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
S2Vec uses S2 Geometry partitions and rasterised features to turn buildings and roads into multilayer images, learning general embeddings with masked autoencoding (MAE). It beats SATCLIP and GEOCLIP on zero-shot geographic socioeconomic extrapolation such as US population density and median income, but needs improvement on tree cover and elevation.
Mistral AI released its first text-to-speech model, Voxtral TTS, with 4B parameters and support for nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It is available through the API and Mistral Studio at $0.016 per 1k characters.
Google Research announced healthcare AI advances at The Check Up, including its Personal Health Agent (PHA) research with Fitbit, which found it supported long-term health better than single-task apps.
The AIMS collaboration between Google Research and several NHS organisations published two companion studies in Nature Cancer evaluating an AI breast cancer detection system within NHS screening workflows.
Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on proprietary knowledge, supporting pre-training, post-training and reinforcement learning. It handles dense and MoE architectures and, when needed, multimodal inputs, with training and governance on companies' own infrastructure.
Mistral AI released Mistral Small 4, the next major version in the series and its first model to combine Magistral reasoning, Pixtral multimodal capabilities and Devstral agentic coding in one model, under Apache 2.0.
Mistral AI joined NVIDIA Nemotron Coalition as a founding member. They plan to jointly develop frontier open AI models, with Mistral contributing architecture, multimodal capabilities and fine-tuning tools, and NVIDIA providing compute, model development tools and synthetic data pipelines.
Google and Cornell published a PNAS study asking GPT-4o, Perplexity, Claude 3.5, Gemini Advanced Pro 1.5, NotebookLM and a custom RAG system 67 expert superconductivity questions. Twelve international experts blindly assessed balance, comprehensiveness, concision, evidence, image relevance and qualitative feedback.
Mistral AI released Leanstral, the first open-source coding agent for Lean 4, using a sparse architecture with 6B active parameters. Its weights are available under Apache 2.0, and it is integrated into Mistral vibe and the free labs-leanstral-2603 API endpoint.
Berkeley AI Research proposed SPEX and ProxySPEX to identify key interactions driving LLM outputs at scale in feature, data and model-component attribution. SPEX turns interaction search into sparse recovery using sparsity and low order; ProxySPEX exploits hierarchy to match SPEX with roughly ten times fewer ablations.
Google introduced urban flash-flood forecasts on Flood Hub, using a new AI approach to predict risks in urban areas up to 24 hours ahead.
Google Research launched Groundsource, using Gemini to extract structured historical disaster data from global news. Its first open dataset contains 2.6 million urban flash-flood records across over 150 countries, spanning 2000 to the present.
Google Research and Beth Israel Deaconess Medical Center conducted a prospective, single-centre feasibility study in which AMIE held text consultations with patients before outpatient primary care visits, under live physician supervision by video.
Mistral built an autonomous agent on its open-source coding assistant Vibe to read Rails source files, generate or improve RSpec tests and run within CI/CD without human intervention.
Google Research released WAXAL, a large open speech dataset initially covering 27 sub-Saharan African languages spoken by more than 100 million people, under CC-BY-4.0.
Mistral released Voxtral Transcribe 2, comprising Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for real-time use.
Mistral released the terminal coding agent Mistral Vibe 2.0, powered by the Devstral 2 model family, adding custom subagents, multiple-choice clarifications, slash-command skills, unified agent modes and automatic updates.
Mistral AI investigated a vLLM memory leak in a disaggregated Prefill/Decode deployment of Mistral Medium 3.1. System memory grew linearly at 400 MB per minute, causing out-of-memory failures within hours. It occurred only on the decode side with graph compilation and KVCache transfer via NIXL. Heaptrack showed stable heap memory; the issue was outside the heap.
Berkeley AI Research proposed a mutual-information framework to evaluate and optimise imaging from noisy measurements and noise models. A NeurIPS 2025 paper validates decoder performance predictions for colour photography, radio astronomy, lensless imaging and microscopy. Designs match end-to-end state-of-the-art approaches with less memory and compute, without task-specific decoder design.
Mistral AI released Mistral OCR 3 with an overall 74% win rate against Mistral OCR 2 on forms, scans, complex tables and handwriting. The company says its accuracy exceeds enterprise document-processing and AI-native OCR solutions.
Mistral AI released the next-generation Devstral 2 coding family, including 123B Devstral 2 under a modified MIT licence and 24B Devstral Small 2 under Apache 2.0, both open-source.
Mistral released the Mistral 3 family, comprising 14B, 8B and 3B small dense models and its strongest yet Mistral Large 3, a sparse MoE with 41B active and 675B total parameters. All are open-source under Apache 2.0.
Mistral AI announced a multi-year SAP partnership to integrate its models into AI Foundation and jointly develop tailored solutions for complex European industries and the public sector. It will also accelerate vision-language-action model research with Helsing for defence and security, open a German office in the coming months and significantly expand its local team.
Berkeley AI Research proposed a divide-and-conquer off-policy reinforcement learning algorithm without temporal-difference (TD) learning. It reduces Bellman recursions from linear to logarithmic counts, scaling to long-horizon tasks.
Mistral released Mistral AI Studio, a production AI platform for enterprise teams built on three pillars: Observability, Agent Runtime and AI Registry.
Mistral AI announced a €1.7 billion Series C at a €11.7 billion post-money valuation, led by semiconductor equipment maker ASML. Existing investors including DST Global, Andreessen Horowitz, Bpifrance, General Catalyst, Index Ventures, Lightspeed and NVIDIA participated.
Le Chat Memories combines automatic saving, intelligent recall and visible sources. Users can disable memory, start memory-free incognito chats, edit or delete individual entries, export and import from other platforms. Memory Insights highlights trends and summaries from personal data, with categorisation and instant forgetting planned.
Mistral launched a beta directory of more than 20 secure MCP-based connectors in Le Chat across data, productivity, development, automation and business. It connects tools including Databricks, Snowflake, GitHub, Atlassian, Asana, Outlook, Box, Stripe and Zapier, and supports custom connectors or any remote MCP server.
Berkeley AI researchers proposed the first theory quantitatively predicting word2vec learning. Under realistic training conditions, it reduces to unweighted least-squares matrix factorisation with a closed-form gradient-flow solution, and the resulting representation is equivalent to PCA of target matrix M*. From small initialisation, word2vec learns orthogonal linear subspaces in discrete steps, increasing embedding rank and reducing loss. These features are M*'s leading eigenvectors and can be calculated in advance from corpus statistics and hyperparameters.
Mistral fine-tuned Pixtral-12B with LoRA, substantially outperforming the untuned baseline on Aerial Image Dataset (AID) satellite classification and reducing hallucinated invalid class names. Fine-tuning uses Mistral's API or LaPlateforme UI without extensive hyperparameter tuning, with 8,000 training and 2,000 test samples.
Mistral AI released Codestral 25.08 and a complete enterprise coding stack comprising Codestral, Codestral Embed, Devstral and the Mistral Code IDE plugin.
Mistral AI, Carbone 4 and France's ADEME completed the first full-lifecycle AI model analysis. By January 2025, Large 2 training and 18 months of use produced 20.4 ktCO₂e, consumed 281,000 cubic metres of water and 660 kg Sb eq of resources.
Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.
Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.
Mistral AI partnered with All Hands AI to release Devstral Medium and upgrade Devstral Small 1.1.
Mistral AI launched AI for Citizens, helping governments and public institutions strategically apply AI to transform public services, innovate and safeguard competitiveness. It offers open-source models, self-hosting, data sovereignty and tailored R&D, with government and public-sector partners in France, Luxembourg, Singapore, the Netherlands, the UK and Switzerland.
Mistral AI launched Mistral Compute, a private integrated AI infrastructure stack covering GPUs, orchestration, APIs, products and services, ranging from bare-metal servers to fully managed PaaS.
Mistral AI released its first reasoning model, Magistral, comprising the 24B open-source Magistral Small and enterprise-focused Magistral Medium.
Mistral released Mistral Code, an AI coding assistant combining Codestral, Codestral Embed, Devstral and Mistral Medium. It supports cloud, dedicated-capacity and local air-gapped GPU deployment, keeping code within enterprise boundaries.