Co-Scientist: A multi-agent AI partner to accelerate research
Google DeepMind published Co-Scientist research in Nature, describing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves scientific hypotheses.
Google DeepMind published Co-Scientist research in Nature, describing a Gemini-based multi-agent AI system that iteratively generates, debates and evolves scientific hypotheses.
Google DeepMind published a year’s progress for AlphaEvolve. The Gemini-powered coding agent has expanded from open problems in mathematics and computer science into genomics, power grids, quantum physics and AI infrastructure.
Jack Clark estimates in Import AI 455 a roughly 60% chance of AI R&D autonomously training successor models without human involvement by the end of 2028, and around 30% in 2027.
Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.
Google Research described how scientists use Empirical Research Assistance (ERA) to advance research in four areas: epidemic forecasting, cosmology, carbon monitoring and neuroscience.
Mistral AI released Workflows in public preview, positioning it as an enterprise AI orchestration layer with durable execution, observability and fault tolerance to take AI workflows from proof of concept to production.
Google Research released the ICLR paper ReasoningBank, an open-source agent memory framework that distils high-level reasoning patterns from successful and failed experiences.
Google DeepMind announced partnerships with Accenture, Bain & Company, BCG, Deloitte and McKinsey to scale frontier AI in enterprises. Partners get early access to models including Gemini and develop industry-specific solutions for finance, manufacturing, retail and media and entertainment. Currently only 25% of organisations have scaled AI into production.
Berkeley AI Research proposed GRASP, a gradient-based planner for learned world models. It lifts trajectories into virtual states for parallel optimisation across time, injects randomness directly into state iterations for exploration and reshapes gradients to give actions clear signals. Avoiding fragile state-input gradients in high-dimensional visual models makes long-horizon planning more practical and robust.
Google released Vantage, a research experiment using generative AI simulated conversations to assess future skills such as problem-solving and collaboration in school and university students, with English registration open on Google Labs. An Executive LLM dynamically guides dialogue and an AI Evaluator scores against rubrics. A study with New York University involving 188 US participants aged 18–25 found AI–expert scoring agreement close to agreement between two human experts.
Google Research released ConvApparel, with over 4,000 human–machine multi-turn clothing-shopping conversations, nearly 15,000 turns, to quantify realism gaps between LLM user simulators and real people.
Google released PaperVizAgent for paper illustrations and ScholarPeer for automated peer review. PaperVizAgent combines retrieval, planning, style, visualisation and review agents, scoring 60.2 on a 0–100 evaluation, ahead of GPT-Image-1.5, Nano-Banana-Pro and Paper2Any. It was the only framework above the 50.0 human baseline.
Mistral AI released Spaces CLI for human developers and coding agents. Commands such as spaces init and spaces dev handle scaffolding, development environments and deployment, with a corresponding flag for every interactive prompt so agents can work autonomously end to end. Each init generates context.json and AGENTS.md to provide structure and rules.
Google DeepMind unveiled an experimental Gemini-powered AI pointer that understands not only what it points at but what it means to the user. The team proposed four interaction principles: avoiding workflow interruption, capturing nearby visual and semantic context, supporting natural shorthand such as “this” and “that”, and turning pixels into actionable entities such as places, dates and objects.
Google released Vibe Coding XR, combining Gemini with the XR Blocks framework based on WebXR, three.js and LiteRT.js to turn natural language prompts directly into physically aware Android XR apps, reportedly in under 60 seconds.
Google Research announced healthcare AI advances at The Check Up, including its Personal Health Agent (PHA) research with Fitbit, which found it supported long-term health better than single-task apps.
Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on proprietary knowledge, supporting pre-training, post-training and reinforcement learning. It handles dense and MoE architectures and, when needed, multimodal inputs, with training and governance on companies' own infrastructure.
Mistral built an autonomous agent on its open-source coding assistant Vibe to read Rails source files, generate or improve RSpec tests and run within CI/CD without human intervention.
Mistral released the terminal coding agent Mistral Vibe 2.0, powered by the Devstral 2 model family, adding custom subagents, multiple-choice clarifications, slash-command skills, unified agent modes and automatic updates.
Mistral AI released the next-generation Devstral 2 coding family, including 123B Devstral 2 under a modified MIT licence and 24B Devstral Small 2 under Apache 2.0, both open-source.
Mistral released Mistral AI Studio, a production AI platform for enterprise teams built on three pillars: Observability, Agent Runtime and AI Registry.
Le Chat Memories combines automatic saving, intelligent recall and visible sources. Users can disable memory, start memory-free incognito chats, edit or delete individual entries, export and import from other platforms. Memory Insights highlights trends and summaries from personal data, with categorisation and instant forgetting planned.
Mistral launched a beta directory of more than 20 secure MCP-based connectors in Le Chat across data, productivity, development, automation and business. It connects tools including Databricks, Snowflake, GitHub, Atlassian, Asana, Outlook, Box, Stripe and Zapier, and supports custom connectors or any remote MCP server.
Mistral AI released Codestral 25.08 and a complete enterprise coding stack comprising Codestral, Codestral Embed, Devstral and the Mistral Code IDE plugin.
Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.
Mistral AI partnered with All Hands AI to release Devstral Medium and upgrade Devstral Small 1.1.
Mistral released Mistral Code, an AI coding assistant combining Codestral, Codestral Embed, Devstral and Mistral Medium. It supports cloud, dedicated-capacity and local air-gapped GPU deployment, keeping code within enterprise boundaries.
Mistral AI released Agents API, combining its language models with built-in connectors for code execution, image generation, document libraries and web search, plus MCP tools, persistent memory across conversations and agent orchestration.
Mistral AI and All Hands AI jointly released Devstral, an agentic LLM for software engineering tasks, open-source under Apache 2.0.
Mistral released enterprise AI assistant Le Chat Enterprise, powered by the new Mistral Medium 3. It offers enterprise search, agent building, custom data and tool connectors, document libraries, custom models and hybrid deployment, with all features rolling out over the next two weeks.
Mistral AI introduced TranscriptToPRDTicket, comprising PRDAgent and TicketCreationAgent powered by Mistral Large 2. It turns meeting transcripts into PRDs and development tickets, automatically creating them in tools such as Linear and Jira.