Start building with Nano Banana 2 Lite and Gemini Omni Flash
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image) and Gemini Omni Flash (gemini-omni-flash-preview).
Google DeepMind released Nano Banana 2 Lite (gemini-3.1-flash-lite-image) and Gemini Omni Flash (gemini-omni-flash-preview).
Dharma AI discusses Goldfeder, Wyder, LeCun and Shwartz-Ziv's 2026 paper AI Must Embrace Specialization via Superhuman Adaptable Intelligence.
Google Research released TabFM for tabular classification and regression. It reframes table prediction as in-context learning, producing predictions in one forward pass without manual training, hyperparameter tuning or feature engineering.
Every Eval Ever (EEE) and Hugging Face Community Evals are now interoperable, allowing cross-publication of evaluations with links to complete records. EEE stores around 229,000 results across more than 22,000 models and 2,200 benchmarks from 31 reporting formats.
Hugging Face proposed DiScoFormer (Density and Score Transformer), estimating a distribution's density and score simultaneously in one forward pass from a set of data points, without retraining for new distributions.
NVIDIA's ENPIRE lets coding agents drive autonomous experiments and iteration with real robot arms. Dual YAM arms and an RTX 5090 workstation achieved 99% success on PushT, pin organisation and cutting cable ties. GPT-5.5/Codex and Opus 4.7/Claude Code traded wins, while Kimi-2.6 lagged. Eight agents found higher-scoring solutions faster.
Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.
Google Research proposed linear elastic caching, modelling page eviction as a ski-rental problem and using shallow decision trees to predict page TTLs, dynamically resizing caches to minimise total ownership cost. After months in Spanner production, memory fell 15.5%, cache misses rose only 5.5%, TCO dropped around 5% and actual I/O cost impact was just 0.5%. It was also validated on several public cache traces.
A Google Research paper at COLM 2026 finds reasoning traces unlock factual knowledge LLMs otherwise cannot recall, even for simple single-hop questions. Tests on Gemini-2.5 Flash and Pro and Qwen3-32B identify two mechanisms: extra tokens act as a 'computation buffer', and 'factual priming' produces related facts to semantically prepare the correct answer. Self-generated intermediate facts can also introduce hallucination risks.
Google DeepMind integrated computer use as a built-in tool in Gemini 3.5 Flash. Previously, it was available only as the standalone Gemini 2.5 computer use model.
Mistral AI added new Connectors capabilities. Admin controls for workspace- or organisation-level access and individual tool toggles are generally available, as are API keys scoped to connectors.
Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.
Researchers from Oxford, the UK AI safety institute, Stanford and LSE found AI consistently more persuasive in text across four experiments with 6,923 participants and 18,978 conversations, even when experts chose topics, researched in advance and were incentivised with £1,000 prizes.
Google DeepMind announced work with the UK government, Google Cloud, Faculty and planning departments in Barnet, Camden and Dorset on an AI planning-approval prototype, aiming to help officers reduce householder application processing time by 50%.
Google Research released vector data converting Farmscapes 2020 high-resolution raster maps into usable inventories of hedgerows, stone walls and coppices across over 130,000 square kilometres of the UK. The framework fine-tunes an RSF Vision-Transformer pretrained on over 300 million global satellite images with around 247 square kilometres of labelled data. Polsby–Popper compactness distinguishes woodland, clusters of trees and hedgerows, with a threshold below 0.5 for linear features.
Google DeepMind released AI Control Roadmap, a framework for building and managing advanced AI deployed inside Google. It applies defence in depth, adding system-level safety layers beyond model alignment to provide protection even when alignment is imperfect.
The UK AISI alignment team and Timaeus co-founded nonprofit Sequent, arguing alignment research is unprepared for superintelligence. It plans 40–80 full-time staff within two to three years and initial fundraising of $100–150 million.
Google published two studies on AI-assisted understanding of skin problems. In a survey of 2,345 participants, over 62% using AI tried to name a skin condition versus 41% in the control group, with 23% accuracy, nearly triple the control group's 8%. A mixed-methods study examined how people use these tools for their own skin concerns, their understanding and differences in doctor communication. AI offered limited help in deciding the next steps for seeking care.
With Google's support, UC San Diego plans a data centre using motherboards from 2,000 retired Pixel phones to provide low-cost, low-carbon cloud computing for hundreds of students and faculty. SPEC benchmarks show 25–50 phones roughly equal one modern server. Phones form Kubernetes-managed clusters of 25–50 devices, with launch expected in autumn 2026.
At AISTATS 2026, Google Research proposed Regularized f-Divergence Kernel Tests for machine unlearning audits, using relative distances to determine whether models are closer to safely retrained versions or original compromised models.
Google DeepMind released experimental open-source model DiffusionGemma, generating text blocks in parallel through text diffusion and achieving up to 4 times faster inference on dedicated GPUs.
Google DeepMind, Schmidt Sciences, Cooperative AI Foundation and ARIA, supported by Google.org, launched up to $10 million in global technical research funding focused on collective behaviour and safety risks in large-scale multi-agent AI systems.
Google DeepMind released Gemini 3.5 Live Translate, a near-real-time speech-to-speech translation model supporting more than 70 languages, with automatic language detection and preservation of speakers' intonation, rhythm and pitch.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Google DeepMind launched a three-month Robotics accelerator for 15 early-stage European robotics start-ups, providing technical guidance and access to its AI stack and Gemini robotics models. Covering logistics, manufacturing, healthcare, climate and navigation, it starts in London this week.
Google DeepMind published results from a preregistered randomised controlled trial in Sierra Leone. Over eight weeks, students using Guided Learning improved maths scores by 0.258 standard deviations over controls, equivalent to around 1.2–1.7 years of conventional progress.
Import AI covers three studies. King's College London, Fudan and Alan Turing Institute built SocioHack with 72 sandbox social environments to test RL exploiting institutional loopholes while complying with rules. Models rediscovered patched historical vulnerabilities with 61.25% recall and 90.85% precision.
Google launched a public preview of Agentic RAG-powered cross-corpus retrieval on Gemini Enterprise Agent Platform. Roles including Orchestrator, Planner, Query Rewriter, Search Fanout and Sufficient Context Agent collaborate on multi-source, multi-hop queries.
Google Research published a study in Nature introducing PHRM, which records video with a phone’s front camera in the seconds after face unlock and uses deep learning to estimate heart rate and resting heart rate in the background.
Google Research open-sourced its hydrological modelling framework on GitHub under Apache 2.0, enabling national weather and hydrology agencies to integrate AI flood forecasting. The Python package uses PyTorch and an LSTM architecture, can train or fine-tune on Caravan data, and includes interactive tutorial notebooks and videos.
Import AI 459 focuses on three studies: a UK AI Safety Institute paper says automated alignment research using AI to supervise AI training is not a universal solution and is harder to implement than expected; other work covers scaling laws for protein-folding models and pricing AI-system extinction risks.
Google Research unveiled the experimental Gemini for Science toolkit at I/O 2026, including Computational Discovery based on ERA and AlphaEvolve, Hypothesis Generation based on Co-Scientist and Literature Insights based on NotebookLM.
Mistral released an AI stack for industrial engineering at AI Now Summit 2026, partnering with Airbus, BMW and ASML to optimise design, simulation and production while retaining control over proprietary data and IP.
Mistral upgraded Le Chat to Vibe, a unified AI agent covering work and coding, retaining all existing conversations, settings and subscription plans.
Mistral AI released the public preview of Search Toolkit, a composable framework for production search pipelines in AI applications. It unifies ingestion, retrieval and evaluation through shared interfaces, is open-source and deploys in cloud, local or edge environments.
Google combined a new single-message encrypted aggregation protocol with TEEs, letting devices submit without multiple online rounds. Google receives only anonymised group insights; raw data is neither exposed nor reconstructed even within hardware protection. TEE attestation proves execution follows public code.
After incorporating Emmi AI, Mistral launched physical AI capabilities for AI-native industrial engineering with partners including ASML, Airbus, Safran and Siemens Energy. The model predicts physical fields directly from geometry and boundary conditions in seconds through one forward pass on a single GPU, versus hours to weeks per design variant in traditional CFD/FEM. It accelerates design iteration while retaining traditional solvers for validation and edge cases.
Mistral acquired Emmi AI and doubled down on fundamental Physics AI research for aerospace, automotive, semiconductor and energy industries. Results include AB-UPT handling 9 million surface and 140 million volume elements on one GPU, and around 30,000 CFD simulations of transonic flow around 3D wings.
Import AI 458 includes an essay based on a 2026 Oxford HAI Lab talk and fiction imagining a 'positive singularity'. Using the Epoch Capabilities Index (ECI, covering over 40 benchmarks), the talk argues continued AI progress presents a choice between exploring the future and avoiding the present.
Mistral AI signed a definitive agreement this week to acquire Physics AI pioneer Emmi AI, strengthening its AI transformation services for industrial companies. Founded in Austria, Emmi AI has over 30 researchers and engineers focusing on large engineering models that replace days of computation with real-time simulation and build digital twins. Its co-founders and team will join Mistral's Science and Applied AI teams in May.