Google Research introduced a new method in Google Photos’ Auto frame to change camera viewpoints after a photo is taken. An internal 3D point-map estimation model reconstructs the scene and focal length, then a generative latent diffusion model fills gaps exposed by the new view. It automatically detects faces and wide-angle distortion to suggest ideal framing. The feature already applies automatically to photos containing people, with reframed versions available among Auto frame candidates.
Google DeepMind published a paper proposing Decoupled DiLoCo, splitting large-scale training into decoupled compute islands that communicate through asynchronous data streams, isolating local hardware failures.
Google Research released the ICLR paper ReasoningBank, an open-source agent memory framework that distils high-level reasoning patterns from successful and failed experiences.
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Google DeepMind released Gemini Robotics-ER 1.6, an upgraded reasoning-first robotics model with stronger spatial reasoning and multi-view understanding. New instrument reading capabilities cover circular pressure gauges, level indicators and digital readouts.
Google Quantum AI published a white paper saying future quantum computers would require fewer qubits and gates than previously estimated to break the elliptic-curve cryptography protecting cryptocurrencies.
Google DeepMind unveiled an experimental Gemini-powered AI pointer that understands not only what it points at but what it means to the user. The team proposed four interaction principles: avoiding workflow interruption, capturing nearby visual and semantic context, supporting natural shorthand such as “this” and “that”, and turning pixels into actionable entities such as places, dates and objects.
Google released Vibe Coding XR, combining Gemini with the XR Blocks framework based on WebXR, three.js and LiteRT.js to turn natural language prompts directly into physically aware Android XR apps, reportedly in under 60 seconds.
Google Research released TurboQuant, compressing the KV cache to three bits without loss of model accuracy or training and fine-tuning, alongside the QJL and PolarQuant methods.
The AIMS collaboration between Google Research and several NHS organisations published two companion studies in Nature Cancer evaluating an AI breast cancer detection system within NHS screening workflows.
Google Research launched Groundsource, using Gemini to extract structured historical disaster data from global news. Its first open dataset contains 2.6 million urban flash-flood records across over 150 countries, spanning 2000 to the present.
Google Research and Beth Israel Deaconess Medical Center conducted a prospective, single-centre feasibility study in which AMIE held text consultations with patients before outpatient primary care visits, under live physician supervision by video.
Google Research released WAXAL, a large open speech dataset initially covering 27 sub-Saharan African languages spoken by more than 100 million people, under CC-BY-4.0.