Google Research published SymptomAI research, using five Gemini Flash 2.0 agents with different questioning strategies for symptom interviews and differential diagnosis, involving 13,917 participants.
Thinking Machines released Inkling on Hugging Face, a multimodal MoE model with around one trillion parameters, a one-million-token context and native image, text and audio inputs, alongside Inkling-Small with 276 billion total and 12 billion active parameters.
Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.
Hugging Face and Cerebras jointly demonstrated a real-time speech-to-speech pipeline, accelerating Gemma 4 31B inference with Cerebras and combining Nvidia Parakeet speech recognition with Alibaba Qwen3TTS synthesis.
Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.
Google DeepMind released Gemini 3.5 Live Translate, a near-real-time speech-to-speech translation model supporting more than 70 languages, with automatic language detection and preservation of speakers' intonation, rhythm and pitch.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Google Research published a study in Nature introducing PHRM, which records video with a phone’s front camera in the seconds after face unlock and uses deep learning to estimate heart rate and resting heart rate in the background.
Google DeepMind added Street View grounding to experimental prototype Project Genie. Users select a real US location with a Maps pin, choose a style and describe a character, then Genie creates an interactive world whose starting location is grounded in real imagery.
Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family, generating high-quality video from combined image, audio, video and text inputs and supporting iterative video editing through natural-language conversations.
Google announced the expansion of content transparency and verification tools to Search, Gemini, Chrome, Pixel and Cloud. SynthID has watermarked over 100 billion images and videos and 60,000 years of audio. SynthID verification in the Gemini app has been used 50 million times and will reach Search and Chrome in the coming weeks.
Google DeepMind launched the Gemini 3.5 family with 3.5 Flash, focusing on agents and coding. It is available from today in the Gemini app, Google Search AI Mode, Google Antigravity, Gemini API and Gemini Enterprise.
Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.
Google DeepMind launched Gemini 3.1 Flash TTS, a new text-to-speech model emphasising stronger controllability, expressiveness and audio quality. From today, developers can preview it through Gemini API and Google AI Studio, enterprises through Vertex AI, and Workspace users can use it in Google Vids.
Google DeepMind released Gemini Robotics-ER 1.6, an upgraded reasoning-first robotics model with stronger spatial reasoning and multi-view understanding. New instrument reading capabilities cover circular pressure gauges, level indicators and digital readouts.
Google released Vibe Coding XR, combining Gemini with the XR Blocks framework based on WebXR, three.js and LiteRT.js to turn natural language prompts directly into physically aware Android XR apps, reportedly in under 60 seconds.
The AIMS collaboration between Google Research and several NHS organisations published two companion studies in Nature Cancer evaluating an AI breast cancer detection system within NHS screening workflows.