Skip to content

Multimodal

Abilities beyond text: visual understanding, mixed image and text, and audio and video input and output.

Latest selected

1–20 of 47

Oct 1

ThursdayToday
  1. Gemini 4 Argon: our next era of frontier intelligence

    Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

Sep 29

Tuesday

Sep 25

Friday

Sep 24

Thursday

Sep 23

Wednesday

Sep 22

Tuesday

Sep 3

Thursday
  1. Introducing WeatherNext 3, our most advanced and accurate global weather AI model

    Google DeepMind and Google Research released WeatherNext 3, calling it the most advanced and accurate global weather model available. It learns directly from real-time geostationary Earth-observation satellite data and generates hourly forecasts, with five-kilometre resolution for key surface variables and roughly five times the overall detail of WeatherNext 2.

Sep 2

Wednesday

Aug 28

Friday

Aug 17

Monday

Aug 12

Wednesday

Aug 11

Tuesday

Aug 10

Monday

Aug 4

Tuesday
  1. Introducing Shieldstral.

    Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.

Jul 30

Thursday

Jul 28

Tuesday