Skip to content

#Multimodal

1 today

Oct 1

ThursdayToday1 items
  1. Gemini 4 Argon: our next era of frontier intelligence

    Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

Sep 29

Tuesday
  1. One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact

    MSRA – Singapore, Microsoft's first research lab in Southeast Asia, spent its first year focusing on next-generation AI models and agent systems, domain AI, AI-native research practices and talent ecosystems. It works with Singapore's healthcare ecosystem on multimodal medical AI and self-evolving diagnostic agents, and jointly held a logistics and transport AI executive roundtable with EDB in February 2026.

Sep 26

Saturday

Sep 25

Friday

Sep 24

Thursday

Sep 23

Wednesday

Sep 3

Thursday
  1. Introducing WeatherNext 3, our most advanced and accurate global weather AI model

    Google DeepMind and Google Research released WeatherNext 3, calling it the most advanced and accurate global weather model available. It learns directly from real-time geostationary Earth-observation satellite data and generates hourly forecasts, with five-kilometre resolution for key surface variables and roughly five times the overall detail of WeatherNext 2.

Sep 2

Wednesday

Aug 28

Friday

Aug 26

Wednesday
  1. AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

    Google released research prototype AgentHands, published at CHI 2026, mapping LLM reasoning to speech-synchronised hand animations in XR headsets so agents can point and demonstrate object operations in 3D. It combines environment perception, a gesture event library, gesture-embedding reasoning and local synchronised execution, supporting deictic, iconic and expressive gestures. In an N=12 user study, it significantly improved spatial reference over voice-only interaction.

Aug 17

Monday

Aug 13

Thursday
  1. MindTopo reveals VLMs’ spatial reasoning abilities

    Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.

Aug 12

Wednesday
  1. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.

Aug 11

Tuesday

Aug 10

Monday

Aug 4

Tuesday
  1. Introducing Shieldstral.

    Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.

Jul 30

Thursday

Jul 28

Tuesday

Jul 27

Monday

Jul 23

Thursday

Jul 21

Tuesday

Jul 16

Thursday
  1. Newer Models, Same Advantage

    DharmaOCR scored 0.925 on a Portuguese OCR benchmark, ahead of Mistral OCR4's 0.798 and Unlimited-OCR's 0.7587. It specialised through two-stage training: supervised fine-tuning on Portuguese corpora, then DPO to stabilise inference. The author argues that concentrating parameters on one language remains a structural advantage despite emerging architectures.

Jul 15

Wednesday

Jul 8

Wednesday
  1. Introducing Robostral Navigate

    Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.

Jul 1

Wednesday

Jun 23

Tuesday
  1. Introducing Mistral OCR 4

    Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.

Jun 17

Wednesday
  1. From pixels to planning: Earth AI for nature restoration

    Google Research released vector data converting Farmscapes 2020 high-resolution raster maps into usable inventories of hedgerows, stone walls and coppices across over 130,000 square kilometres of the UK. The framework fine-tunes an RSF Vision-Transformer pretrained on over 300 million global satellite images with around 247 square kilometres of labelled data. Polsby–Popper compactness distinguishes woodland, clusters of trees and hedgerows, with a threshold below 0.5 for linear features.

Jun 13

Saturday
  1. Research into how AI can help users understand skin conditions

    Google published two studies on AI-assisted understanding of skin problems. In a survey of 2,345 participants, over 62% using AI tried to name a skin condition versus 41% in the control group, with 23% accuracy, nearly triple the control group's 8%. A mixed-methods study examined how people use these tools for their own skin concerns, their understanding and differences in doctor communication. AI offered limited help in deciding the next steps for seeking care.

Jun 9

Tuesday

Jun 5

Friday