Skip to content

#Multimodal

2 today

Jul 8

Wednesday
  1. Introducing Robostral Navigate

    Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.

Jul 1

Wednesday

Jun 23

Tuesday
  1. Introducing Mistral OCR 4

    Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.

Jun 17

Wednesday
  1. From pixels to planning: Earth AI for nature restoration

    Google Research released vector data converting Farmscapes 2020 high-resolution raster maps into usable inventories of hedgerows, stone walls and coppices across over 130,000 square kilometres of the UK. The framework fine-tunes an RSF Vision-Transformer pretrained on over 300 million global satellite images with around 247 square kilometres of labelled data. Polsby–Popper compactness distinguishes woodland, clusters of trees and hedgerows, with a threshold below 0.5 for linear features.

Jun 15

Monday

Jun 13

Saturday
  1. Research into how AI can help users understand skin conditions

    Google published two studies on AI-assisted understanding of skin problems. In a survey of 2,345 participants, over 62% using AI tried to name a skin condition versus 41% in the control group, with 23% accuracy, nearly triple the control group's 8%. A mixed-methods study examined how people use these tools for their own skin concerns, their understanding and differences in doctor communication. AI offered limited help in deciding the next steps for seeking care.

Jun 9

Tuesday

Jun 5

Friday

May 29

Friday

May 18

Monday
  1. Introducing Gemini Omni

    Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family, generating high-quality video from combined image, audio, video and text inputs and supporting iterative video editing through natural-language conversations.

May 17

Sunday

May 16

Saturday
  1. Strengthening Singapore’s AI Future: A New National Partnership

    Google DeepMind and Singapore's government formed a national AI partnership with new projects in healthcare, research, education and climate. It explores an AI-assisted triadic-care model, uses AlphaFold and Google Earth for Southeast Asian infectious-disease research and develops a Gemma-based running assistant to help visually impaired athletes train independently.

Apr 30

Thursday
  1. Enabling a new model for healthcare with AI co-clinician

    Google DeepMind announced AI co-clinician research to explore AI agents assisting patient care under clinical supervision. In blinded assessments of 98 real primary-care queries, 97 responses had no critical errors, and doctors preferred them to existing evidence-synthesis tools. Across 140 consultation skills, AI matched or exceeded primary-care doctors on 68, but expert doctors were better overall at recognising red flags and guiding key physical examinations.

Apr 16

Thursday

Apr 13

Monday

Apr 9

Thursday

Apr 3

Friday

Mar 25

Wednesday

Mar 18

Wednesday

Mar 17

Tuesday

Dec 17

Wednesday

Dec 3

Wednesday

Aug 1

Friday

Jul 17

Thursday
  1. Le Chat dives deep.

    Mistral added features to Le Chat including preview Deep Research, Voxtral-powered voice mode, multilingual reasoning with Magistral, Projects for organising conversations and image editing with Black Forest Labs.

May 7

Wednesday

Mar 17

Monday

Mar 7

Friday
  1. Mistral OCR

    Mistral AI introduced Mistral OCR, an optical character recognition API that reads images and PDFs and outputs text and images interleaved in order. It is already the default document understanding model for millions of Le Chat users.