Skip to content

All AI news

4 today

Aug 21

Friday

Aug 14

Friday

Aug 12

Wednesday

Aug 11

Tuesday

Aug 10

Monday

Aug 6

Thursday

Aug 4

Tuesday
  1. Introducing Shieldstral.

    Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.

Jul 30

Thursday
  1. We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control

    Google DeepMind released its next-generation music model Lyria 3.5 on Google Flow Music. It claims improvements in musicality, lyrics, vocals and creative control, including more complex and natural melodic structures, lyrics with better prompt adherence and structural awareness, more expressive vocals with clearer pronunciation, and easier control of rhythm and duration.

Jul 28

Tuesday

Jul 27

Monday

Jul 21

Tuesday

Jul 17

Friday

Jul 16

Thursday
  1. Newer Models, Same Advantage

    DharmaOCR scored 0.925 on a Portuguese OCR benchmark, ahead of Mistral OCR4's 0.798 and Unlimited-OCR's 0.7587. It specialised through two-stage training: supervised fine-tuning on Portuguese corpora, then DPO to stabilise inference. The author argues that concentrating parameters on one language remains a structural advantage despite emerging architectures.

Jul 15

Wednesday

Jul 8

Wednesday
  1. Introducing Robostral Navigate

    Mistral AI released its first embodied navigation model, Robostral Navigate. The 8B model uses one ordinary RGB camera without LiDAR or depth sensors, achieving 76.6% success in unseen R2R-CE validation environments, 9.7 percentage points above the best single-camera approach and 4.5 points above the best depth- or multi-camera system.

Jul 2

Thursday

Jul 1

Wednesday

Jun 30

Tuesday

Jun 27

Saturday

Jun 23

Tuesday
  1. Introducing Mistral OCR 4

    Mistral released OCR 4, returning bounding boxes, block-type classifications and per-page and per-word confidence alongside text extraction. It supports 170 languages and single-container self-hosting. Independent annotators preferred OCR 4 on average 72% of the time in blind evaluations of over 600 documents. It scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench, though Mistral warns both benchmarks have known scoring limitations.

Jun 11

Thursday

Jun 9

Tuesday

May 22

Friday

May 18

Monday
  1. Introducing Gemini Omni

    Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family, generating high-quality video from combined image, audio, video and text inputs and supporting iterative video editing through natural-language conversations.

May 16

Saturday

Apr 16

Thursday

Apr 13

Monday

Apr 3

Friday

Mar 24

Tuesday
  1. Speaking of Voxtral

    Mistral AI released its first text-to-speech model, Voxtral TTS, with 4B parameters and support for nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It is available through the API and Mistral Studio at $0.016 per 1k characters.

Mar 17

Tuesday

Feb 5

Thursday

Dec 17

Wednesday

Dec 9

Tuesday

Dec 3

Wednesday

Jul 15

Tuesday
  1. Voxtral

    Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.

Jul 11

Friday

Jun 10

Tuesday