Skip to content

Speech and audio

Progress in AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

21selectedRelated topicsMultimodalAI videoProduct updates

Latest selected

21–21 of 21

Jul 15

Tuesday
  1. Voxtral

    Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.