Skip to content
Mistral AI·· 2025-07-15

Voxtral

Voxtral

AI summary

Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.

Selection record

AdmittedSum of both 160 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:发布Voxtral语音理解AI模型及评测
Why it was chosen
Open-source licensing, API prices and multilingual benchmark comparisons for both sizes inform speech-understanding choices and costs.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Mistral AI · mistral.ai