Mistral AI·· 2025-07-15
Voxtral
Voxtral
AI summary
Mistral AI released Voxtral speech-understanding models in 24B and 3B versions, both under Apache 2.0 and available via its API. With a 32k-token context, they handle up to 30 minutes of transcription or 40 minutes of audio understanding, including built-in question answering and summaries, automatic multilingual detection and voice-triggered function calls, inheriting Mistral Small 3.1's text capabilities.
Selection record
Threshold 60Official, first-handFirst 78Second 82
AdmittedSum of both 160 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:发布Voxtral语音理解AI模型及评测
- Why it was chosen
- Open-source licensing, API prices and multilingual benchmark comparisons for both sizes inform speech-understanding choices and costs.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Mistral AI · mistral.ai