Voxtral Small
Best open-weight transcription model — also understands audio.
Voxtral Small builds on Mistral Small 3 with audio input, excelling at speech transcription, translation and audio understanding while keeping strong text performance. Apache 2.0 weights.
- Provider
- Mistral AI
- Type
- Speech-to-text
- Released
- Jul 15, 2025
- Context window
- 33K tokens
- Price
- $0.10 input / $0.30 output per 1M tokens
- Input
- text, audio, file
- Output
- text
- Open weights
- Yes
- License
- Apache 2.0
Best for
- Extraction & classification
- Multimodal
- Local & on-device
Strengths
- Open weights
- ≈2.8% WER
- Audio Q&A, not just transcripts
More from Mistral AI
- Mistral Large 4 — Mistral AI. Mistral’s 1T-parameter multimodal MoE flagship (preview; weights promised).
- Mistral Medium 3.5 — Mistral AI. Fast enterprise-grade multimodal model.
- Mistral Small 4 — Mistral AI. Apache-licensed open model unifying reasoning, vision and coding.