StepAudio 3 ASR
Lowest word error rate on the Artificial Analysis transcription benchmark.
StepAudio 3 ASR is StepFun’s speech recognition model; it ties for the lowest AA-WER v2 word error rate on Artificial Analysis’ non-streaming speech-to-text leaderboard.
- Provider
- StepFun
- Type
- Speech-to-text
- Released
- Sep 14, 2026
- Price
- ≈$6.67 per 1,000 minutes
- Input
- audio
- Output
- text
- Open weights
- No
Best for
- Extraction & classification
- Multimodal
Strengths
- Lowest WER (≈1.7%)
More from StepFun
- Step 5 Preview — StepFun. StepFun’s new agentic flagship — 600B MoE with video input.