StepAudio 3 ASR

Lowest word error rate on the Artificial Analysis transcription benchmark.

StepAudio 3 ASR is StepFun’s speech recognition model; it ties for the lowest AA-WER v2 word error rate on Artificial Analysis’ non-streaming speech-to-text leaderboard.

Provider
StepFun
Type
Speech-to-text
Released
Sep 14, 2026
Price
≈$6.67 per 1,000 minutes
Input
audio
Output
text
Open weights
No

Best for

Strengths

More from StepFun

Sources