StepFun Ships Five StepAudio 3 Models, Topping Independent Leaderboards for Voice Reasoning and Conversation
Technology📅 September 15, 2026👤 FreeReadText Team

StepFun Ships Five StepAudio 3 Models, Topping Independent Leaderboards for Voice Reasoning and Conversation

Shanghai AI lab StepFun has released the StepAudio 3 series — five models covering real-time conversation, transcription, speech synthesis, general audio generation, and music on a single platform. The Realtime model tops Artificial Analysis rankings for speech reasoning and conversational dynamics, though independent testing flags an 8.83-second time-to-first-audio.

On September 15, 2026, Shanghai AI lab StepFun released the StepAudio 3 series, shipping five models at once on its open platform: StepAudio 3 Realtime for native full-duplex conversation and tool calls, ASR for transcription, TTS for speech generation, Gen for unified audio generation, and Music for song creation. The release packages most of the audio stack a voice-agent or media-production workflow needs under one model family. StepFun was founded in 2023 by Daxin Jiang, a former Microsoft global vice president who worked on speech and Azure AI — a background that makes the audio series a direct extension of the founder's earlier work rather than a side project attached to a text-model lab.

StepFun's strongest evidence comes from independent testing. Artificial Analysis ranks StepAudio 3 Realtime first for speech reasoning with 99.7% on its 1,000-question Big Bench Audio test, and first on conversational dynamics at 98.9% — ahead of Alibaba Cloud's Qwen Audio 3.0 Realtime Plus at 98.4% and OpenAI's GPT-Realtime-2 High at 95.3%. Conversational dynamics measures pauses, interruptions, turn-taking, and short acknowledgments — the behaviors that decide whether a voice agent feels usable after the demo ends. StepAudio 3 ASR ties Alibaba's Fun-Realtime-ASR-preview for the lowest word error rate among 59 models tested, at 1.7%, and the model uses context and domain knowledge to recognize names, homophones, and technical vocabulary across medicine, finance, law, and software development.

The same evaluation exposes the launch's weakness. StepAudio 3 Realtime took an average 8.83 seconds to produce its first audio on the benchmark, against 0.44 seconds for the fastest model listed and roughly one to three seconds for models from Google, OpenAI, Alibaba, and xAI. StepFun's technical reports, submitted to arXiv in the days before the launch, describe a 'Think-While-Speaking' system that runs private reasoning alongside spoken output — a design targeting the tension between deliberation and latency that the independent result shows remains unresolved. On pricing, TTS costs $0.36 per 10,000 characters, nearly 58% below the $0.85 of StepAudio 2.5 TTS, with voice cloning at $1.50 per voice; ASR Max runs $0.24 per audio hour, nearly 11 times the budget StepAudio 2.5 ASR; and Realtime, Gen, and Music are available as free limited-time preview APIs while final pricing is settled.

The five-model release puts StepFun in two crowded businesses at once: Realtime, ASR, and TTS target developers building customer-service agents and voice interfaces, while Gen and Music compete for media-production work. StepAudio 3 Gen consolidates speech, vocals, sound effects, and music through a shared discrete audio representation, so a single prompt can specify voices, timing, ambience, and music in one clip. StepFun is selling a single API stack for voice agents and media generation with credible benchmark results on quality — but production adoption will depend on whether it can cut real-time latency and convert the free previews into competitive pricing once the launch window closes.

StepFunStepAudio 3Text-to-SpeechSpeech RecognitionVoice AIAudio GenerationArtificial Analysis

Lähde

← Back to News