Alibaba's Qwen-Audio-3.0-TTS Tops Global Leaderboard: 16 Languages, Voice Cloning, and Sub-200ms Latency
Technology📅 July 20, 2026👤 FreeReadText Team

Alibaba's Qwen-Audio-3.0-TTS Tops Global Leaderboard: 16 Languages, Voice Cloning, and Sub-200ms Latency

Alibaba releases Qwen-Audio-3.0-TTS in Flash and Plus variants, claiming the #1 spot on the Artificial Analysis Speech Arena with an Elo score of 1,234 — surpassing Google, ElevenLabs, and Cartesia in blind listening tests while adding support for 16 languages and 20 Chinese dialects.

On July 14, 2026, Alibaba Cloud released Qwen-Audio-3.0-TTS on its Model Studio platform, marking the company's most ambitious entry into the global text-to-speech market. Available in two variants — Flash for real-time applications with sub-200ms time-to-first-audio and Plus for studio-quality professional use — the model quickly claimed the #1 position on the Artificial Analysis Speech Arena with an Elo score of 1,234, surpassing Google's Gemini 3.1 Flash TTS, ElevenLabs v3, and Cartesia Sonic 3.5 in blind human preference comparisons.

The model supports 16 languages including seven newly added for this release — Arabic, Indonesian, Portuguese, Thai, Vietnamese, Malay, and Tagalog — alongside Chinese, English, Japanese, Korean, and German. It also covers 20 Chinese dialects including Cantonese, Shanghainese, Sichuan, and Northeastern varieties. A standout feature is the 86 inline emotional tags such as [gasp], [giggles], [angry], [whispers], and [sighing] that give developers per-word and per-phrase control over vocal expression — a level of granularity few commercial TTS systems offer. Voice cloning works from as little as 5 seconds of reference audio, with improved noise robustness for imperfect source recordings. The model outputs at up to 48kHz and can synthesize up to 3 minutes of continuous audio in a single pass, eliminating the splice artifacts common in chunked TTS generation.

Pricing positions Qwen-Audio-3.0-TTS competitively: Flash costs approximately $15 per million characters and Plus at $20 per million characters, undercutting ElevenLabs' enterprise tier and roughly matching OpenAI Voice Engine pricing. Both versions include 10,000 free characters for the first 90 days after activation. The model is available exclusively through Alibaba Cloud Model Studio via WebSocket API in Singapore and Beijing regions — no open-source weights have been released, distinguishing it from Meta's Llama-Voice and NetEase Youdao's Confucius4-TTS, both of which published full model weights under Apache 2.0.

The same week, Alibaba also previewed Qwen-Audio-3.0-Realtime, a speech-to-speech conversational model with full-duplex interaction, function calling, and multi-speaker handling. The dual release — one model for TTS quality and one for real-time conversation — signals that Alibaba intends to compete across the full voice AI stack rather than targeting a single niche. With Qwen-Audio-3.0-TTS now topping independent benchmarks, Alibaba joins the small group of companies — alongside Google, OpenAI, ElevenLabs, and Microsoft — whose TTS models can credibly claim best-in-class performance.

AlibabaQwen-AudioTTSArtificial AnalysisVoice CloningMultilingual TTSChinese AI

来源

← Back to News