Speechify's Simba 3.2 Tops Global TTS Leaderboard: Consumer-First Voice Model Beats OpenAI, Google, and ElevenLabs in Blind Tests
Technology📅 July 10, 2026👤 FreeReadText Team

Speechify's Simba 3.2 Tops Global TTS Leaderboard: Consumer-First Voice Model Beats OpenAI, Google, and ElevenLabs in Blind Tests

Speechify's Simba 3.2 reaches #1 on the Artificial Analysis TTS leaderboard, beating ElevenLabs, OpenAI, Google DeepMind, and Cartesia in independent blind listening tests — while priced at just $6–10 per million characters, roughly 15x cheaper than rivals in the top ten.

In mid-July 2026, Speechify's Simba 3.2 text-to-speech model claimed the #1 overall position on the Artificial Analysis TTS leaderboard — the industry's most widely cited independent benchmark — with an Elo score of 1,234 based on approximately 1,355 blind evaluations. The model also secured the top spot for real-time models on Voice Arena, a blind-listener benchmark modeled on Chatbot Arena. Speechify CEO Cliff Weitzman announced the achievement on July 10, calling Simba 3.2 'our best model yet' and emphasizing that it was 'built to power voice agents at scale and perfected from millions of A/B tests we run in our consumer platform.'

The benchmark achievement is notable because both Artificial Analysis and Voice Arena use blind pairwise comparisons — native speakers hear pairs of clips and vote on naturalness without knowing which model produced them. Artificial Analysis tests live API endpoints four times daily at random times with random voices and standardized prompts, accepting no self-reported scores. Simba 3.2's Elo of 1,234 placed it at the top of a field that included ElevenLabs v3, Google Gemini 3.1 Flash TTS, Cartesia Sonic 3.5, Inworld Realtime TTS-2, and OpenAI Voice Engine. By late July, Alibaba's Qwen-Audio-3.0-TTS-Plus had edged slightly ahead at 1,236, but Simba 3.2 remained the cheapest model in the top ten by a wide margin.

Weitzman framed the model's advantage as a 'trifecta' of quality, cost, and latency. Simba 3.2 is streaming-native with sub-100ms time-to-first-byte, priced at $10 per million characters on the standard tier and $6 per million on the Scale tier — roughly 15x cheaper than ElevenLabs and 6x cheaper than Cartesia at comparable or better quality. Speechify's consumer platform, which serves millions of users for reading aloud documents, articles, and ebooks, provides a continuous stream of A/B testing data that the company says gives it an edge in optimizing for naturalness — a feedback loop that API-only TTS providers lack.

The moment marks a milestone for consumer-first TTS. Speechify started as a reading accessibility app and has incrementally built competitive voice AI infrastructure in-house. Its rise to the top of independent benchmarks alongside infrastructure giants like Google, Alibaba, and Meta signals that the TTS market is no longer dominated by a handful of cloud providers — and that consumer-scale usage data, combined with aggressive pricing, can produce quality that competes with deep research budgets. Industry analysts note that the speed of leaderboard turnover — with new #1 models appearing nearly every month in 2026 — reflects accelerating TTS innovation and the absence of a durable moat based on model quality alone.

SpeechifySimba 3.2TTS BenchmarkArtificial AnalysisVoice AIConsumer TTS

Fuente

← Back to News