Cartesia Launches Sonic 3.5 and Ink 2: SSM Architecture Makes It the First Provider to Top Both TTS and STT Leaderboards
Technology📅 June 17, 2026👤 FreeReadText Team

Cartesia Launches Sonic 3.5 and Ink 2: SSM Architecture Makes It the First Provider to Top Both TTS and STT Leaderboards

Cartesia releases Sonic 3.5 (TTS) and Ink 2 (STT) built on State Space Model architecture rather than Transformers, achieving sub-90ms latency and 42-language support — becoming the only provider to simultaneously hold the #1 ranking for both speaking and listening on Artificial Analysis benchmarks.

On June 17, 2026, Cartesia publicly launched Sonic 3.5 and Ink 2, a paired text-to-speech and speech-to-text model suite built on State Space Model (SSM) architecture rather than the Transformer architecture used by virtually every competing voice AI system. The launch made Cartesia the only provider to simultaneously hold the #1 position for both TTS naturalness and STT accuracy on the Artificial Analysis leaderboards — a dual-leaderboard achievement no other company has matched. Sonic 3.5 replaced the now-retired Sonic Turbo as Cartesia's flagship TTS model, available through Cartesia's API and the Together AI serverless platform.

The SSM architecture is the headline technical story. While Transformers scale quadratically with context length due to attention mechanisms, SSMs — stemming from the S4 and Mamba research lineage created by Cartesia co-founders Albert Gu and Chris Ré at Stanford's AI Lab — process audio as a continuous stream with constant memory consumption. This design eliminates the growing KV cache that burdens Transformer-based TTS during long-running conversations, enabling sub-90ms time-to-first-audio latency and efficient on-device deployment on laptops and mobile hardware. A multi-stream SSM design separates text conditioning and audio generation into distinct state representations, further improving both quality and throughput. Sonic 3.5 also natively handles alphanumerics — order numbers, phone numbers, confirmation codes — without preprocessing, alongside context-aware English heteronym pronunciation.

The model supports 42 languages at native quality, spanning English, Hindi, Spanish, French, German, Japanese, Chinese, Korean, Portuguese, Italian, Dutch, Polish, Russian, Arabic, Hebrew, and 27 more. Ink 2, the speech-to-text counterpart, was released alongside Sonic 3.5 to form a complete real-time voice AI stack that developers can access through a single unified API. The combined offering targets conversational AI platforms, voice agents, and real-time transcription use cases — a market where sub-100ms latency and high accuracy on both sides of the conversation are increasingly table stakes. Pricing is credit-based with tiers starting at $5 per month for Pro and $299 per month for Scale.

Cartesia's trajectory reflects the growing conviction that architectural diversity — not just model scale — will define the next phase of voice AI. The company had raised approximately $191 million by late 2025 from investors including Kleiner Perkins, Index Ventures, Lightspeed, and NVIDIA, and counts ServiceNow, Decagon, Zomato, and Retell AI among its customers. While Transformer-based models from OpenAI, Google, and ElevenLabs continue to dominate voice AI headlines, Cartesia's dual-leaderboard achievement with a fundamentally different architecture signals that the SSM approach — long considered a research curiosity — is now production-competitive at the highest tier of commercial voice AI.

CartesiaSonic 3.5Ink 2State Space ModelSSMReal-Time TTSSpeech-to-Text

Sumber

← Back to News