Inworld Launches Realtime TTS-2: Conversation-Aware Voice Model Tops the Artificial Analysis Voice Arena
Technology📅 August 31, 2026👤 FreeReadText Team

Inworld Launches Realtime TTS-2: Conversation-Aware Voice Model Tops the Artificial Analysis Voice Arena

Inworld AI has released Realtime TTS-2, a voice model built specifically for realtime conversation that hears the full audio of an exchange and takes delivery direction in plain English. The launch lands the company at #1 on the Artificial Analysis Controlled Voice Arena with an Elo score of 1123, four points ahead of Cartesia Sonic 3.6.

On August 31, 2026, Inworld AI announced the general availability of Realtime TTS-2, the successor to its Realtime TTS 1.5 model and the first voice model the company built from the ground up for realtime conversation rather than narration. The model hears the full audio of the exchange — the user's tone, pacing, and emotional state — and takes voice direction in plain English the way developers prompt an LLM, while holding one voice identity across over 100 languages. It ships through the Inworld API and the Inworld Realtime API, and customers on Realtime TTS 1.5 upgrade by changing the model identifier with no other code changes.

Four capabilities anchor the launch. Voice Direction accepts a natural-language description of how a line should be delivered — such as 'tired but warm, like she just got home' — with inline non-verbal markers like [laugh], [sigh], [breathe], [clear_throat], and [cough] rendered as audio events rather than spoken words. Conversational Awareness conditions each response on the actual audio of prior turns, so the same line lands differently after a joke than after bad news. Crosslingual generation preserves one speaker identity across more than 100 languages, including mid-utterance switches. Advanced Voice Design generates a saved voice from a written prompt with no reference audio, with Expressive, Balanced, and Stable stability modes. Voice cloning from 5–15 seconds of authorized reference audio is available through a two-step API, and median time-to-first-audio for the TTS layer is sub-200 milliseconds.

On the Artificial Analysis Controlled Voice Arena, Realtime TTS-2 ranks first with an Elo score of 1123, four points above Cartesia Sonic 3.6's 1119, in a blind test where different models clone the same set of eight voices. Inworld's previous model, Realtime TTS 1.5, already held the #1 spot on the Artificial Analysis Speech Arena ahead of Google and ElevenLabs, and the company's comparison table positions Realtime TTS-2 as the only model combining multi-turn-aware speech, advanced free-form voice direction, crosslingual identity across 100+ languages, and voice profiling in a single stack. The Realtime API speaks the OpenAI Realtime protocol with Inworld extensions, so existing OpenAI Realtime clients can connect with a single URL change, and the model is available through partners including Cloudflare, DeepInfra, GMI Cloud, LiveKit, Stream, and VoiceRun.

Early partners emphasized expressiveness over raw benchmarks. 'Inworld's TTS-2 marks a real step forward in emotionally expressive voice synthesis,' said LiveKit co-founder and CTO David Zhao, while Latitude CEO Nick Walton called it 'a significant advance' for AI-native games with emotionally complex characters. CEO Kylan Gibbs framed the company's thesis simply: 'We are obsessed with how Voice AI feels, not just how it sounds.' Pricing is metered per character — roughly 1,000 characters per minute of audio — under the same metering as Realtime TTS 1.5, with pay-as-you-go volume tiers.

InworldRealtime TTS-2Voice AIConversational AIText-to-SpeechVoice CloningArtificial Analysis

المصدر

← Back to News