Deepgram Launches Flux TTS: First Conversation-Native Text-to-Speech Built for Real-Time Voice Agents
Technology📅 August 12, 2026👤 FreeReadText Team

Deepgram Launches Flux TTS: First Conversation-Native Text-to-Speech Built for Real-Time Voice Agents

Deepgram releases Flux TTS, a conversation-native text-to-speech model that retains conversational state across entire calls, handles barge-in interruptions natively, and starts speaking in as little as 80 milliseconds — free until September 12.

In August 2026, Deepgram launched Flux TTS, the second model in its Flux family and what the company calls 'the first conversation-native TTS built for real-time voice agents.' The announcement, published on Deepgram's blog by product marketing lead Yael Scarlett and product manager Jeff Liu, follows the October 2025 debut of Flux STT, which reimagined speech-to-text for live conversations, and the April 2026 Flux Multilingual release, which extended the architecture to 10 languages. Flux TTS is free until September 12 to encourage developers to build with it.

The core design departure is that Flux TTS holds the whole conversation in memory as it speaks rather than starting fresh on every line. Tone, pacing, pronunciation, and emotional register stay consistent across an entire call without developers hand-feeding context through SSML, style tags, or prompt engineering. The model handles interruptions natively: when a caller barges in, the server reports exactly what the caller heard via text_spoken and text_remaining, so the agent resumes cleanly instead of guessing. Time to first audio is as low as 80 milliseconds, and on the inputs that break production voice agents — account numbers, drug names, alphanumeric identifiers, currency, and dates — Flux TTS reports a median word error rate of 2.2%, roughly half of ElevenLabs and a third of Cartesia, and 3.4% on hard prompts, beating the next-best model by 47%.

Three architectural choices underpin the model: a high-fidelity neural codec from the same family as recent research from Kyutai, Meta, and Google; interleaved text-to-audio generation that keeps first-audio latency under 200ms regardless of response length; and a Mamba state-space backbone with fixed-size memory that holds context across a whole session without the quadratic cost of transformers. The model launches with conversational English voices, with additional languages, voice cloning, and finer emotional controls on the roadmap. It deploys in the cloud, self-hosted, or on-prem, and ships with native integrations for voice agent platforms including Vapi, Pipecat, LiveKit, Jambonz, and Cloudflare.

Early partners frame the release as closing the gap between demo-quality TTS and production voice agents. Kwindla Hultman Kramer, CEO of Daily, said Flux TTS makes TTS 'part of the agent's conversational architecture rather than a layer developers have to tune line by line,' while jambonz founder Dave Horton highlighted how conversation-aware speech removes the need to 'hand-tune every turn.' According to Deepgram CEO Scott Stephenson, IBM and Coval are two examples of how Flux TTS is being used across the voice AI ecosystem, with the IBM integration extending the technology into watsonx Orchestrate. The same week saw competitors like Soniox ship a multilingual TTS update with 60 languages and voice cloning — Deepgram's counter-bet is that conversation-native design, not raw language count, is what production voice agents actually need.

DeepgramFlux TTSVoice AgentsConversational AIText-to-SpeechLow LatencyMamba

Forrás

← Back to News