Deepdub Launches Phantom Z 3.4 Conversational: Enterprise Multilingual TTS Built for Live Voice Agents
Technology📅 September 5, 2026👤 FreeReadText Team

Deepdub Launches Phantom Z 3.4 Conversational: Enterprise Multilingual TTS Built for Live Voice Agents

Deepdub has released Phantom Z 3.4 Conversational, an enterprise-grade multilingual text-to-speech model delivering 150-millisecond time-to-first-audio at full 48 kHz fidelity, with production-grade text normalization and extended Hebrew support. The model targets voice agents and contact centers where misread account numbers and dates decide whether a call is contained or escalated.

On September 3, 2026, voice AI company Deepdub announced Phantom Z 3.4 Conversational, a new multilingual text-to-speech model designed for production voice agents and contact centers, available to all Deepdub clients immediately. The release centers on reliability rather than demo polish: an end-to-end p95 time-to-first-audio of 150 milliseconds in real-time mode at full-range 48 kHz audio, and cross-language voice transfer from under three seconds of reference audio. Deepdub builds and trains its own speech models in-house rather than licensing them, which the company says lets it bring a new language into production in about two weeks.

The headline technical work is in text normalization, the step that converts written text into spoken words. A delivery date written 2024-12-31 is read as 'December thirty first, twenty twenty-four,' an invoice total of $1,240 as 'one thousand two hundred forty dollars,' and an appointment at 14:30 as 'two thirty.' Deepdub says its testing puts the model ahead of the other systems it measured against on these categories. The release also extends Hebrew support significantly: because Hebrew is written without vowels, the same letters can spell different words, so Phantom Z 3.4 resolves pronunciation from the whole sentence rather than word by word. Deepdub reports that every instance of the ambiguous word שלט in its Hebrew test set was read correctly, that the model ranks first for Hebrew text-to-speech on the ivrit.ai TTS Arena leaderboard hosted on Hugging Face, and that it was preferred over the company's previous Hebrew model in 71 percent of decisive blind listening comparisons.

The model covers more than 50 locales and dialects verified by local voice and language experts. CEO and co-founder Ofir Krakowski framed the positioning bluntly: 'Every voice model sounds impressive for two minutes in a demo. Very few survive two weeks with real customers.' He added that deployments don't stall on the 95 percent a model gets right — 'they stall on the misread account number, the mangled surname, the one wrong digit on a live call. We built this model for that last few percent, because in production, the last few percent is the whole product.' The launch lands as the voice-agent market consolidates around production metrics such as containment rate and cost per resolved call, pushing TTS vendors to compete on numeric accuracy and latency rather than raw naturalness.

Early production customers backed the claims. Adir Haziza, CTO at Voiceman, said the company runs Deepdub for live, real-time phone calls 'where latency and naturalness aren't nice-to-haves,' and that with 3.4 'our callers show it: they stay longer, talk more, and engage with our agents like we've never seen before.' Dor Levy, Head of Jeen Talk at Jeen AI, said the details came out right and the Hebrew was the most natural his team had heard. For enterprises running voice agents across multiple languages, Deepdub's pitch is a single model with one set of behaviors to test — and one contract to hold.

DeepdubPhantom Z 3.4Text-to-SpeechVoice AgentsMultilingual TTSHebrew TTSText Normalization

Forrás

← Back to News