PolyAI Debuts Dialog-RSN-1: An Audio-Native Voice Model That Hears Calls the Way Humans Do, With Sub-300ms Latency
Technology📅 July 30, 2026👤 FreeReadText Team

PolyAI Debuts Dialog-RSN-1: An Audio-Native Voice Model That Hears Calls the Way Humans Do, With Sub-300ms Latency

PolyAI launches Dialog-RSN-1, an audio-native dialog model that reasons directly over raw call audio instead of text transcripts — fusing turn-taking, speech recognition, function calling, and response generation into a single LLM that responds in under 300 milliseconds, outperforming cascaded and speech-to-speech architectures on call center benchmarks.

On July 30, 2026, PolyAI — a London-based enterprise voice AI company — introduced Dialog-RSN-1, an audio-native dialog model that represents a deliberate middle path between traditional cascaded voice architectures and fully end-to-end speech-to-speech systems. The model fuses turn-taking, speech recognition, function calling, and response generation into a single large language model that reasons directly over raw call audio rather than relying on an intermediate text transcript, while handing speech generation off to a separate TTS system that enterprises can control independently.

Dialog-RSN-1 is audio-aware on the input side only — it processes raw audio to understand tone, emotion, hesitations, and environmental context, but uses a separate TTS for output. PolyAI argues this design captures most of the benefits of audio-native models — detecting frustration from vocal tone, catching misrecognitions using conversation context, and handling interruptions naturally — while preserving enterprise control over the output voice, which is critical for brand consistency in contact center deployments. Turn-taking is baked into the model itself: the first predicted token classifies each moment as EMPTY (no speech), ONGOING (caller hasn't finished), or COMPLETE (ready for a reply), eliminating the need for manually tuned silence windows or probability thresholds.

Performance benchmarks position Dialog-RSN-1 as a significant step forward. The model reliably responds in under 300 milliseconds, with a range of 280 to 500ms — compared to 860–1,900ms for OpenAI's GPT Realtime 2.1. On PolyAI's internal Dialog-Eval benchmark, Dialog-RSN-1 achieved the highest overall quality score among real-time-capable models, and it outperformed dedicated ASR models on word error rate for in-call audio. Built via supervised fine-tuning and reinforcement fine-tuning on millions of in-house call center conversations, the model uses open-weight multimodal base architectures — including Gemma, GPT-OSS, Qwen, and Mistral — optimized for A100 GPUs in the 8B dense to 30B sparse parameter range.

Dialog-RSN-1 launches initially in English only, available immediately to existing PolyAI customers with broader language support planned. The launch arrives in an intensely competitive month: Smallest.ai's Voice 4.0 with its async Hydra model, Omilia's Lexis native TTS, and Five9's Voice AI Agents all targeted the enterprise contact center market within weeks of each other. PolyAI's bet is that audio-native understanding — capturing the full richness of a caller's voice rather than reducing it to a transcript — will prove the decisive advantage for complex, emotionally charged customer conversations where tone carries as much meaning as words.

PolyAIDialog-RSN-1Audio-Native AICall CenterVoice AIContact CenterLow Latency

Forrás

← Back to News