Kakao Upgrades Kanana-o AI Voice: Natural-Language Prompts Now Control Emotion, Accent, and Delivery — Outperforms OpenAI on Korean TTS Benchmark
Technology📅 August 4, 2026👤 FreeReadText Team

Kakao Upgrades Kanana-o AI Voice: Natural-Language Prompts Now Control Emotion, Accent, and Delivery — Outperforms OpenAI on Korean TTS Benchmark

South Korean internet giant Kakao announces a major upgrade to its Kanana-o omni AI model, enabling users to control voice emotion, tone, accent, and delivery style through plain-language text prompts — scoring 94.50 on the Korean InstructTTSEval benchmark, ahead of OpenAI's GPT-4o-mini-tts at 91.10.

On August 4, 2026, South Korean internet giant Kakao announced a major upgrade to the voice generation capabilities of its proprietary omni AI model, Kanana-o. The update enables users to control tone, emotion, accent, speed, and delivery style through simple natural-language text prompts — no SSML markup or coding required. Users can type instructions such as 'read it in a sad voice,' 'read it in a Gyeongsang Province accent,' 'like a sports broadcast,' or compound directions like 'lower the tone and read it quickly in a sad voice,' and the model adjusts speed, volume, pitch, emotion, intonation, and intensity to match the request.

The technical backbone of the upgrade is a new in-house voice tokenizer called LM-SPT (LM-aligned SPeech Tokenizer), which compresses speech into fewer tokens by decoupling linguistic content from speaker identity. This reduces the amount of data the model must process, improving generation speed and naturalness simultaneously. In internal evaluations, LM-SPT outperformed global speech codecs including Mimi, DualCodec, and CosyVoice2 in expert listening tests for both naturalness and speaker similarity. On the Korean InstructTTSEval benchmark — which measures how accurately models follow spoken-delivery instructions — Kanana-o scored 94.50 points, placing it ahead of OpenAI's GPT-4o-mini-tts (91.10) and within competitive range of Google's Gemini-2.5-flash-preview-tts (95.38).

While primarily trained on Korean data, Kanana-o can follow the same style of natural-language voice instructions in English, suggesting Kakao has ambitions beyond the domestic market. The company has stated plans to extend the model with non-verbal expressions — including laughter, sighs, and exclamations — and to unify voice understanding and generation into a single integrated architecture. Kakao has not yet announced a public API or specific service integration timeline, but the upgrade positions Kanana-o as the most advanced Korean-language voice AI model available and among the most expressively controllable TTS systems globally.

The Kanana-o upgrade reflects a broader industry paradigm shift from SSML-based voice control — where developers write markup tags to adjust pitch and speed — to natural-language prompting, a direction also embraced by OpenAI's GPT-4o-mini-tts and Google's Gemini TTS. For the APAC voice AI ecosystem, Kakao's progress carries particular significance: it demonstrates that world-class voice synthesis research is not confined to US West Coast labs, and that models optimized for specific languages and cultural contexts can outperform general-purpose global models on regional benchmarks. As voice interfaces expand across Asian markets with diverse linguistic and cultural expectations, Kakao's investment in Korean-first expressive voice AI positions it as a regional leader in a category where nuance directly determines user adoption.

KakaoKanana-oKorean AIVoice AIEmotion ControlText-to-SpeechLM-SPT

Sursă

← Back to News