Resemble AI Releases Chatterbox Nano and Flash: Open-Source TTS Goes Edge-Native with 110M-Parameter Model That Runs on CPU
Technology📅 July 6, 2026👤 FreeReadText Team

Resemble AI Releases Chatterbox Nano and Flash: Open-Source TTS Goes Edge-Native with 110M-Parameter Model That Runs on CPU

Resemble AI ships Chatterbox Nano (110M parameters, runs at 3x realtime on CPU) and Chatterbox Flash (diffusion-LLM architecture for high-throughput production) as MIT-licensed open-source TTS models, alongside multilingual V3 supporting 23+ languages — bringing enterprise-grade voice synthesis to edge devices and on-premise deployments.

In July 2026, Resemble AI released a major expansion of its open-source Chatterbox TTS family, shipping three new models under the MIT license: Chatterbox Nano, Chatterbox Flash, and Chatterbox Multilingual V3. The release marks one of the most comprehensive open-source TTS suites available, spanning the full spectrum from CPU-only edge deployment to high-throughput cloud production, all with built-in audio watermarking for provenance and regulatory compliance.

Chatterbox Nano is the headline achievement for edge deployment. At just 110 million parameters, it achieves 10x realtime inference on GPU and 3x realtime on an 8-thread CPU — meaning a 10-second utterance generates in just over 3 seconds on consumer laptop hardware with no GPU required. The model supports paralinguistic tags including [laugh], [sigh], and [cough] for expressive control, voice cloning from a 5-second reference audio clip, and ships with baked-in audio provenance watermarking by default. Chatterbox Flash targets the other end of the deployment spectrum: rebuilt on a diffusion-LLM (dLLM) architecture, it achieves 2x the throughput of Resemble AI's autoregressive baseline on vLLM using a novel 'prior-subtraction' technique that reduces diffusion steps without compromising output quality, making it suitable for high-volume production TTS workloads.

Chatterbox Multilingual V3 rounds out the release with 500 million parameters, support for 23+ languages, improved speaker similarity scores, and reduced hallucination rates compared to V2. Together, the three models form a deployment gradient that lets developers choose the right trade-off between quality, speed, and hardware requirements — from a Raspberry Pi running Nano for an offline accessibility device to a GPU cluster serving Flash for a million-call-per-day contact center. The MIT license imposes no commercial restrictions, meaning enterprises can fine-tune, modify, and deploy the models in proprietary products without royalty obligations.

The release reinforces a broader industry trend toward open-source voice AI infrastructure. Meta's Llama-Voice (April 2026, Apache 2.0), NetEase Youdao's Confucius4-TTS (June 2026, Apache 2.0), and Fish Audio's open-source models (July 2026) have collectively created a credible open alternative to closed commercial TTS APIs from ElevenLabs, OpenAI, and Google. Resemble AI's contribution is distinctive for its focus on the deployment spectrum — Nano for edge, Flash for throughput, Multilingual V3 for breadth — and its inclusion of watermarking as a default feature, a design choice that anticipates the EU AI Act's mandatory synthetic audio labeling requirements now in force. As one industry observer noted, the Chatterbox family effectively 'open-sources the full TTS deployment playbook' — from edge to cloud, in one consistent model family under one license.

Resemble AIChatterboxOpen Source TTSEdge AIMIT LicenseVoice CloningWatermarking

Источник

← Back to News