Suno Launches Speech: The First Audio Model Generating Voice and Music Together in One Cohesive Track
Technology📅 October 1, 2026👤 FreeReadText Team

Suno Launches Speech: The First Audio Model Generating Voice and Music Together in One Cohesive Track

AI music company Suno opened its Speech feature to public beta on October 1, 2026, letting users type a script or idea and generate spoken audio with original background music in a single end-to-end pass. Simple and Advanced modes cover everything from a one-line prompt to a full custom script, with an optional music toggle and clips up to about eight minutes.

On October 1, 2026, Suno announced Speech, its first product beyond song generation, in a blog post by chief product officer Jack Brody. The company describes it as 'the first audio model that generates voice and music together as one cohesive track' — end-to-end generation rather than the traditional workflow of stitching a text-to-speech output to separately produced background music. The feature opened to public beta across Suno's web and mobile platforms after roughly a month of testing with a small group of users, available from the Create tab.

To use Speech, you type an idea, a poem, or something you've written, then describe the voice and musical style you have in mind. Simple mode takes a short prompt — The Verge's example is 'a pirate captain rallying his crew' — while Advanced mode accepts a custom script with the exact wording plus settings for the AI voice's gender, speech style, and how much variety each generation has. Background music is optional: a toggle switches it off for clean speech-only output, and clips can run up to about eight minutes. Suggested uses include poems with relaxing music, meditations, pep talks, bedtime stories, and dramatized readings of everyday messages.

Suno is candid that the feature is a work in progress. 'Beta really does mean beta,' Brody wrote. 'Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic.' The company says it will keep improving Speech around user feedback. As The Verge noted, AI-generated speech itself is not new — DeepMind, Adobe's text-to-speech tool, and ElevenLabs are established players — but Suno's spin is pairing generated voices with optional music in one workflow, and the launch is widely read as an attempt to diversify a platform whose music generator has attracted numerous copyright lawsuits.

Brody positioned Speech within Suno's broader 'creative entertainment' vision: 'Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression.' The launch lands a month after Suno introduced its v6 music models built in partnership with the music industry, and extends the platform toward narrated storytelling — meditations, children's stories, and dramatized readings — territory that overlaps directly with audiobook and TTS platforms.

SunoSpeech SynthesisAI AudioMusic GenerationVoice GenerationPublic Beta

Kilde

← Back to News