AI music maker Suno now generates spoken words

The new Speech feature provides synthetic voiceovers embellished with AI background music.
If you buy something from a link, The Verge may earn a commission. See our ethics statement.


Suno is branching out from the world of AI music, launching a new feature that generates spoken voices based on scripts or prompted descriptions. Speech is now available in public beta across Suno’s web and mobile platforms, and allows you to simultaneously generate voiceovers and background music to accompany them.
“Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression,” Suno chief product officer, Jack Brody, said in the announcement. “Today, we’re expanding what’s possible in Suno with Speech: the first audio model that generates voice and music together as one cohesive track.”
AI-generated speech is hardly new — DeepMind has been experimenting with deep learning speech synthesis for a decade, Adobe has a text-to-speech tool, and ElevenLabs has become one of the most recognizable platforms for it since launching in 2023. Suno is just throwing its hat into the ring — likely in an attempt to diversify the platform, given its music generator has attracted so many lawsuits .
Pairing AI music with generated voices is Suno’s spin on text-to-speech tools. It’s optional, meaning you can easily turn off the background music with a toggle if you just want clean speech, but the idea is that it’ll compliment certain use cases for generative spoken word — such as having a calming soundtrack for poems, or something more energetic for dramatic voiceovers and encouraging speeches.
To use the feature, select the “Create” tab, and navigate to the Speech option. There are two modes: Simple, which allows you to describe what you want to create via the provided prompt box (such as “a pirate captain rallying his crew”), or the Advanced mode that lets you add a custom script if you already know exactly what you want it to say. Advanced settings also let you adjust the gender of the AI voice, speech style, and how much variety each voice generation will have. Speech has a maximum duration of around eight minutes.
Suno admits that the feature is far from perfect, but says it’ll keep improving Speech around user feedback. “Beta really does mean beta,” said Brody. “Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic. You will almost certainly discover uses for this that never occurred to us.”
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Image & Video
All news
Image & VideoSean Parker is rebuilding Stability AI around music
Sean Parker, who once taught the music industry what asking for forgiveness looks like, is now back with the labels' blessing and money.
Read at TechCrunch
Image & VideoNacon’s new PS5 controller can mix audio from your phone and console
Nacon announced what the company is claiming is the world's first officially licensed PlayStation 5 controller with a built-in screen for adjusting settings like joystick sensitivity or remapping buttons right on the gamepad. The Revolution 5 Unlimited's screen can also be used…
Read at The Verge
Image & VideoApple’s reportedly developing a smart home camera that doesn’t record video
Apple's rumored push into smart home tech could include a smart home security camera that only gives users text event descriptions instead of video footage. Mark Gurman said in the first episode of the Power On podcast that the camera will be part of "a new Apple smart home ecos…
Read at The Verge
Image & VideoGoogle’s new Guided Vision feature can help you read the fine print
Guided Vision is launching in Gemini Live on compatible Android devices today to use AI to give real-time audio descriptions of anything you point your phone's camera at. By sharing your camera in Gemini Live, you can have Google's AI help with things like reading small text, de…
Read at The Verge