The company that taught 100 million people to generate songs just handed them the ability to generate the voice that introduces, narrates, or sells them.

The Summary

  • Suno launched Speech, a new AI feature that generates spoken voiceovers from scripts or prompted descriptions, now in public beta on web and mobile.
  • The feature generates voice and music simultaneously, letting users create podcasts, audiobooks, or narrated content with background tracks in one workflow.
  • CEO Mikey Shulman says Suno has 100 million users and is exploring licensing deals with artists for AI-generated remixes and voice models.
  • The move signals Suno's shift from pure music generation to broader audio creation, positioning it against podcast tools, audiobook platforms, and text-to-speech services.

The Signal

Suno didn't just add a feature. It redrew the map of what its platform is for. Speech, now live in public beta, generates spoken voice from text prompts or scripts and layers it with synchronized background music. This isn't TTS with a Spotify track playing underneath. It's a unified generative workflow where the voice and the score emerge together from the same system.

Chief product officer Jack Brody framed it as an expansion of "human expression," but the real story is business model expansion. Suno built a user base of 100 million people who make songs. Now those same people can make podcast intros, audiobook narration, guided meditations, YouTube voiceovers, sales videos, and anything else that requires a voice and a vibe.

"Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression."

The implications get sharper when you look at Shulman's comments to Bloomberg. He describes AI transforming music from passive listening to active creation, remix, and collaboration. More telling: he mentions partnerships across the music industry and opportunities around licensing artists' voices and music. Translation: Suno is building the rails for artists to license their voice models, letting fans generate new content in their voice, with their sound, under terms the artist controls.

This is the Web4 playbook. Users generate. Artists license. Platforms facilitate. Everyone gets a cut, or at least a say. It's also a direct challenge to the current standoff between generative AI companies and rights holders. Suno is betting that artists will license into the system rather than sue their way out of it.

What sets Speech apart:

  • Generates voice and music in one pass, not two tools duct-taped together
  • Works from both structured scripts and loose natural language prompts
  • Available now on web and mobile, no waitlist or enterprise gating

The timing matters. OpenAI's Advanced Voice Mode showed people what conversational AI sounds like. ElevenLabs proved demand for high-fidelity voice cloning. Suno is threading the middle: not conversational, not cloned, but generative and musical. It's voice-as-content-layer, not voice-as-interface.

The Implication

If you're a creator who's been stitching together voiceover tools, music libraries, and editing software, this collapses your stack. If you're an artist wondering how AI licensing works in practice, watch what Suno builds next. The company is positioning itself as the platform where artists opt in, users create, and everyone's incentives align.

The broader read: we're moving past the "AI will replace X" panic and into the "AI will let anyone do X" phase. Suno isn't trying to replace voice actors or musicians. It's trying to make 100 million people feel like they can be both. Whether that expands the pie or just floods the zone is the open question. But the tools are live, the users are there, and the licensing conversations are happening. This is what the agent economy looks like when it's aimed at culture instead of spreadsheets.

Sources

The Verge AI | Bloomberg Tech