Introducing voice, music, image, and video generation in the ElevenLabs MCP
- Published
ListenListen to this article
The ElevenLabs MCP already lets you manage your agents from the assistants you work in every day. And now, it does more than that - it lets you create. With one install and a single OAuth sign in, your assistant can generate speech, transcripts, dubs, music, sound effects, images, and video, and everything you make lands in your ElevenCreative workspace.
It is live now in Claude, ChatGPT, Cursor, Grokbot, Hermes, and more. One connector covers the whole stack, drawing on over 50 models, with no server to run and no API keys to manage.
One connector for the whole stack
Most creative work spans several tools, and moving between them is where time gets lost. The ElevenLabs MCP brings the models you already rely on into the conversation you are already having. Ask for what you need in plain language and it generates inline, from a single voiceover to a finished video with a score behind it.
Here is what you can create once you connect.

Text to Speech
Ask for a voiceover in any voice from the library and Text to Speech generates it in the conversation, ready to use. There is no export step and no second tab. You describe the line and the delivery you want, and the audio comes back where you asked for it.
Speech to Text with Scribe
Pass in a recording and Scribe returns a clean transcript, complete with speaker labels and timestamps across 99 languages. From there it is a short step to a script, captions, or a dub, without leaving the assistant.
Dubbing
Ask for a track in another language and dubbing runs through the same connector, preserving the original speaker's voice, tone, and delivery. A finished dub comes back in the conversation.
Music and sound effects
Generate a track in any genre or style, with or without vocals, to sit under a piece of video or carry a spot on its own. Sound effects work the same way, from a short prompt to a set of variations you can choose between.
Image and video
Image and video sit alongside the audio tools. Generate an image, edit it, animate it into video, and add lipsync, all from the same connector and all in one thread. The image and video models are the leading ones, and they work next to the voices, music, and sound effects you already trust.
From a brief to a finished piece
The tools are useful on their own, and they are more useful together. A single brief can produce a script, a voiceover, a music bed, sound design, and the video to carry it. You describe the campaign once and assemble the parts in the same place, rather than stitching them together across separate applications afterward.
The same connector that powers ElevenLabs Agents
This is the same MCP that powers ElevenLabs Agents, so a single install covers both. If your team already connected ElevenLabs to review agent performance or spin up new agents, the creative tools are already available to you. There is nothing else to set up.
Everything lands in your ElevenCreative workspace
Every generation lands in your ElevenCreative workspace, where you can open it in Studio to finish the work. Adjust the timeline, refine narration, layer in music and sound effects, and export when it is ready. The conversation gets you to a strong first version quickly, and Studio is there when you want precise control over the final cut.
Workspace admins decide which tools the connection can call, and data residency is selected at connection time, so the people responsible for the account stay in control of what it can access.
Available in Claude, ChatGPT, and Cursor
Installing takes a few clicks. Find ElevenLabs in the connectors directory of Claude, ChatGPT, or Cursor, select connect, and sign in with your ElevenLabs account. Access is scoped to your workspace through OAuth, with no API keys to manage and no manual configuration.
Every node available in Flows is available through the MCP, so the range you can reach from the assistant matches what you can build on the canvas.
Get started
Add the ElevenLabs MCP today and start creating.
.png&w=3840&q=80)


