Every AI music tool of the last two years has been able to generate a voice. The hard part was never making one sing. It was being able to use the result without a lawyer in the room.
Here is the direct answer. On July 23, 2026, ElevenLabs updated ElevenMusic with three things: Vocals, which generates an original song sung in a voice you choose, including your own; References, which lets you set a track's style and mood from a piece of reference audio; and an update bringing Styles onto the Music v2 model so a track holds its character the whole way through. Everything you upload, voice or song, is screened for copyright compliance before it generates anything. That last clause is doing more work than the other three combined.
What Actually Shipped
Three features arrived together and they are easy to confuse, partly because they overlap and partly because ElevenLabs already had a feature called Finetunes doing something adjacent. Here is the clean version.
| Feature | What it does | What you give it |
|---|---|---|
| Vocals | Generates an original song sung in a chosen voice | Your own voice, or one from the voice library |
| References | Sets the style and mood of a generated track | A short reference audio clip |
| Styles | Holds a consistent sound across a full track, now on Music v2 | A style you define or pick |
| Finetunes (existing) | Trains a custom version of the music model on your catalogue | Up to 50 tracks, 250 minutes of audio total |
The distinction that matters in practice: References and Styles shape a sound, Finetunes builds you a model. A reference clip nudges one generation. A Finetune produces a version of the music model that carries your tone, texture, and arrangement choices across everything you generate afterwards, which is the difference between a filter and an instrument.
ElevenLabs is explicit that a reference does not copy or remix the audio you upload. It influences production style, instrumentation, tempo, and mood. That is a deliberate architectural choice and, again, a legal one.
One wrinkle if you go looking for the spec. References is the general-availability version of the reference-audio feature that was previewed when Music v2 was announced, and the numbers have moved: launch coverage puts the usable clip range at 10 seconds to 5 minutes, while the older product documentation still describes a reference as a short track of up to about 30 seconds. Treat the longer range as current and the docs as catching up, but test before you build a pipeline that assumes either.
The Part Everyone Skips: Consent Is the Product
It is tempting to read "AI clones your singing voice" as one more escalation in a technology that has spent two years being sued. The more useful reading is that ElevenLabs has spent those two years building the opposite posture, and Vocals is what that posture makes possible.
Three pieces of evidence, in ascending order of how much they cost the company.
One: the model is trained on licensed data. Music v2 is trained only on licensed material and cleared for commercial use, which is why output comes with no sync fees and no clearance delays. That is a straightforward commercial promise, and it is also the entire reason a brand's legal team will let you use it.
Two: uploads are screened before they are used. Every reference is checked for copyright compliance, and for Finetunes the screening is run by a third-party compliance system before training starts. You are not permitted to upload copyrighted music you do not own, and if your upload is rejected, you do not get a refund. That is a friction the company chose to add.
Three: they proved it with real artists, six months before this launch. In January 2026 ElevenLabs released The Eleven Album on Spotify, billed as the first large-scale multi-artist AI collaboration built on a rights-secure framework, with Liza Minnelli, Art Garfunkel, and Michael Feinstein among the participants. The artists retain full ownership, commercial rights, and all streaming revenue. Some tracks use an AI vocal likeness, some were sung the traditional way; Minnelli's track uses her real voice rather than a recreation.
Why this matters commercially. The competitive question in AI music stopped being "whose output sounds best" some time ago. It is now "whose output can I put behind a paid ad, in a client deliverable, or on a storefront, without discovering in eight months that I cannot." Licensed training data plus pre-generation screening is a slower, more expensive way to build. It is also the only version that survives contact with a procurement process.
What You Can Actually Do With It
Four uses that are real rather than demo-shaped, roughly in order of how quickly they pay for themselves.
- 01A consistent sonic identity across content. Not one good track, the same recognisable sound across fifty. This is what Finetunes are for, and it is the use case most people miss because they are busy generating one-off songs.
- 02Scratch vocals and demos. Hearing a topline in a voice close to the intended one, before booking a session, collapses the slowest loop in songwriting from days to minutes.
- 03Localised versions of the same track. Music v2 generates vocals across multiple languages without the pronunciation artefacts that make most AI singing unusable outside English.
- 04Programmatic catalogue work. The Finetunes API landed on July 20 with five endpoints for creating, listing, checking status, updating, and deleting a Finetune, and a
finetune_idparameter on the generation methods. That turns a creative tool into infrastructure.
The Limits, Before You Plan Around It
Specifications people find out about after they have committed.
| Constraint | Value |
|---|---|
| Generation length | 3 seconds minimum, 5 minutes maximum |
| Finetune input | Up to 50 tracks, each 10 to 600 seconds, 30MB per track |
| Finetune total audio | 250 minutes maximum |
| Finetune build time | Roughly 5 to 10 minutes |
| Export quality | MP3 at 44.1kHz, 128 to 192kbps; studio-grade on higher tiers |
| Rejected upload | No refund |
The five-minute ceiling and the MP3 default are the two that catch people. If you are scoring long-form video or delivering masters to a client who expects lossless, check your tier before you promise anything. And the no-refund rule on rejected uploads means the copyright screen is not a formality you can iterate against cheaply.
Where This Leaves the Competition
License first, ship second
Slower to market, narrower training data, and a compliance step that costs users money when it fails. In exchange: output that clears legal, and artists willing to lend their names.
Ship first, settle later
Faster, broader, and cheaper to build, with the licensing question deferred into litigation and negotiated afterwards. Fine for a demo. Harder to build a client business on.
I do not think the second approach is illegitimate, and plenty of excellent tools sit in that column. But if your work involves shipping audio on behalf of someone else, the column you sit in is not an aesthetic preference. It is a risk you are quietly transferring to your client, and most of them have not been told.
The Honest Part
Four things worth saying plainly.
Licensed does not mean risk-free. "Trained on licensed data" addresses the training set. It does not resolve every question about a generated output that happens to resemble something. It is a much better position than the alternative and it is not a legal opinion.
Your own voice is still a permission question. Vocals makes it trivial to generate singing in a cloned voice. Whose voice, and with what documented consent, is your problem and not the platform's. If you are cloning anyone other than yourself, get it in writing before you generate, not before you publish.
Screening is a filter, not a guarantee. An automated third-party compliance check catches the obvious. It is not an indemnity, and the no-refund policy means the cost of a false positive lands on you.
The demo will beat your real output. This is true of every generative tool and it is especially true of music, where the gap between "impressive in a fifteen-second clip" and "holds up across a full track under repeat listening" is enormous. Generate ten, not one, before you form a view.
The interesting engineering here is not the voice. It is that someone built the consent step into the pipeline before the generate button, where it is expensive, instead of after it, where it is merely someone else's problem.
If You Only Do Three Things, Do These
- 01Generate the same brief ten times before judging it. One clip tells you nothing about consistency, which is the only property that matters for real use.
- 02If you want a sound rather than a song, build a Finetune, not a prompt. Prompts give you variety. Finetunes give you identity.
- 03Write down whose voice you are cloning and who agreed to it, before you generate. The platform screens copyright, not your permissions.
The broader pattern is worth noticing. The models that win commercial adoption are increasingly not the most capable ones, they are the ones whose provenance can be explained to a legal team in one sentence. That was true of open weights this month, and it is true of music now. Capability is becoming table stakes. Provenance is becoming the moat.
Questions People Ask
What is ElevenLabs Vocals?
A feature added to ElevenMusic on July 23, 2026 that generates a completely original song sung in a voice you choose. You can use your own voice or select one from ElevenLabs' voice library, and it is designed to be usable by people with no music production experience.
What is the difference between References, Styles, and Finetunes?
References let you upload an audio clip to set the style and mood of a generation, influencing instrumentation, tempo, and production without copying the clip. Styles keep a consistent character across a full track and now run on Music v2. Finetunes are heavier: you upload up to 50 tracks and get a custom version of the music model that carries your sound across every future generation.
Can I use ElevenLabs music commercially?
Yes. Music v2 is trained only on licensed data and cleared for commercial use, with no sync fees or clearance delays on the output. That covers the training set and the licence to use what you generate; it is not a substitute for legal advice on a specific release.
Does ElevenLabs check what I upload for copyright?
Yes, and before generation rather than after. Uploaded voices and reference tracks are screened for copyright compliance, and Finetune uploads go through a third-party compliance system before training begins. You must not upload copyrighted music you do not own, and a rejected upload is not refunded.
How long can an ElevenLabs generated track be?
Between 3 seconds and 5 minutes per generation. Finetune source tracks must each be 10 to 600 seconds and under 30MB, with a 250 minute total across a maximum of 50 tracks. Output is MP3 at 44.1kHz and 128 to 192kbps, with studio-grade exports on higher tiers.
Is there an API for this?
Partly. A Music Finetunes API shipped on July 20, 2026 with five endpoints covering create, list, status, update, and delete, plus a finetune_id parameter on the music generation methods. Finetunes were available in the ElevenCreative interface at launch, with broader API access on request for enterprise customers.
Working out where generative audio fits in what you are building?
I help founders and brands build generative AI systems that ship: model routing, agentic workflows, coding-agent setups, and content engines that are production-grade down to the plumbing. If you want a second pair of eyes on yours, book a session.
Book a Growth ChatWritten July 26, 2026, three days after the ElevenMusic update. The July 23 feature set is from ElevenLabs' own announcement and from Digital Today's report of it. Music v2's capabilities, licensing position, and launch date are from ElevenLabs' introduction post, published May 26 and updated July 22. Finetune limits, copyright screening, export formats, and generation length come from ElevenLabs' product documentation. The Finetunes API details are from the July 20 release notes. The Eleven Album details, including the January 21 release, artist participation, and the retention of ownership and streaming revenue, are from ElevenLabs' announcement and coverage in Variety, Billboard, NBC News, Music Ally, and Deadline. Two caveats on precision. The reference-clip duration is genuinely inconsistent between the launch coverage and the current documentation, and I have flagged that above rather than pick one. And one claim circulating in social coverage of this launch, that a single recording can be one-shot into a finished vocal, I could not corroborate against ElevenLabs' documentation or any independent source, so it is not asserted anywhere in this post. Nothing here is sponsored, and I have no relationship with ElevenLabs.