Análise do ElevenLabs Music v2: Desempenho real, API e prompt
Chloe Murphy
Última atualização em 10 de agosto de 2026
AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?
We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.
Quick Answer: How Good Is ElevenLabs Music v2?
OUR SHORT VERDICT
ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.
It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.
Best observed strength: lyrics and audible song structure.
Main limitation: fine timing and ending instructions were not always exact.
Test scope: three controlled API runs, not an industry benchmark.
ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.
Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.
O que é o ElevenLabs Music v2?
ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.
Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.
MODEL ID music_v2
MAXIMUM DURATION 5 minutes
PUBLISHED API RATE $0.15/minute
AUDIO OUTPUT 44.1 kHz
PUBLISHED BITRATE 128–192 kbps
API ACCESS Usuários pagos
ElevenLabs’ official Music v2 release page, checked August 2026.
ElevenLabs Music v2 Test Results
We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.
Test 1: Pop Song With Supplied Lyrics
TEST 01Pop Song With Supplied Lyrics
Meta de 90 segundos112 BPMUm menorVoz de contralto
Configuração da tarefa: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.
Show the full copyable prompt
Crie uma música de synth-pop contemporânea totalmente original, com 90 segundos de duração e um vocal principal em tom de contralto, quente e expressivo. Use 112 BPM e a menor. O clima deve ser esperançoso, íntimo e cinematográfico. Comece com um piano elétrico abafado e uma batida suave de sintetizador analógico. Adicione um baixo redondo e uma bateria eletrônica precisa no primeiro verso, crie tensão no pré-refrão e abra caminho para um refrão amplo e memorável com harmonias de fundo contidas. Mantenha a voz principal clara e centralizada, com separação nítida entre os instrumentos. Sem introdução falada, rap, sons de multidão, imitação de um artista real ou fade-out. Termine com um acorde final claro.
Interprete exatamente a letra abaixo. Não adicione, remova, substitua ou reordene palavras ou seções.
[Verso 1]
As luzes da cidade se desvanecem atrás de nós
Ruas silenciosas agora nos guiam
Sussurros da meia-noite, sem olhar para trás
Um novo horizonte rompe o cinza
[Pré-refrão]
De mãos dadas, caminhamos rumo ao desconhecido
Perseguindo sonhos que chamamos de nossos
[Refrão]
Estamos vivos na noite silenciosa
Encontrando estrelas além do horizonte
Perdidos nas sombras, mas tão brilhantes
Juntos, vamos redefinir
[Verso 2]
Passos suaves em ruas vazias
Cada respiração, uma história contada
Sem promessas, apenas céus infinitos
Em seus olhos, o futuro se revela
[Refrão final]
Estamos vivos na noite tranquila
Encontrando estrelas além do horizonte
Perdidos nas sombras, mas tão brilhantes
Juntos, vamos redefinir
Test 01 resultStrong lyric consistency with a longer electronic intro
90,0 sSAÍDA
42,0 sGENERATION TIME
$0.225TASK CHARGE
Não há desvio claro na letra was detected in the supplied words.
The output had a polished electronic synth-pop character and clean vocal/instrument separation.
The intro felt longer than requested before the vocal arrived.
Test 2: Cinematic Instrumental
TEST 02Cinematic Instrumental
Meta de 90 segundos82 BPMRé menorSem vocais
Configuração da tarefa: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.
Show the full copyable prompt
Crie uma trilha instrumental cinematográfica totalmente original de 90 segundos para uma cena dramática de exploração de ficção científica. Use 82 BPM e D menor. Sem vocais, falas, sussurros, coro ou sons semelhantes à voz humana. Comece suavemente com um motivo discreto de piano com feltro e cordas graves sustentadas. Introduza uma batida analógica sutil após a abertura e, em seguida, desenvolva gradualmente com camadas de cordas e percussão grave contida. Adicione metais apenas no terço final. Alcance um clímax emocional claro perto do final, sem volume excessivo ou distorção. Mantenha o motivo de piano reconhecível e preserve uma separação clara entre piano, cordas, percussão, pulso de sintetizador e metais. Use timbres orquestrais realistas com uma mixagem espaçosa, porém controlada. Evite efeitos de impacto típicos de trailers na primeira metade. Não imite nenhuma trilha sonora, compositor ou franquia já existente. Termine com um acorde sustentado e resolvido, em vez de um fade-out.
Test 02 resultA convincing cinematic build, with one early impact
90,0 sSAÍDA
20,9 sGENERATION TIME
$0.225TASK CHARGE
No voice-like texture was detected.
The piano motif remained recognizable as the arrangement built around it.
An impact arrived earlier than the brief allowed, weakening strict timing compliance.
Test 3: Eight-Part Song Structure
TEST 03Eight-Part Song Structure
Meta de 90 segundos104 BPMRé maiorOrdem das seções corrigida
Configuração da tarefa: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.
Show the full copyable prompt
Crie uma música de pop alternativo totalmente original, com duração de 90 segundos, com um vocal principal de tenor claro e expressivo. Use 104 BPM e tom de Ré maior, com produção moderna de pop alternativo, pronúncia natural e efeitos vocais moderados. Sem narração falada, rap, imitação de um artista real ou fade-out. Mantenha exatamente essa ordem das seções e faça com que cada mudança na produção seja audível: Introdução — cerca de 6 segundos, apenas guitarra elétrica com filtro e sintetizador distante; Verso — adicione baixo sem efeito e percussão com batida no aro do tambor; Pré-refrão — adicione um pad crescente sem bateria completa; Refrão — introduza a bateria completa, sintetizadores mais amplos e um gancho melódico claro; Pausa instrumental — cerca de 8 segundos, repita a melodia do refrão sem vocais; Ponte — reduza para piano e um vocal íntimo e próximo; Refrão final — restaure a banda completa, adicione harmonias de apoio e uma linha de guitarra em oitava, e torne-o mais grandioso do que o primeiro refrão; Outro — remova a bateria, volte à guitarra com filtro e termine com um acorde final limpo.
Interprete exatamente a letra abaixo. Não adicione, remova, substitua, reordene nem cante as instruções instrumentais entre parênteses como se fossem letra.
[Intro]
(Instrumental)
[Verso]
Mostradores rachados zumbem uma melodia quebrada
Os dedos traçam as linhas desbotadas
Sussurros estáticos enchem a sala
Procurando entre os sinais emaranhados
[Pré-refrão]
Sinais piscam, som indistinto
Esperando por uma voz profunda
[Refrão]
Você consegue me ouvir em meio ao barulho?
Um chamado distante, uma escolha frágil
A estática se dissipa, uma voz mais clara
Segure firme, vamos encontrar o ruído
[Interlúdio instrumental]
(Instrumental)
[Ponte]
Esperança frágil na luz fragmentada
Ecos suaves, eles me puxam para perto
No silêncio, você aparece
[Refrão final]
Você consegue me ouvir em meio ao barulho?
Um chamado distante, uma escolha frágil
A estática se dissipa, uma voz mais clara
Segure firme, vamos encontrar o barulho
[Outro]
Sussurros se transformam em música
Estamos onde pertencemos
Test 03 resultThe clearest structural result, but not perfect timing control
90,0 sSAÍDA
20,9 sGENERATION TIME
$0.225TASK CHARGE
All eight sections were easy to identify, with strong bridge and final-chorus contrast.
The instrumental break ran longer than the requested eight seconds.
The ending behaved more like a fade than the clean final chord in the prompt.
What the Three Tests Show
Performance Findings
3/3Usable outputs
Every run returned a playable 90-second file that fit the broad musical task.
1Clear lyric pass
The supplied pop lyrics showed no clear deviation in our listening checks.
8Audible sections
The fixed-structure song made all requested sections easy to identify.
Vocals: clear, centered, and separated well from the electronic backing in Test 1.
Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.
Efficiency Findings
Generation Time Across Three 90-Second Runs
Pop vocals
42,0 s
Instrumental
20,9 s
Eight-part song
20,9 s
Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.
Total output: 270 seconds across three files.
Total task-reported charge: $0.675.
Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
No reruns: all three completed tasks produced usable audio on the first recorded attempt.
How to Access ElevenLabs Music v2
Access Through ElevenLabs
Use ElevenMusic for browser-based composition.
Use ElevenAPI for programmatic generation and editing.
The official SDK exposes the model through elevenlabs.music.compose com model_id="music_v2". A basic request can set a natural-language prompt e music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.
Generation: create a complete song or instrumental from a prompt.
Detailed response: request more information about the generated composition.
Transmissão: begin receiving audio without waiting for the complete file.
Inpainting: regenerate a selected region instead of replacing the whole track.
Reference matching: guide a generation with reference material where supported.
Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_prompt ou bad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.
Example API Request
Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
audio = client.music.compose(
model_id="music_v2",
prompt=(
"Create an original cinematic instrumental with felt piano, "
"low strings, a gradual build, and a resolved final chord."
),
music_length_ms=90_000,
)
save(audio, "music-v2-test.mp3")
Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.
Maximum published output length for one music request.
Our 90-second task charge
$0.225
1.5 minutes × $0.15; reported by each completed test task.
Our three-test total
$0.675
Three 90-second outputs; no rerun cost included.
The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.
How to Prompt ElevenLabs Music v2
Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:
1SaídaSong or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5ExclusõesNo speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.
Useful English Music Terms
Click any term or descriptor to copy it. Exemplo: warm, intimate lead vocal with clear diction.