ElevenLabs Music v2 Review: Real Performance, API & Prompt
Chloe Murphy
Last Updated 2026-08-10
AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?
We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.
Quick Answer: How Good Is ElevenLabs Music v2?
OUR SHORT VERDICT
ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.
It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.
Best observed strength: lyrics and audible song structure.
Main limitation: fine timing and ending instructions were not always exact.
Test scope: three controlled API runs, not an industry benchmark.
ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.
Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.
What Is ElevenLabs Music v2?
ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.
Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.
MODEL ID music_v2
MAXIMUM DURATION 5 minutes
PUBLISHED API RATE $0.15/minute
AUDIO OUTPUT 44.1 kHz
PUBLISHED BITRATE 128–192 kbps
API ACCESS Paid users
ElevenLabs’ official Music v2 release page, checked August 2026.
ElevenLabs Music v2 Test Results
We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.
Test 1: Pop Song With Supplied Lyrics
TEST 01Pop Song With Supplied Lyrics
90-second target112 BPMA minorAlto vocal
Task setup: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.
Show the full copyable prompt
Create one completely original 90-second contemporary synth-pop song with a warm, expressive alto lead vocal. Use 112 BPM and A minor. The mood is hopeful, intimate, and cinematic. Begin with muted electric piano and a soft analog-synth pulse. Add round bass and tight electronic drums in the first verse, build tension in the pre-chorus, and open into a wide, memorable chorus with restrained background harmonies. Keep the lead vocal clear and centered with clean instrument separation. No spoken introduction, rap, crowd sounds, imitation of a real artist, or fade-out. End on a clear final chord.
Perform the exact lyrics below. Do not add, remove, replace, or reorder words or sections.
[Verse 1]
City lights fade behind our backs
Silent streets now lead the way
Midnight whispers, no looking back
A new horizon breaks the gray
[Pre-Chorus]
Hand in hand, we step unknown
Chasing dreams we call our own
[Chorus]
We’re alive in the quiet night
Finding stars beyond the skyline
Lost in shadows, yet so bright
Together, we’ll redefine
[Verse 2]
Footsteps soft on empty roads
Every breath a story told
No promises, just endless skies
In your eyes, the future lies
[Final Chorus]
We’re alive in the quiet night
Finding stars beyond the skyline
Lost in shadows, yet so bright
Together, we’ll redefine
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0sOUTPUT
42.0sGENERATION TIME
$0.225TASK CHARGE
No clear lyric deviation was detected in the supplied words.
The output had a polished electronic synth-pop character and clean vocal/instrument separation.
The intro felt longer than requested before the vocal arrived.
Test 2: Cinematic Instrumental
TEST 02Cinematic Instrumental
90-second target82 BPMD minorNo vocals
Task setup: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.
Show the full copyable prompt
Create one completely original 90-second cinematic instrumental cue for a dramatic science-fiction exploration scene. Use 82 BPM and D minor. No vocals, spoken words, whispers, choir, or human voice-like sounds. Begin quietly with a sparse felt-piano motif and low sustained strings. Introduce a subtle analog pulse after the opening, then build gradually with layered strings and restrained low percussion. Add brass only in the final third. Reach one clear emotional climax near the end without excessive loudness or distortion. Keep the piano motif recognizable and preserve clear separation between piano, strings, percussion, synth pulse, and brass. Use realistic orchestral timbres with a spacious but controlled mix. Avoid trailer impacts in the first half. Do not imitate an existing score, composer, or franchise. End on a resolved sustained chord rather than a fade-out.
Test 02 resultA convincing cinematic build, with one early impact
90.0sOUTPUT
20.9sGENERATION TIME
$0.225TASK CHARGE
No voice-like texture was detected.
The piano motif remained recognizable as the arrangement built around it.
An impact arrived earlier than the brief allowed, weakening strict timing compliance.
Test 3: Eight-Part Song Structure
TEST 03Eight-Part Song Structure
90-second target104 BPMD majorFixed section order
Task setup: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.
Show the full copyable prompt
Create one completely original 90-second alternative-pop song with a clear, expressive tenor lead vocal. Use 104 BPM and D major with modern alternative-pop production, natural pronunciation, and restrained vocal effects. No spoken narration, rap, imitation of a real artist, or fade-out. Preserve this exact section order and make every production change audible: Intro—about 6 seconds, filtered electric guitar and distant synth only; Verse—add dry bass and rim-click percussion; Pre-Chorus—add a rising pad without full drums; Chorus—bring in full drums, wider synths, and a clear melodic hook; Instrumental Break—about 8 seconds, echo the chorus melody without vocals; Bridge—reduce to piano and a close intimate vocal; Final Chorus—restore the full band, add backing harmonies and an octave guitar line, and make it larger than the first chorus; Outro—remove drums, return to filtered guitar, and finish on a clean final chord.
Perform the exact lyrics below. Do not add, remove, replace, reorder, or sing the parenthetical instrumental directions as lyrics.
[Intro]
(Instrumental)
[Verse]
Cracked dials hum a broken tune
Fingers trace the faded lines
Static whispers fill the room
Searching through the tangled signs
[Pre-Chorus]
Signals flicker, unclear sound
Waiting for a voice profound
[Chorus]
Can you hear me through the noise?
A distant call, a fragile choice
Static fades, a clearer voice
Hold on tight, we’ll find the noise
[Instrumental Break]
(Instrumental)
[Bridge]
Fragile hope in fractured light
Echoes soft, they pull me near
In the silence, you appear
[Final Chorus]
Can you hear me through the noise?
A distant call, a fragile choice
Static fades, a clearer voice
Hold on tight, we’ll find the noise
[Outro]
Whispers turn to song
We’re where we belong
Test 03 resultThe clearest structural result, but not perfect timing control
90.0sOUTPUT
20.9sGENERATION TIME
$0.225TASK CHARGE
All eight sections were easy to identify, with strong bridge and final-chorus contrast.
The instrumental break ran longer than the requested eight seconds.
The ending behaved more like a fade than the clean final chord in the prompt.
What the Three Tests Show
Performance Findings
3/3Usable outputs
Every run returned a playable 90-second file that fit the broad musical task.
1Clear lyric pass
The supplied pop lyrics showed no clear deviation in our listening checks.
8Audible sections
The fixed-structure song made all requested sections easy to identify.
Vocals: clear, centered, and separated well from the electronic backing in Test 1.
Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.
Efficiency Findings
Generation Time Across Three 90-Second Runs
Pop vocals
42.0s
Instrumental
20.9s
Eight-part song
20.9s
Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.
Total output: 270 seconds across three files.
Total task-reported charge: $0.675.
Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
No reruns: all three completed tasks produced usable audio on the first recorded attempt.
How to Access ElevenLabs Music v2
Access Through ElevenLabs
Use ElevenMusic for browser-based composition.
Use ElevenAPI for programmatic generation and editing.
The official SDK exposes the model through elevenlabs.music.compose with model_id="music_v2". A basic request can set a natural-language prompt and music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.
Generation: create a complete song or instrumental from a prompt.
Detailed response: request more information about the generated composition.
Streaming: begin receiving audio without waiting for the complete file.
Inpainting: regenerate a selected region instead of replacing the whole track.
Reference matching: guide a generation with reference material where supported.
Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_prompt or bad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.
Example API Request
Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
audio = client.music.compose(
model_id="music_v2",
prompt=(
"Create an original cinematic instrumental with felt piano, "
"low strings, a gradual build, and a resolved final chord."
),
music_length_ms=90_000,
)
save(audio, "music-v2-test.mp3")
Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.
Maximum published output length for one music request.
Our 90-second task charge
$0.225
1.5 minutes × $0.15; reported by each completed test task.
Our three-test total
$0.675
Three 90-second outputs; no rerun cost included.
The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.
How to Prompt ElevenLabs Music v2
Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:
1OutputSong or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5ExclusionsNo speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.
Useful English Music Terms
Click any term or descriptor to copy it. Example: warm, intimate lead vocal with clear diction.
Replace the words in brackets. These are starter templates based on the controls that mattered in our tests, not additional test results.
01Lyrics-First Pop SongFor an original vocal song using your exact words
Create one completely original [DURATION]-second cinematic instrumental for [SCENE OR USE]. Use [BPM] BPM and [KEY]. No vocals, spoken words, choir, whispers, or human voice-like sounds. Begin quietly with [MAIN MOTIF] and [LOW FOUNDATION]. Add [PULSE OR TEXTURE] after the opening, then build gradually with [INSTRUMENTS]. Introduce percussion in the middle and brass only in the final third. Reach one clear emotional climax near the end. Keep the main motif recognizable and preserve clean instrument separation. Avoid trailer impacts in the first half. End on a resolved sustained chord rather than a fade-out.
03Background Music BedFor videos, podcasts, product demos, or cafés