AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?
We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.
Quick Answer: How Good Is ElevenLabs Music v2?
OUR SHORT VERDICT
ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.
It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.
Best observed strength: lyrics and audible song structure.
Main limitation: fine timing and ending instructions were not always exact.
Test scope: three controlled API runs, not an industry benchmark.
ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.
Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.
ElevenLabs Music v2란 무엇인가요?
ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.
Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.
MODEL ID music_v2
MAXIMUM DURATION 5 minutes
PUBLISHED API RATE $0.15/minute
AUDIO OUTPUT 44.1 kHz
PUBLISHED BITRATE 128–192 kbps
API ACCESS 유료 사용자
ElevenLabs’ official Music v2 release page, checked August 2026.
ElevenLabs Music v2 Test Results
We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.
Test 1: Pop Song With Supplied Lyrics
TEST 01Pop Song With Supplied Lyrics
90초 목표112 BPM미성년자알토 성부
작업 설정: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.
Show the full copyable prompt
따뜻하고 표현력 풍부한 알토 리드 보컬이 돋보이는, 완전히 독창적인 90초 분량의 현대식 신스팝 곡을 하나 만들어 주세요. BPM은 112, 조는 A 단조로 설정하세요. 곡의 분위기는 희망적이고, 친밀하며, 영화 같은 느낌을 주어야 합니다. 뮤트 처리된 일렉트릭 피아노와 부드러운 아날로그 신스 비트로 시작하세요. 첫 번째 절에서는 둥글게 울리는 베이스와 타이트한 일렉트로닉 드럼을 추가하고, 프리코러스에서 긴장감을 고조시킨 뒤, 절제된 배경 하모니가 어우러진 넓고 기억에 남는 코러스로 이어지게 하세요. 리드 보컬은 선명하게 중앙에 위치하도록 하고, 악기 간의 분리감을 명확히 유지하세요. 말로 된 인트로, 랩, 군중 소리, 실제 아티스트의 모방, 페이드 아웃은 포함하지 마세요. 명확한 마지막 코드로 끝맺으세요.
아래의 가사를 정확히 따라 부르세요. 단어나 구절을 추가, 삭제, 대체하거나 순서를 바꾸지 마세요.
[1절]
우리 뒤로 도시의 불빛이 희미해져 가네
이제 고요한 거리가 길을 안내해
한밤중의 속삭임, 뒤를 돌아보지 않아
새로운 지평이 회색을 뚫고 나타난다
[프리 코러스]
손을 맞잡고, 우리는 미지의 길을 걷는다
우리만의 꿈이라 부르는 것을 쫓으며
[코러스]
우리는 고요한 밤 속에 살아 숨 쉬네
스카이라인 너머 별들을 찾아
그림자 속에 휩싸여도, 여전히 빛나네
함께, 우리는 새로운 의미를 찾아낼 거야
[2절]
텅 빈 길 위를 부드럽게 내딛는 발걸음
숨 쉴 때마다 이야기가 펼쳐져
약속은 없지만, 끝없이 펼쳐진 하늘만
네 눈 속에 미래가 깃들어 있어
[Final Chorus]
우리는 고요한 밤 속에서 살아 숨 쉬고 있어
스카이라인 너머 별들을 찾아가며
그림자 속에 휩싸여도, 여전히 빛나
함께, 우리는 새로운 의미를 찾아낼 거야
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0초출력
42.0초GENERATION TIME
$0.225TASK CHARGE
가사 내용에 뚜렷한 일탈이 없음 was detected in the supplied words.
The output had a polished electronic synth-pop character and clean vocal/instrument separation.
The intro felt longer than requested before the vocal arrived.
Test 2: Cinematic Instrumental
TEST 02Cinematic Instrumental
90초 목표82 BPMD 단조보컬 없음
작업 설정: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.
Show the full copyable prompt
극적인 SF 탐사 장면을 위한 완전히 독창적인 90초 분량의 영화적 기악곡을 하나 제작해 주세요. 82 BPM과 D 단조를 사용해 주십시오. 보컬, 대사, 속삭임, 합창, 또는 사람의 목소리와 유사한 소리는 사용해서는 안 됩니다. 희박한 느낌의 피아노 모티프와 낮은 음역의 지속되는 현악기로 조용하게 시작하십시오. 오프닝 이후 미묘한 아날로그 박동을 도입한 뒤, 겹쳐진 현악기와 절제된 저음 타악기를 통해 서서히 분위기를 고조시키십시오. 금관악기는 마지막 3분의 1 구간에만 추가하십시오. 과도한 음량이나 왜곡 없이 끝부분 근처에서 명확한 감정적 절정을 이끌어내십시오. 피아노 모티프가 명확하게 인식되도록 하고, 피아노, 현악기, 타악기, 신스 맥박, 금관악기 간의 명확한 분리감을 유지하십시오. 사실적인 오케스트라 음색을 사용하고, 공간감이 있으면서도 통제된 믹싱을 적용하십시오. 전반부에서는 트레일러 특유의 강렬한 임팩트 사운드를 피하십시오. 기존 악보, 작곡가, 또는 프랜차이즈를 모방하지 마십시오. 페이드아웃 대신 해결된 지속 화음으로 끝맺으십시오.
Test 02 resultA convincing cinematic build, with one early impact
90.0초출력
20.9초GENERATION TIME
$0.225TASK CHARGE
No voice-like texture was detected.
The piano motif remained recognizable as the arrangement built around it.
An impact arrived earlier than the brief allowed, weakening strict timing compliance.
Test 3: Eight-Part Song Structure
TEST 03Eight-Part Song Structure
90초 목표104 BPMD 장조섹션 순서 수정
작업 설정: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.
Show the full copyable prompt
명확하고 표현력 있는 테너 리드 보컬이 돋보이는, 완전히 독창적인 90초 분량의 얼터너티브 팝 곡을 하나 제작하십시오. 104 BPM과 D 장조를 사용하며, 현대적인 얼터너티브 팝 프로덕션 스타일, 자연스러운 발음, 절제된 보컬 이펙트를 적용하십시오. 내레이션, 랩, 실제 아티스트의 모방, 페이드 아웃은 포함하지 마십시오. 정확히 이 순서대로 구성을 유지하고, 모든 프로덕션 변경 사항이 뚜렷하게 들리도록 하십시오: 인트로—약 6초, 필터링된 일렉트릭 기타와 멀리서 들리는 신스만 사용; 절—드라이 베이스와 림 클릭 퍼커션 추가; 프리 코러스—전체 드럼 없이 상승하는 패드 추가; 코러스—전체 드럼, 더 넓게 펼쳐지는 신스, 그리고 선명한 멜로디 후크 도입; 인스트루멘탈 브레이크—약 8초, 보컬 없이 코러스 멜로디를 에코 처리; 브릿지—피아노와 가까이서 녹음된 친밀한 보컬로 축소; 피날레 코러스—풀 밴드 사운드를 복원하고, 백킹 하모니와 옥타브 기타 라인을 추가하며, 첫 번째 코러스보다 더 웅장하게 표현; 아웃트로—드럼을 제거하고, 필터링된 기타 사운드로 돌아간 뒤, 깔끔한 마지막 코드로 마무리한다.
아래 가사를 정확히 따라 부르십시오. 가사를 추가, 삭제, 대체, 순서 변경하거나, 괄호 안에 있는 악기 연주 지침을 가사로 부르지 마십시오.
[인트로]
(악기 연주)
[버스]
갈라진 다이얼이 부서진 선율을 윙윙거린다
손가락이 바랜 선을 따라간다
정전기 소리가 방을 채운다
얽힌 신호들 사이를 헤매며
[프리 코러스]
신호가 깜빡이고, 소리는 불분명해
깊이 있는 목소리를 기다리며
[코러스]
소음 속에서도 내 목소리가 들리나요?
멀리서 들려오는 부름, 아슬아슬한 선택
잡음이 사라지고, 더 선명한 목소리
꽉 붙잡아, 우리가 그 소음을 찾아낼 거야
[악기 연주 브레이크]
(악기 연주)
[브릿지]
부서진 빛 속의 연약한 희망
부드러운 메아리가 나를 끌어당겨
침묵 속에서, 네가 나타난다
[마지막 코러스]
소음 속에서도 내 목소리가 들리나요?
멀리서 들려오는 부름, 아슬아슬한 선택
잡음이 사라지고, 더 선명한 목소리
꽉 붙잡아, 우리가 그 소음을 찾아낼 거야
[아웃트로]
속삭임이 노래로 변해
우리는 마땅히 있어야 할 곳에 있어
Test 03 resultThe clearest structural result, but not perfect timing control
90.0초출력
20.9초GENERATION TIME
$0.225TASK CHARGE
All eight sections were easy to identify, with strong bridge and final-chorus contrast.
The instrumental break ran longer than the requested eight seconds.
The ending behaved more like a fade than the clean final chord in the prompt.
What the Three Tests Show
Performance Findings
3/3Usable outputs
Every run returned a playable 90-second file that fit the broad musical task.
1Clear lyric pass
The supplied pop lyrics showed no clear deviation in our listening checks.
8Audible sections
The fixed-structure song made all requested sections easy to identify.
Vocals: clear, centered, and separated well from the electronic backing in Test 1.
Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.
Efficiency Findings
Generation Time Across Three 90-Second Runs
Pop vocals
42.0초
Instrumental
20.9초
Eight-part song
20.9초
Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.
Total output: 270 seconds across three files.
Total task-reported charge: $0.675.
Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
No reruns: all three completed tasks produced usable audio on the first recorded attempt.
How to Access ElevenLabs Music v2
Access Through ElevenLabs
Use ElevenMusic for browser-based composition.
Use ElevenAPI for programmatic generation and editing.
The official SDK exposes the model through elevenlabs.music.compose 와 함께 model_id="music_v2". A basic request can set a natural-language 프롬프트 그리고 music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.
Generation: create a complete song or instrumental from a prompt.
Detailed response: request more information about the generated composition.
스트리밍: begin receiving audio without waiting for the complete file.
Inpainting: regenerate a selected region instead of replacing the whole track.
Reference matching: guide a generation with reference material where supported.
Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_prompt 또는 bad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.
Example API Request
Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
audio = client.music.compose(
model_id="music_v2",
prompt=(
"Create an original cinematic instrumental with felt piano, "
"low strings, a gradual build, and a resolved final chord."
),
music_length_ms=90_000,
)
save(audio, "music-v2-test.mp3")
Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.
Maximum published output length for one music request.
Our 90-second task charge
$0.225
1.5 minutes × $0.15; reported by each completed test task.
Our three-test total
$0.675
Three 90-second outputs; no rerun cost included.
The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.
How to Prompt ElevenLabs Music v2
Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:
1출력Song or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5제외 사항No speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.
Useful English Music Terms
Click any term or descriptor to copy it. 예시: warm, intimate lead vocal with clear diction.