ElevenLabs Music v2 评测:实际表现、API 及提示词

elevenlabs-music-v2-评测-hero-3x2-v2

AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?

We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.

Quick Answer: How Good Is ElevenLabs Music v2?

OUR SHORT VERDICT

ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.

It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.

  • Best observed strength: lyrics and audible song structure.
  • Main limitation: fine timing and ending instructions were not always exact.
  • Test scope: three controlled API runs, not an industry benchmark.

ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.

ElevenLabs Music v2 Overview

Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.

什么是 ElevenLabs Music v2?

ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.

Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.

MODEL ID
music_v2
MAXIMUM DURATION
5 minutes
PUBLISHED API RATE
$0.15/minute
AUDIO OUTPUT
44.1 kHz
PUBLISHED BITRATE
128–192 kbps
API ACCESS
付费用户
ElevenLabs 官方《Introducing Music v2》发布页面
ElevenLabs’ official Music v2 release page, checked August 2026.

ElevenLabs Music v2 Test Results

We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.

Test 1: Pop Song With Supplied Lyrics

TEST 01Pop Song With Supplied Lyrics
90秒目标112 BPM未成年人女高音

任务设置: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.

Show the full copyable prompt
创作一首完全原创的90秒当代合成器流行歌曲,主唱为温暖且富有表现力的中音女声。采用112 BPM和A小调。整体氛围应充满希望、亲密且富有电影感。开场以弱音电钢琴和柔和的模拟合成器脉冲声开始。 在第一段歌词中加入圆润的贝斯和紧凑的电子鼓点,在前副歌部分营造张力,随后过渡到宽广且令人难忘的副歌,并配以克制的背景和声。保持主唱声音清晰且居中,乐器声部分离分明。不得包含说唱式引子、说唱、人群声、模仿真实艺人的演唱,也不得使用淡出效果。 以一个清晰的终结和弦收尾。

请严格按照以下歌词演唱。不得添加、删除、替换或调整词语或段落的顺序。

[第一段]
城市灯光在我们身后渐渐淡去
寂静的街道如今指引着方向
午夜的低语,不再回首
新的地平线划破灰蒙
[前副歌]
手牵手,我们踏上未知的旅程
追寻属于我们自己的梦想
[副歌]
在这寂静的夜里,我们活着
在天际线彼端寻觅星辰
虽迷失在阴影中,却依然璀璨
携手同行,我们将重新定义
[第二段]
空荡的道路上,脚步轻柔
每一口气,都诉说着一个故事
没有承诺,只有无垠的天空
你的眼中,未来正等待
[终结副歌]
在这静谧的夜里,我们生机勃勃
在天际线彼端寻觅星辰
虽迷失在阴影中,却依然璀璨
携手同行,我们将重新定义
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0秒OUTPUT
42.0秒GENERATION TIME
$0.225TASK CHARGE
  • 歌词没有明显的偏差 was detected in the supplied words.
  • The output had a polished electronic synth-pop character and clean vocal/instrument separation.
  • The intro felt longer than requested before the vocal arrived.

Test 2: Cinematic Instrumental

TEST 02Cinematic Instrumental
90秒目标82 BPMD小调无人声

任务设置: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.

Show the full copyable prompt
为一段充满戏剧张力的科幻探索场景创作一段完全原创的90秒电影风格器乐配乐。节奏为82 BPM,调性为D小调。不得包含人声、白话、低语、合唱或任何类似人声的声音。 以稀疏的毛毡钢琴主题和低音弦乐的持续音静谧开场。开场后引入微妙的模拟脉冲,随后通过分层弦乐和克制的低音打击乐逐渐铺陈。仅在最后三分之一处加入铜管乐器。在接近结尾处达到一个清晰的情感高潮,避免过度的音量或失真。 保持钢琴主题的辨识度,并确保钢琴、弦乐、打击乐、合成器脉冲和铜管乐器之间保持清晰的分离度。使用真实的管弦乐音色,混音效果应开阔但受控。前半部分避免使用预告片式的突击音效。不要模仿现有的配乐、作曲家或影视系列。结尾应以一个解决后的持续和弦收尾,而非淡出。.
Test 02 resultA convincing cinematic build, with one early impact
90.0秒OUTPUT
20.9秒GENERATION TIME
$0.225TASK CHARGE
  • No voice-like texture was detected.
  • The piano motif remained recognizable as the arrangement built around it.
  • An impact arrived earlier than the brief allowed, weakening strict timing compliance.

Test 3: Eight-Part Song Structure

TEST 03Eight-Part Song Structure
90秒目标104 BPMD大调固定章节顺序

任务设置: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.

Show the full copyable prompt
创作一首完全原创的90秒另类流行歌曲,主唱为音色清晰、富有表现力的男高音。采用104 BPM和D大调,以现代另类流行风格制作,发音自然,人声效果克制。不得包含旁白、说唱、模仿真实艺人或淡出处理。 请严格保留以下段落顺序,并确保每次制作变化都能被清晰听出:前奏——约6秒,仅含经过滤波处理的电吉他和远处的合成器;主歌——加入干声贝斯和鼓边敲击声;前副歌——加入渐强的垫音,但不使用完整鼓组;副歌——加入完整鼓组、更宽广的合成器音色以及清晰的旋律钩子; 器乐间奏——约8秒,无人声,回响副歌旋律;过渡段——简化为钢琴伴奏和近距离亲密的人声;最终副歌——恢复完整乐队编曲,加入和声及八度吉他旋律线,并使其气势比第一个副歌更宏大; 尾奏——移除鼓组,回归经过滤波处理的吉他声,并以一个清澈的终结和弦收尾。

请严格按照以下歌词演唱。不得添加、删除、替换、调整顺序,也不得将括号内的器乐指示当作歌词演唱。

[前奏]
(器乐)
[主歌]
龟裂的表盘哼着破碎的旋律
指尖轻抚褪色的线条
静电的低语充盈整个房间
在错综复杂的征兆中搜寻
[前副歌]
信号闪烁,声音模糊
等待着那深邃的声音
[副歌]
你能穿过噪音听见我吗?
遥远的呼唤,脆弱的选择
杂音渐消,声音愈发清晰
紧握不放,我们会找到那片喧嚣
[器乐间奏]
(器乐)
[过渡段]
破碎光线中的脆弱希望
柔和的回声,将我拉近
在寂静中,你悄然出现
[终曲副歌]
你能穿过喧嚣听见我吗?
遥远的呼唤,脆弱的选择
杂音渐消,声音愈发清晰
紧握不放,我们会找到那喧嚣
[尾声]
低语化作歌声
我们回到了该在的地方
Test 03 resultThe clearest structural result, but not perfect timing control
90.0秒OUTPUT
20.9秒GENERATION TIME
$0.225TASK CHARGE
  • All eight sections were easy to identify, with strong bridge and final-chorus contrast.
  • The instrumental break ran longer than the requested eight seconds.
  • The ending behaved more like a fade than the clean final chord in the prompt.

What the Three Tests Show

Performance Findings

3/3Usable outputs

Every run returned a playable 90-second file that fit the broad musical task.

1Clear lyric pass

The supplied pop lyrics showed no clear deviation in our listening checks.

8Audible sections

The fixed-structure song made all requested sections easy to identify.

  • Vocals: clear, centered, and separated well from the electronic backing in Test 1.
  • Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
  • Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
  • Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.

Efficiency Findings

Generation Time Across Three 90-Second Runs
Pop vocals
42.0秒
Instrumental
20.9秒
Eight-part song
20.9秒

Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.

  • Total output: 270 seconds across three files.
  • Total task-reported charge: $0.675.
  • Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
  • No reruns: all three completed tasks produced usable audio on the first recorded attempt.

How to Access ElevenLabs Music v2

Access Through ElevenLabs

  • Use ElevenMusic for browser-based composition.
  • Use ElevenAPI for programmatic generation and editing.
  • "(《世界人权宣言》) Music API is available to paid users.
  • Commercial-use licensing is listed for Starter and higher plans.

Access Through GlobalGPT

  • Use Music v2 alongside other audio and AI models in one account.
  • No separate VPN or provider setup is required.
  • A single platform can cost less to maintain than several standalone subscriptions.
  • You can move to another model without rebuilding your full tool stack.
Try the GlobalGPT audio generator

API Access and Available Controls

The official SDK exposes the model through elevenlabs.music.composemodel_id="music_v2". A basic request can set a natural-language 推动music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.

  • Generation: create a complete song or instrumental from a prompt.
  • Detailed response: request more information about the generated composition.
  • 流媒体: begin receiving audio without waiting for the complete file.
  • Inpainting: regenerate a selected region instead of replacing the whole track.
  • Reference matching: guide a generation with reference material where supported.

Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_promptbad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.

Example API Request

Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

audio = client.music.compose(
    model_id="music_v2",
    prompt=(
        "Create an original cinematic instrumental with felt piano, "
        "low strings, a gradual build, and a resolved final chord."
    ),
    music_length_ms=90_000,
)

save(audio, "music-v2-test.mp3")

Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.

Pricing and Our Test Cost

项目Verified value是什么意思
Official Music API rate$0.15 per minutePublished on the ElevenLabs API pricing page, excluding taxes, levies, and duties.
Maximum duration5 minutesMaximum published output length for one music request.
Our 90-second task charge$0.2251.5 minutes × $0.15; reported by each completed test task.
Our three-test total$0.675Three 90-second outputs; no rerun cost included.

The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.

How to Prompt ElevenLabs Music v2

Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:

1输出Song or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5除外条款No speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.

Useful English Music Terms

Click any term or descriptor to copy it. 例如 warm, intimate lead vocal with clear diction.

01Lyrics-First Pop SongFor an original vocal song using your exact words
03Background Music BedFor videos, podcasts, product demos, or cafés