ElevenLabs Music v2 評測:實際表現、API 與提示詞

elevenlabs-music-v2-評測-hero-3x2-v2

AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?

We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.

Quick Answer: How Good Is ElevenLabs Music v2?

OUR SHORT VERDICT

ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.

It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.

  • Best observed strength: lyrics and audible song structure.
  • Main limitation: fine timing and ending instructions were not always exact.
  • Test scope: three controlled API runs, not an industry benchmark.

ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.

ElevenLabs Music v2 Overview

Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.

什麼是 ElevenLabs Music v2?

ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.

Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.

MODEL ID
music_v2
MAXIMUM DURATION
5 minutes
PUBLISHED API RATE
$0.15/minute
AUDIO OUTPUT
44.1 kHz
PUBLISHED BITRATE
128–192 kbps
API ACCESS
付費用戶
ElevenLabs 官方《Introducing Music v2》發行頁面
ElevenLabs’ official Music v2 release page, checked August 2026.

ElevenLabs Music v2 Test Results

We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.

Test 1: Pop Song With Supplied Lyrics

TEST 01Pop Song With Supplied Lyrics
90 秒目標每分鐘 112 次未成年人女高音

任務設定: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.

Show the full copyable prompt
創作一首完全原創的 90 秒當代合成器流行歌曲,主唱為溫暖且富有表現力的中音女聲。節拍為 112 BPM,調性為 A 小調。整體氛圍應充滿希望、親密且具電影感。開頭以弱音電鋼琴搭配柔和的類比合成器脈動聲開始。 在第一段主歌中加入圓潤的貝斯與緊湊的電子鼓聲,於前副歌部分營造張力,並過渡至寬廣且令人難忘的副歌,搭配克制的背景和聲。保持主唱聲音清晰且居中,樂器聲部分層分明。不得包含說唱式引言、饒舌、群眾聲效、模仿真實藝人,或淡出效果。 以清晰的終結和弦收尾。

請嚴格按照以下歌詞演唱。不得增刪、替換或重新排列任何詞句或段落。

[第一段]
城市燈火在我們身後漸漸淡去
寂靜的街道如今引領前路
午夜的低語,不回頭
嶄新的地平線劃破灰濛
[前副歌]
十指相扣,我們踏上未知的旅程
追尋屬於我們自己的夢想
[副歌]
在寂靜的夜裡,我們活著
在天際線彼端尋覓星辰
迷失於陰影之中,卻依然璀璨
攜手同行,我們將重新定義
[第二段]
空蕩道路上輕柔的腳步聲
每一次呼吸都訴說著一個故事
沒有承諾,只有無垠的天空
未來,盡在你的眼底
[終結副歌]
我們在寂靜的夜裡活著
在天際線彼端尋覓星辰
雖迷失在陰影中,卻依然璀璨
我們將攜手重新定義
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0 秒輸出
42.0 秒GENERATION TIME
$0.225TASK CHARGE
  • 歌詞並無明顯偏差 was detected in the supplied words.
  • The output had a polished electronic synth-pop character and clean vocal/instrument separation.
  • The intro felt longer than requested before the vocal arrived.

Test 2: Cinematic Instrumental

TEST 02Cinematic Instrumental
90 秒目標每分鐘 82 次D小調無人聲

任務設定: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.

Show the full copyable prompt
為一幕充滿戲劇張力的科幻探索場景,創作一首完全原創、長度為 90 秒的電影風格器樂配樂。節拍為 82 BPM,調性為 D 小調。不得包含人聲、對白、低語、合唱或任何類似人聲的音效。 以稀疏的軟質鋼琴主題和低音弦樂的長音靜謐開場。開場後引入微妙的類比節拍,隨後透過層疊的弦樂與克制的低音打擊樂逐步鋪陳。僅在最後三分之一處加入銅管樂器。接近結尾時達到一個清晰的情感高潮,但避免過度響亮或失真。 保持鋼琴主題的辨識度,並確保鋼琴、弦樂、打擊樂、合成器脈動與銅管樂器之間有清晰的分離度。採用真實的管弦樂音色,混音效果應寬廣但受控。前半段應避免使用預告片常見的衝擊音效。請勿模仿現有的配樂、作曲家或系列作品。結尾應以一個解決後的長音和弦收尾,而非淡出。.
Test 02 resultA convincing cinematic build, with one early impact
90.0 秒輸出
20.9 秒GENERATION TIME
$0.225TASK CHARGE
  • No voice-like texture was detected.
  • The piano motif remained recognizable as the arrangement built around it.
  • An impact arrived earlier than the brief allowed, weakening strict timing compliance.

Test 3: Eight-Part Song Structure

TEST 03Eight-Part Song Structure
90 秒目標每分鐘 104 次D大調固定章節順序

任務設定: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.

Show the full copyable prompt
創作一首完全原創的 90 秒另類流行歌曲,主唱需為清晰且富有表現力的男高音。節拍為 104 BPM,調性為 D 大調,採用現代另類流行風格的製作手法,發音自然,且人聲效果應適度克制。不得包含口白、饒舌、模仿真實藝人,亦不得使用淡出效果。 請嚴格維持以下段落順序,並確保每項製作變更皆能清晰可辨:前奏——約6秒,僅含經過濾波處理的電吉他與遙遠的合成器聲;主歌——加入未經處理的貝斯與鼓邊敲擊聲;前副歌——加入漸強的墊音,但暫不使用完整鼓組;副歌——加入完整鼓組、更寬廣的合成器聲,以及清晰的主旋律鉤子; 器樂間奏——約 8 秒,無人聲,僅以迴響效果重現副歌旋律;過門——簡化為鋼琴與近距離的親密人聲;最終副歌——恢復完整樂團編制,加入和聲與八度吉他旋律線,並使其氣勢比第一個副歌更為宏大; 結尾段—移除鼓組,回歸經過濾波處理的吉他聲,並以一個清澈的終結和弦收尾。

請精準演唱以下歌詞。切勿增刪、替換、重新排序,亦不得將括號內的器樂指示當作歌詞來演唱。

[前奏]
(器樂)
[主歌]
龜裂的錶盤低鳴著破碎的旋律
指尖撫過褪色的線條
靜電的低語充斥整個房間
在糾結的徵兆中搜尋
[前副歌]
訊號閃爍,聲響模糊
等待著那深邃的聲音
[副歌]
你能穿過這喧囂聽見我嗎?
遙遠的呼喚,脆弱的抉擇
雜訊漸淡,聲音更清晰
緊緊抓住,我們將找到那噪音
[器樂間奏]
(器樂)
[過門]
破碎光線中的脆弱希望
柔和的迴聲,將我拉近
在寂靜中,你現身
[終曲副歌]
你能穿過這喧囂聽見我嗎?
遙遠的呼喚,脆弱的抉擇
雜訊漸消,聲音更清晰
緊緊抓住,我們將找到那喧囂
[尾聲]
低語化作歌聲
我們已回到歸屬之地
Test 03 resultThe clearest structural result, but not perfect timing control
90.0 秒輸出
20.9 秒GENERATION TIME
$0.225TASK CHARGE
  • All eight sections were easy to identify, with strong bridge and final-chorus contrast.
  • The instrumental break ran longer than the requested eight seconds.
  • The ending behaved more like a fade than the clean final chord in the prompt.

What the Three Tests Show

Performance Findings

3/3Usable outputs

Every run returned a playable 90-second file that fit the broad musical task.

1Clear lyric pass

The supplied pop lyrics showed no clear deviation in our listening checks.

8Audible sections

The fixed-structure song made all requested sections easy to identify.

  • Vocals: clear, centered, and separated well from the electronic backing in Test 1.
  • Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
  • Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
  • Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.

Efficiency Findings

Generation Time Across Three 90-Second Runs
Pop vocals
42.0 秒
Instrumental
20.9 秒
Eight-part song
20.9 秒

Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.

  • Total output: 270 seconds across three files.
  • Total task-reported charge: $0.675.
  • Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
  • No reruns: all three completed tasks produced usable audio on the first recorded attempt.

How to Access ElevenLabs Music v2

Access Through ElevenLabs

  • Use ElevenMusic for browser-based composition.
  • Use ElevenAPI for programmatic generation and editing.
  • Music API is available to paid users.
  • Commercial-use licensing is listed for Starter and higher plans.

Access Through GlobalGPT

  • Use Music v2 alongside other audio and AI models in one account.
  • No separate VPN or provider setup is required.
  • A single platform can cost less to maintain than several standalone subscriptions.
  • You can move to another model without rebuilding your full tool stack.
Try the GlobalGPT audio generator

API Access and Available Controls

The official SDK exposes the model through elevenlabs.music.composemodel_id="music_v2". A basic request can set a natural-language 提示music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.

  • Generation: create a complete song or instrumental from a prompt.
  • Detailed response: request more information about the generated composition.
  • 串流: begin receiving audio without waiting for the complete file.
  • Inpainting: regenerate a selected region instead of replacing the whole track.
  • Reference matching: guide a generation with reference material where supported.

Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_promptbad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.

Example API Request

Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

audio = client.music.compose(
    model_id="music_v2",
    prompt=(
        "Create an original cinematic instrumental with felt piano, "
        "low strings, a gradual build, and a resolved final chord."
    ),
    music_length_ms=90_000,
)

save(audio, "music-v2-test.mp3")

Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.

Pricing and Our Test Cost

項目Verified value這代表什麼
Official Music API rate$0.15 per minutePublished on the ElevenLabs API pricing page, excluding taxes, levies, and duties.
Maximum duration5 minutesMaximum published output length for one music request.
Our 90-second task charge$0.2251.5 minutes × $0.15; reported by each completed test task.
Our three-test total$0.675Three 90-second outputs; no rerun cost included.

The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.

How to Prompt ElevenLabs Music v2

Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:

1輸出Song or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5除外事項No speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.

Useful English Music Terms

Click any term or descriptor to copy it. 範例: warm, intimate lead vocal with clear diction.

01Lyrics-First Pop SongFor an original vocal song using your exact words
03Background Music BedFor videos, podcasts, product demos, or cafés