ElevenLabs Music v2 レビュー:実際のパフォーマンス、API、およびプロンプト

elevenlabs-music-v2-レビュー-ヒーロー-3x2-v2

AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?

We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.

Quick Answer: How Good Is ElevenLabs Music v2?

OUR SHORT VERDICT

ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.

It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.

  • Best observed strength: lyrics and audible song structure.
  • Main limitation: fine timing and ending instructions were not always exact.
  • Test scope: three controlled API runs, not an industry benchmark.

ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.

ElevenLabs Music v2 Overview

Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.

「ElevenLabs Music v2」とは何ですか?

ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.

Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.

MODEL ID
music_v2
MAXIMUM DURATION
5 minutes
PUBLISHED API RATE
$0.15/minute
AUDIO OUTPUT
44.1 kHz
PUBLISHED BITRATE
128–192 kbps
API ACCESS
有料ユーザー
ElevenLabs公式『Introducing Music v2』リリースページ
ElevenLabs’ official Music v2 release page, checked August 2026.

ElevenLabs Music v2 Test Results

We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.

Test 1: Pop Song With Supplied Lyrics

TEST 01Pop Song With Supplied Lyrics
90秒の目標112 BPM未成年者アルト声部

タスクの設定: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.

Show the full copyable prompt
温かみがあり、表現力豊かなアルトのリードボーカルをフィーチャーした、完全にオリジナルの90秒間のコンテンポラリー・シンセポップ曲を1曲制作してください。BPMは112、調はAマイナーとします。ムードは希望に満ち、親密で、映画的な雰囲気を醸し出すものにしてください。ミュートをかけたエレクトリックピアノと、柔らかなアナログシンセのパルスから曲を始めます。 最初のヴァースでは、丸みのあるベースとタイトな電子ドラムを加え、プリコーラスで緊張感を高め、控えめなバックグラウンド・ハーモニーを伴った、広がりのある印象的なコーラスへと展開させてください。リードボーカルは明瞭に、センターに配置し、各楽器の分離を明確に保ってください。ナレーションによるイントロ、ラップ、観衆の歓声、実在するアーティストの模倣、フェードアウトは一切行わないでください。 明確な最終コードで終わらせてください。

以下の歌詞を正確に歌ってください。単語やセクションの追加、削除、置換、順序の変更は行わないでください。

[ヴァース1]
背後で街の明かりが消えていく
静まり返った通りが今、道を導く
真夜中のささやき、振り返ることはない
新しい地平線が、灰色の空を切り裂く
[プリコーラス]
手を取り合って、未知の道を歩む
自分たちだけの夢を追い求めて
[コーラス]
静かな夜に、私たちは生きている
地平線の彼方に星を見つける
影に迷いながらも、とても輝いている
共に、新たな定義を築こう
[ヴァース2]
人影のない道に響く柔らかな足音
息づかいひとつひとつが物語を紡ぐ
約束などない、ただ果てしない空があるだけ
君の瞳の中に、未来が宿っている
[Final Chorus]
静かな夜の中で、私たちは生きている
地平線の彼方に輝く星を見つける
影に迷いながらも、それでもとても輝いている
共に、新たな定義を切り拓こう
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0秒出力
42.0秒GENERATION TIME
$0.225TASK CHARGE
  • 歌詞に明らかな逸脱は見られない was detected in the supplied words.
  • The output had a polished electronic synth-pop character and clean vocal/instrument separation.
  • The intro felt longer than requested before the vocal arrived.

Test 2: Cinematic Instrumental

TEST 02Cinematic Instrumental
90秒の目標82 BPMニ短調ボーカルなし

タスクの設定: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.

Show the full copyable prompt
ドラマチックなSF探検シーン向けに、完全にオリジナルの90秒間のシネマティック・インストゥルメンタル楽曲を1曲制作してください。BPMは82、調性はDマイナーとしてください。ボーカル、ナレーション、ささやき声、合唱、および人間の声に似た音は一切使用しないでください。 冒頭は、控えめなフェルトピアノのモチーフと低音域の持続する弦楽器で静かに始めます。オープニングの後、繊細なアナログのパルスを取り入れ、その後、重なり合う弦楽器と抑制の効いた低音パーカッションで徐々に盛り上げていきます。金管楽器は最後の3分の1の部分でのみ追加してください。終盤近くで、過度な音量や歪みを生じさせることなく、明確な感情的なクライマックスに達するようにしてください。 ピアノのモチーフは一聴してそれと分かるようにし、ピアノ、弦楽器、パーカッション、シンセのパルス、金管楽器の間の明確な分離を保つこと。リアルなオーケストラの音色を使用し、空間的でありながらもコントロールされたミックスにすること。前半ではトレーラー特有のインパクト音は避けること。既存の楽曲、作曲家、またはフランチャイズを模倣しないこと。フェードアウトではなく、解決された持続和音で終わらせること。.
Test 02 resultA convincing cinematic build, with one early impact
90.0秒出力
20.9秒GENERATION TIME
$0.225TASK CHARGE
  • No voice-like texture was detected.
  • The piano motif remained recognizable as the arrangement built around it.
  • An impact arrived earlier than the brief allowed, weakening strict timing compliance.

Test 3: Eight-Part Song Structure

TEST 03Eight-Part Song Structure
90秒の目標104 BPMニ長調セクションの順序を固定しました

タスクの設定: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.

Show the full copyable prompt
明確で表現力豊かなテノールのリードボーカルを特徴とする、完全にオリジナルの90秒間のオルタナティブ・ポップ曲を1曲制作してください。BPMは104、調性はDメジャーとし、モダンなオルタナティブ・ポップのサウンドプロダクション、自然な発音、控えめなボーカルエフェクトを採用してください。ナレーション、ラップ、実在するアーティストの模倣、フェードアウトは一切含まないこと。 以下のセクション順序を厳守し、プロダクション上の変更がすべて聴き取れるようにしてください:イントロ—約6秒、フィルターをかけたエレキギターと遠くで響くシンセのみ;ヴァース—ドライなベースとリムクリックのパーカッションを追加;プレコーラス—フルドラムなしで上昇するパッドを追加;コーラス—フルドラム、広がりのあるシンセ、そして明瞭なメロディックなフックを導入; インストゥルメンタル・ブレイク—約8秒間、ボーカルなしでコーラスのメロディをエコーさせる;ブリッジ—ピアノと親密で近距離感のあるボーカルのみに縮小;ファイナル・コーラス—フルバンドを復活させ、バックコーラスとオクターブ・ギター・ラインを追加し、最初のコーラスよりもスケールを大きくする; アウトロ—ドラムを削除し、フィルターをかけたギターに戻し、クリーンな最終コードで締めくくる。

以下の歌詞を正確に歌ってください。追加、削除、置換、順序変更を行わないでください。また、括弧内のインストゥルメンタルに関する指示を歌詞として歌わないでください。

[イントロ]
(インストゥルメンタル)
[ヴァース]
ひび割れたダイヤルが、壊れたメロディーを唸る
指が色あせた線をなぞる
雑音のささやきが部屋を満たす
絡み合った兆しを探し求める
[プレコーラス]
信号がちらつき、不鮮明な音
深遠な声を待ちわびて
[コーラス]
雑音の中、私の声が聞こえる?
遠くからの呼び声、儚い選択
雑音が消え、より鮮明な声が聞こえる
しっかり掴まって、そのノイズを見つけよう
[インストルメンタル・ブレイク]
(インストルメンタル)
[ブリッジ]
砕けた光の中の儚い希望
柔らかな残響が、私を引き寄せる
静寂の中、君が現れる
[ラスト・コーラス]
雑音の中、私の声が聞こえる?
遠くからの呼び声、儚い選択
雑音が消え、声が鮮明に
しっかり掴まって、その雑音を見つけよう
[アウトロ]
ささやきが歌へと変わる
私たちは、あるべき場所にいる
Test 03 resultThe clearest structural result, but not perfect timing control
90.0秒出力
20.9秒GENERATION TIME
$0.225TASK CHARGE
  • All eight sections were easy to identify, with strong bridge and final-chorus contrast.
  • The instrumental break ran longer than the requested eight seconds.
  • The ending behaved more like a fade than the clean final chord in the prompt.

What the Three Tests Show

Performance Findings

3/3Usable outputs

Every run returned a playable 90-second file that fit the broad musical task.

1Clear lyric pass

The supplied pop lyrics showed no clear deviation in our listening checks.

8Audible sections

The fixed-structure song made all requested sections easy to identify.

  • Vocals: clear, centered, and separated well from the electronic backing in Test 1.
  • Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
  • Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
  • Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.

Efficiency Findings

Generation Time Across Three 90-Second Runs
Pop vocals
42.0秒
Instrumental
20.9秒
Eight-part song
20.9秒

Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.

  • Total output: 270 seconds across three files.
  • Total task-reported charge: $0.675.
  • Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
  • No reruns: all three completed tasks produced usable audio on the first recorded attempt.

How to Access ElevenLabs Music v2

Access Through ElevenLabs

  • Use ElevenMusic for browser-based composition.
  • Use ElevenAPI for programmatic generation and editing.
  • について Music API is available to paid users.
  • Commercial-use licensing is listed for Starter and higher plans.

Access Through GlobalGPT

  • Use Music v2 alongside other audio and AI models in one account.
  • No separate VPN or provider setup is required.
  • A single platform can cost less to maintain than several standalone subscriptions.
  • You can move to another model without rebuilding your full tool stack.
Try the GlobalGPT audio generator

API Access and Available Controls

The official SDK exposes the model through elevenlabs.music.composemodel_id="music_v2". A basic request can set a natural-language 迅速 そして music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.

  • Generation: create a complete song or instrumental from a prompt.
  • Detailed response: request more information about the generated composition.
  • ストリーミング: begin receiving audio without waiting for the complete file.
  • Inpainting: regenerate a selected region instead of replacing the whole track.
  • Reference matching: guide a generation with reference material where supported.

Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_prompt または bad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.

Example API Request

Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

audio = client.music.compose(
    model_id="music_v2",
    prompt=(
        "Create an original cinematic instrumental with felt piano, "
        "low strings, a gradual build, and a resolved final chord."
    ),
    music_length_ms=90_000,
)

save(audio, "music-v2-test.mp3")

Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.

Pricing and Our Test Cost

項目Verified valueその意味
Official Music API rate$0.15 per minutePublished on the ElevenLabs API pricing page, excluding taxes, levies, and duties.
Maximum duration5 minutesMaximum published output length for one music request.
Our 90-second task charge$0.2251.5 minutes × $0.15; reported by each completed test task.
Our three-test total$0.675Three 90-second outputs; no rerun cost included.

The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.

How to Prompt ElevenLabs Music v2

Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:

1出力Song or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5適用除外No speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.

Useful English Music Terms

Click any term or descriptor to copy it. warm, intimate lead vocal with clear diction.

01Lyrics-First Pop SongFor an original vocal song using your exact words
03Background Music BedFor videos, podcasts, product demos, or cafés