รีวิว ElevenLabs Music v2: ประสิทธิภาพจริง, API และ Prompt

elevenlabs-music-v2-review-hero-3x2-v2

AI music is moving fast, and ElevenLabs Music v2 is one of the most ambitious new song-generation models. But can it preserve supplied lyrics, follow a detailed arrangement, and deliver a usable track without heavy editing?

We tested it with three original 90-second briefs: a synth-pop vocal song, a cinematic instrumental, and an eight-part alternative-pop arrangement. This review covers the audio, prompt compliance, generation time, task-reported cost, official controls, and API access.

Quick Answer: How Good Is ElevenLabs Music v2?

OUR SHORT VERDICT

ElevenLabs Music v2 was strongest when the brief demanded explicit lyrics and section control. It produced all three requested 90-second files, kept the supplied words intact in our lyric test, and made the eight requested song sections easy to hear.

It was less exact at the edges: one instrumental impact arrived earlier than requested, one instrumental break ran long, and the structured song ended more like a fade than the clean final chord in the prompt.

  • Best observed strength: lyrics and audible song structure.
  • Main limitation: fine timing and ending instructions were not always exact.
  • Test scope: three controlled API runs, not an industry benchmark.

ElevenLabs Music v2 is also available through GlobalGPT’s audio generator. That route puts multiple audio and AI models in one account, avoids a separate VPN or provider setup, and can reduce the cost of maintaining several standalone subscriptions.

ElevenLabs Music v2 Overview

Before looking at our results, it helps to separate the model’s published feature set from what three prompts can actually prove.

ElevenLabs Music v2 คืออะไร?

ElevenLabs introduced Music v2 on May 26, 2026 as an upgraded music-generation model for songs and instrumentals. The company says the release improves vocals, instrumentation, arrangement, multilingual support, dense lyric delivery, genre transitions, and inpainting.

Music v2 powers ElevenMusic, ElevenAPI, and ElevenCreative. Through the API, developers can generate tracks programmatically, edit a selected region with inpainting, and work with reference matching.

MODEL ID
music_v2
MAXIMUM DURATION
5 minutes
PUBLISHED API RATE
$0.15/minute
AUDIO OUTPUT
44.1 kHz
PUBLISHED BITRATE
128–192 kbps
API ACCESS
ผู้ใช้ที่ชำระเงินแล้ว
หน้าประกาศการปล่อยเวอร์ชัน v2 ของ ElevenLabs official Introducing Music
ElevenLabs’ official Music v2 release page, checked August 2026.

ElevenLabs Music v2 Test Results

We used three controlled API runs with a fixed 90-second target. GPT generated the original lyric sets once on our platform, and we froze them before testing Music v2. Listening notes were cross-checked with Gemini 3.6 Flash audio analysis. Times below are local end-to-end generation time; charges are values reported by the completed tasks.

Test 1: Pop Song With Supplied Lyrics

TEST 01Pop Song With Supplied Lyrics
เป้าหมาย 90 วินาที112 BPMผู้เยาว์เสียงอัลโต

การตั้งค่างาน: Create a contemporary synth-pop song about two people leaving a city late at night, using the supplied words exactly. We listened for lyric preservation, vocal clarity, hook strength, arrangement, and the requested final chord.

Show the full copyable prompt
สร้างเพลงซินธ์-ป็อปสมัยใหม่ที่มีความยาว 90 วินาที และมีความเป็นต้นฉบับอย่างสมบูรณ์ พร้อมด้วยเสียงร้องนำอัลโตที่อบอุ่นและแสดงอารมณ์ได้ดี ใช้ความเร็ว 112 BPM และคีย์ A minor บรรยากาศของเพลงควรเป็นความหวัง ความใกล้ชิด และมีความรู้สึกเหมือนในภาพยนตร์ เริ่มต้นด้วยเสียงเปียโนไฟฟ้าที่ลดเสียงลง และจังหวะพัลส์จากซินธ์แอนะล็อกที่นุ่มนวล เพิ่มเบสที่กลมกลืนและกลองอิเล็กทรอนิกส์ที่กระชับในท่อนแรก สร้างความตึงเครียดในส่วนก่อนคอรัส และเปิดเข้าสู่คอรัสที่กว้างและน่าจดจำด้วยฮาร์โมนีพื้นหลังที่ควบคุมได้ รักษาร้องนำให้ชัดเจนและอยู่ตรงกลาง พร้อมการแยกเสียงเครื่องดนตรีที่ชัดเจน ห้ามมีบทนำที่พูด แร็ป เสียงฝูงชน การเลียนแบบศิลปินจริง หรือการเฟดเอาท์ จบด้วยคอร์ดสุดท้ายที่ชัดเจน

ร้องตามเนื้อเพลงด้านล่างอย่างแม่นยำ ห้ามเพิ่ม ลบ เปลี่ยน หรือจัดเรียงคำหรือส่วนใดส่วนหนึ่งใหม่

[Verse 1]
แสงไฟเมืองค่อยๆ จางหายหลังเรา
ถนนที่เงียบสงบนำทางเราไป
เสียงกระซิบยามเที่ยงคืน ไม่หันหลังกลับ
ขอบฟ้าใหม่แตกผ่านความเทา
[Pre-Chorus]
มือจับมือ เราเดินสู่สิ่งที่ไม่รู้จัก
ไล่ตามความฝันที่เราเรียกว่าของเราเอง
[Chorus]
เรายังมีชีวิตอยู่ในคืนที่เงียบสงบ
ค้นพบดวงดาวที่อยู่เหนือเส้นขอบฟ้า
หลงในเงามืด แต่ยังคงสว่างไสว
ร่วมกัน เราจะนิยามใหม่
[Verse 2]
เสียงฝีเท้าเบาๆ บนถนนที่ว่างเปล่า
ทุกการหายใจคือเรื่องราวที่เล่า
ไม่มีคำสัญญา เพียงท้องฟ้าที่ไร้ขอบเขต
ในดวงตาคุณ อนาคตอยู่ตรงนั้น
[Final Chorus]
เรายังมีชีวิตอยู่ในคืนที่เงียบสงบ
ค้นพบดวงดาวที่อยู่ไกลเกินเส้นขอบฟ้า
หลงอยู่ในเงามืด แต่กลับส่องประกายเจิดจ้า
ร่วมกัน เราจะนิยามใหม่
Test 01 resultStrong lyric consistency with a longer electronic intro
90.0 วินาทีผลลัพธ์
42.0 วินาทีGENERATION TIME
$0.225TASK CHARGE
  • ไม่มีความแตกต่างที่ชัดเจนในเนื้อเพลง was detected in the supplied words.
  • The output had a polished electronic synth-pop character and clean vocal/instrument separation.
  • The intro felt longer than requested before the vocal arrived.

Test 2: Cinematic Instrumental

TEST 02Cinematic Instrumental
เป้าหมาย 90 วินาที82 BPMดี ไมเนอร์ไม่มีเสียงร้อง

การตั้งค่างาน: Build a science-fiction cue around a recognizable felt-piano motif, with gradual orchestral growth, brass only in the final third, no human voice texture, and a resolved ending.

Show the full copyable prompt
สร้างเพลงบรรเลงแบบภาพยนตร์ที่มีความยาว 90 วินาที และมีความเป็นต้นฉบับอย่างสมบูรณ์ สำหรับฉากการสำรวจในภาพยนตร์วิทยาศาสตร์แนวระทึกขวัญ ใช้จังหวะ 82 BPM และคีย์ D minor ห้ามใช้เสียงร้อง คำพูด เสียงกระซิบ เสียงร้องประสาน หรือเสียงที่คล้ายเสียงมนุษย์ เริ่มต้นอย่างเงียบสงบด้วยธีมเปียโนแบบเฟลท์ที่เรียบง่ายและเสียงสตริงต่ำที่ยืดยาว นำจังหวะแอนะล็อกที่ละเอียดอ่อนเข้ามาหลังส่วนเปิด จากนั้นค่อยๆ สร้างความเข้มข้นด้วยสตริงที่ซ้อนชั้นและเครื่องตีต่ำที่ควบคุมได้ เพิ่มเครื่องทองเหลืองเฉพาะในส่วนสุดท้ายของเพลง สร้างจุดสูงสุดทางอารมณ์ที่ชัดเจนใกล้ตอนจบ โดยไม่ใช้ความดังหรือการบิดเบือนเสียงที่มากเกินไป รักษาให้ธีมเปียโนยังคงจดจำได้ และรักษาความแยกแยะที่ชัดเจนระหว่างเปียโน เครื่องสาย เครื่องตี จังหวะซินธ์ และเครื่องทองเหลือง ใช้เสียงออร์เคสตราที่สมจริง พร้อมการมิกซ์ที่กว้างขวางแต่ควบคุมได้ หลีกเลี่ยงการใช้เอฟเฟกต์แบบเทรลเลอร์ในครึ่งแรก อย่าเลียนแบบเพลงประกอบที่มีอยู่ นักแต่งเพลง หรือแฟรนไชส์ที่มีอยู่แล้ว จบด้วยคอร์ดที่ยืนยาวและลงตัว แทนที่จะเฟดเอาต์.
Test 02 resultA convincing cinematic build, with one early impact
90.0 วินาทีผลลัพธ์
20.9 วินาทีGENERATION TIME
$0.225TASK CHARGE
  • No voice-like texture was detected.
  • The piano motif remained recognizable as the arrangement built around it.
  • An impact arrived earlier than the brief allowed, weakening strict timing compliance.

Test 3: Eight-Part Song Structure

TEST 03Eight-Part Song Structure
เป้าหมาย 90 วินาที104 BPMD เมเจอร์จัดเรียงส่วนให้คงที่

การตั้งค่างาน: Create an alternative-pop song with eight audible sections in a fixed order: Intro, Verse, Pre-Chorus, Chorus, Instrumental Break, Bridge, Final Chorus, and Outro.

Show the full copyable prompt
สร้างเพลงอัลเทอร์เนทีฟ-ป็อปต้นฉบับใหม่ทั้งหมดความยาว 90 วินาที ด้วยเสียงร้องนำเทเนอร์ที่ชัดเจนและแสดงอารมณ์ได้ดี ใช้ความเร็ว 104 BPM และคีย์ D major พร้อมการผลิตแบบอัลเทอร์เนทีฟ-ป็อปสมัยใหม่ การออกเสียงตามธรรมชาติ และเอฟเฟกต์เสียงร้องที่จำกัด ไม่มีการบรรยายด้วยเสียงพูด แร็ป การเลียนแบบศิลปินจริง หรือการเฟดเอาท์ รักษาลำดับส่วนนี้ให้ตรงตามเดิมและทำให้ทุกการเปลี่ยนแปลงในการผลิตได้ยินได้ชัดเจน: Intro—ประมาณ 6 วินาที ใช้เพียงกีตาร์ไฟฟ้าที่ผ่านฟิลเตอร์และซินธ์ที่ฟังดูห่างไกล; Verse—เพิ่มเบสแบบ dry และเครื่องตีแบบ rim-click; Pre-Chorus—เพิ่ม pad ที่ค่อยๆ เพิ่มระดับเสียงขึ้นโดยไม่ใช้ชุดกลองเต็มรูปแบบ; Chorus—นำชุดกลองเต็มรูปแบบ ซินธ์ที่กว้างขึ้น และ hook ที่ชัดเจนและมีทำนอง; ช่วงพักเครื่องดนตรี—ประมาณ 8 วินาที, ใช้เสียงเอคโค่ของทำนองคอรัสโดยไม่มีเสียงร้อง; ช่วงบริดจ์—ลดเหลือเพียงเปียโนและเสียงร้องที่ใกล้ชิดและเป็นส่วนตัว; คอรัสสุดท้าย—นำวงดนตรีเต็มรูปแบบกลับมา เพิ่มเสียงประสานและไลน์กีตาร์อ็อกเทฟ และทำให้มีขนาดใหญ่กว่าคอรัสแรก; Outro—ลบกลองออก กลับสู่เสียงกีตาร์ที่ผ่านฟิลเตอร์ และจบด้วยคอร์ดสุดท้ายที่ชัดเจน

ร้องตามเนื้อเพลงที่ระบุไว้ด้านล่างอย่างแม่นยำ ห้ามเพิ่ม ลบ เปลี่ยนลำดับ หรือร้องคำสั่งเครื่องดนตรีในวงเล็บเป็นเนื้อเพลง

[Intro]
(เครื่องดนตรี)
[Verse]
หน้าปัดที่แตกร้าวส่งเสียงเพลงที่ขาดๆ หายๆ
นิ้วมือลูบตามเส้นที่จางหาย
เสียงรบกวนกระซิบเติมเต็มห้อง
ค้นหาผ่านสัญญาณที่พันกัน
[Pre-Chorus]
สัญญาณกระพริบ เสียงไม่ชัดเจน
รอคอยเสียงที่ลึกซึ้ง
[Chorus]
คุณได้ยินฉันผ่านเสียงรบกวนได้ไหม?
เสียงเรียกจากไกล การเลือกที่เปราะบาง
เสียงรบกวนจางลง เสียงที่ชัดเจนขึ้น
จับมือกันแน่น เราจะหาเสียงรบกวนนั้นเจอ
[ช่วงพักดนตรี]
(ดนตรี)
[บริดจ์]
ความหวังที่เปราะบางในแสงที่แตกสลาย
เสียงสะท้อนเบาๆ ดึงฉันเข้าใกล้
ในความเงียบ คุณปรากฏตัว
[คอรัสสุดท้าย]
คุณได้ยินฉันผ่านเสียงรบกวนได้ไหม?
เสียงเรียกจากไกลๆ การเลือกที่เปราะบาง
เสียงรบกวนค่อยๆ จางลง เสียงที่ชัดเจนขึ้น
จับมือกันไว้แน่น เราจะหาทางผ่านเสียงรบกวนนี้
[ส่วนท้าย]
เสียงกระซิบเปลี่ยนเป็นเพลง
เราอยู่ที่ที่เราควรอยู่
Test 03 resultThe clearest structural result, but not perfect timing control
90.0 วินาทีผลลัพธ์
20.9 วินาทีGENERATION TIME
$0.225TASK CHARGE
  • All eight sections were easy to identify, with strong bridge and final-chorus contrast.
  • The instrumental break ran longer than the requested eight seconds.
  • The ending behaved more like a fade than the clean final chord in the prompt.

What the Three Tests Show

Performance Findings

3/3Usable outputs

Every run returned a playable 90-second file that fit the broad musical task.

1Clear lyric pass

The supplied pop lyrics showed no clear deviation in our listening checks.

8Audible sections

The fixed-structure song made all requested sections easy to identify.

  • Vocals: clear, centered, and separated well from the electronic backing in Test 1.
  • Instrumental writing: the piano motif and controlled build in Test 2 showed that the model can handle more than vocal pop.
  • Prompt control: broad structure was stronger than frame-level timing; intros, breaks, impacts, and endings still deserve manual review.
  • Editing burden: none of the three outputs was unusable, but Tests 2 and 3 would need small arrangement edits for strict production delivery.

Efficiency Findings

Generation Time Across Three 90-Second Runs
Pop vocals
42.0 วินาที
Instrumental
20.9 วินาที
Eight-part song
20.9 วินาที

Average local end-to-end generation time: 27.9 seconds. Local time includes request handling, polling, and network delay; it is not provider-side latency.

  • Total output: 270 seconds across three files.
  • Total task-reported charge: $0.675.
  • Charge per run: $0.225 for 90 seconds, which equals 1.5 minutes × the official $0.15-per-minute API rate.
  • No reruns: all three completed tasks produced usable audio on the first recorded attempt.

How to Access ElevenLabs Music v2

Access Through ElevenLabs

  • Use ElevenMusic for browser-based composition.
  • Use ElevenAPI for programmatic generation and editing.
  • The Music API is available to paid users.
  • Commercial-use licensing is listed for Starter and higher plans.

Access Through GlobalGPT

  • Use Music v2 alongside other audio and AI models in one account.
  • No separate VPN or provider setup is required.
  • A single platform can cost less to maintain than several standalone subscriptions.
  • You can move to another model without rebuilding your full tool stack.
Try the GlobalGPT audio generator

API Access and Available Controls

The official SDK exposes the model through elevenlabs.music.compose กับ model_id="music_v2". A basic request can set a natural-language คำสั่ง และ music_length_ms. Composition plans add more precise section structure, lyric timing, and arrangement control.

  • Generation: create a complete song or instrumental from a prompt.
  • Detailed response: request more information about the generated composition.
  • สตรีมมิ่ง: begin receiving audio without waiting for the complete file.
  • Inpainting: regenerate a selected region instead of replacing the whole track.
  • Reference matching: guide a generation with reference material where supported.

Prompts that name copyrighted artists, songs, or protected lyrics can trigger a bad_prompt หรือ bad_composition_plan response. Describe musical attributes instead of asking for a living artist’s exact style.

Example API Request

Python SDKModel: music_v2
import os
from elevenlabs import ElevenLabs, save

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

audio = client.music.compose(
    model_id="music_v2",
    prompt=(
        "Create an original cinematic instrumental with felt piano, "
        "low strings, a gradual build, and a resolved final chord."
    ),
    music_length_ms=90_000,
)

save(audio, "music-v2-test.mp3")

Keep the API key in an environment variable, not inside the script. For a production app, add request logging, retry limits, and a clear cost cap before generating long tracks.

Pricing and Our Test Cost

รายการVerified valueหมายความว่า
Official Music API rate$0.15 per minutePublished on the ElevenLabs API pricing page, excluding taxes, levies, and duties.
Maximum duration5 minutesMaximum published output length for one music request.
Our 90-second task charge$0.2251.5 minutes × $0.15; reported by each completed test task.
Our three-test total$0.675Three 90-second outputs; no rerun cost included.

The unit rate in our test record matches the official API rate. GlobalGPT’s value here is consolidated access and fewer separate subscriptions—not a verified lower per-minute Music v2 price.

How to Prompt ElevenLabs Music v2

Music v2 responded best when the brief separated the musical goal from the non-negotiable constraints. Build your prompt in this order:

1ผลลัพธ์Song or instrumental, plus duration.
2Musical frameGenre, mood, tempo, key, and vocal type.
3ArrangementOpening, section changes, climax, and ending.
4Mix prioritiesVocal position, separation, loudness, and space.
5ข้อยกเว้นNo speech, no fade, no crowd, or no voice texture.
6Lyrics and structureSection labels and whether words may change.

Useful English Music Terms

Click any term or descriptor to copy it. ตัวอย่าง: warm, intimate lead vocal with clear diction.

01Lyrics-First Pop SongFor an original vocal song using your exact words
03Background Music BedFor videos, podcasts, product demos, or cafés