Eleven Multilingual v2 is not a new 2026 release, but it is still an active ElevenLabs text-to-speech model worth considering. Launched in August 2023, it remains in the company’s flagship lineup with 29 supported languages, a 10,000-character request limit, and an official emphasis on stable long-form generation.
The catch is that “newer” and “better for your job” are not the same thing. Eleven v3 offers more expressive delivery and broader language coverage, while Flash v2.5 prioritizes speed, scale, and lower API cost. Multilingual v2 sits between them as the established option for narration, courses, localization, and other single-speaker work where consistency matters more than dramatic performance.
Want to try the model before setting up an API request? GlobalGPT’s Audio Generator puts Eleven Multilingual v2 in a browser workspace: choose the model, paste your script, pick a voice, and generate the audio.

What is Eleven Multilingual v2?
Eleven Multilingual v2 is a multilingual text-to-speech model from ElevenLabs. Its native API model ID is eleven_multilingual_v2. You send text and a voice ID to the Text to Speech endpoint, and the service returns synthesized audio in the selected voice.
Its practical appeal is simple: one model covers 29 languages and accepts up to 10,000 characters in a single request. ElevenLabs describes it as its most stable option for long-form generation, although that statement is the company’s positioning rather than an independent guarantee for every voice, script, or language.
Multilingual v2 is designed for work such as narration, audiobooks, e-learning, corporate voiceovers, and localized media. If your end product is a talking-character video, the same voiceover can also feed a broader dialogue and lip-sync workflow.
Is Eleven Multilingual v2 still active in 2026?
Yes. As of August 11, 2026, ElevenLabs lists Multilingual v2 under its flagship Text to Speech models. It does not appear in the documentation’s deprecated-model table. That is the clearest current status: active and documented, but no longer the newest ElevenLabs speech model.
The release date is separate from its present status. ElevenLabs announced Multilingual v2 on August 22, 2023. A page title that includes “2026” should describe the review date, not imply that the model launched this year.
That positioning makes Multilingual v2 an established production option, not abandoned software. It also means the buying decision should compare its steadier, longer-script profile with the specific advantages of v3 and Flash v2.5.
Which languages does Eleven Multilingual v2 support?
ElevenLabs officially lists 29 languages for Multilingual v2:
- English, Japanese, Chinese, German, Hindi, and French
- Korean, Portuguese, Italian, Spanish, Indonesian, and Dutch
- Turkish, Filipino, Polish, Swedish, Bulgarian, and Romanian
- Arabic, Czech, Greek, Finnish, Croatian, Malay, and Slovak
- Danish, Tamil, Ukrainian, and Russian
Language support does not mean every voice will sound equally convincing in every region. ElevenLabs recommends choosing a voice whose accent matches the target language and region. That matters for Spanish, Portuguese, English, and other languages with strong regional variation. Test the exact voice and script you plan to publish rather than treating the model’s language list as an accent guarantee.
Voice quality: naturalness, consistency, and the evidence
ElevenLabs describes Multilingual v2 as lifelike, emotionally aware, and consistent across languages. Those are vendor claims. They are useful for understanding the model’s intended position, but they should not be mistaken for neutral testing.
Independent evidence is narrower. Artificial Analysis listed Multilingual v2 at about 1100.07 Elo in its Provider Voice Arena when checked on August 11, 2026. The arena uses paired listener preferences and eight provider voices across US and UK English, split across male and female voices. That is transparent, model-specific evidence for English voice preference; it does not prove equivalent performance across all 29 languages, regional accents, or long scripts.
Public feedback is mixed, but the sample is too small to describe the user base as a whole. In one 2024 discussion, commenters described v2 as more realistic than v1 while also reporting occasional accent-consistency issues. Another user reported a short-term period of speed and consistency problems. A later ElevenLabs-associated response recommended Multilingual v2 or Turbo v2 for more consistent long-form work. These posts are useful warning signals, not prevalence data or controlled benchmarks.
What can be said confidently?
- The model remains officially supported and is positioned for steady, high-quality long-form speech.
- Independent English preference data gives it a credible external signal, within a limited English voice set.
- Accent quality depends on language, region, voice selection, and script—not just the model ID.
- Objective audio checks can detect truncation, clipping, format problems, and timing anomalies, but they cannot replace qualified listening.
Cross-language consistency
The same scenario was generated in English, Spanish, and Mandarin with the route-default voice. Each player contains the first valid output retained under the frozen test rule.
Listen for: continuity of speaker identity and obvious pronunciation issues. A qualified Spanish listener should assess accent authenticity separately.
These embeddings are an auxiliary identity signal; they do not evaluate naturalness, pronunciation, accent, prosody, or emotional delivery. No different-speaker calibration group was available.
Hands-on test: five first-valid outputs through the Anywhere API route
We tested the Anywhere direct HTTPS route to the model ID eleven-multilingual-v2. This is not the same identifier format as the native ElevenLabs API, which uses eleven_multilingual_v2. The route exposed a task interface with the model and prompt, but it did not expose a reliable voice selector, stability, similarity, style, speed, or output-format controls. Every run therefore used the route defaults.
The pre-registered batch contained five scripts: one 1,399-character English narration, the same scenario in English, Spanish, and Mandarin, and one pronunciation stress test covering names, dates, currency, a phone number, a URL, and acronyms. We retained the first valid playable output and did not rerun for quality.
Route reliability summary
All five pre-registered scripts used the first valid playable output. The route did not expose voice or generation controls, so every run used its defaults.
Alcance: these figures describe the Anywhere route, not native ElevenLabs latency or the GlobalGPT browser experience. File checks do not score naturalness, emotion, accent, or local pacing.
All five formal tasks succeeded on the first attempt and returned decodable mono MP3 files at 44.1 kHz. The files contained no clipped samples and no adjacent-sample jumps above 0.8 in the local waveform scan. Those checks establish file integrity; they do not score naturalness.
Long-form stability, emotion, pronunciation, and latency tradeoffs
Long scripts
The 10,000-character ceiling is generous enough for roughly ten minutes of audio according to ElevenLabs’ character-limit table. For longer chapters, split at natural paragraph or scene boundaries rather than cutting at an arbitrary character count. The API provides previous_text, next_text, and request-ID continuity fields to help adjacent chunks flow together.
Our test proves that one 1,399-character script completed as a valid file on the first attempt through the Anywhere route. It does not prove stability at the full 10,000-character ceiling or across an entire audiobook.
Long-form stability
Listen for: pace drift between the opening and closing passages, awkward pauses, clicks, clipping, or audible artifacts.
Objective result: the file decoded cleanly with no clipped samples or adjacent-sample jumps above 0.8. Machine timing estimated 2.858 words/s in the first half and 2.670 in the second, with no evidence of an overall second-half speed-up.
One run does not prove stability at the 10,000-character limit or rule out local pacing issues that require human listening.
Emotion and stability settings
In the native ElevenLabs API, lower stability allows more variation while higher stability can make delivery more consistent and potentially more monotonous. Style exaggeration can add expression but may also reduce predictability and increase compute. Treat those controls as creative tradeoffs, not universal quality upgrades.
Pronunciación
Multilingual v2 does not support phoneme tags. ElevenLabs recommends pronunciation dictionaries with aliases as a workaround. Normalize ambiguous numbers, abbreviations, dates, and URLs before generation, then keep a small regression script for names or terms that matter to your brand.
Pronunciation stress test
Listening checkpoint: offline speech recognition rendered the phone number as “1-800-5555-0199.” Listen to the original number and URL carefully before publication.
The extra digit is not a confirmed TTS error. Speech recognition can introduce mistakes, so this is a review checkpoint rather than a pronunciation verdict.
Latencia
Multilingual v2 is not ElevenLabs’ low-latency option. ElevenLabs says it has higher latency than Flash models. Our measured route times ranged from 10.161 to 31.489 seconds for the five formal tasks, but those figures include the Anywhere task route, network, queueing, and file delivery. They must not be treated as native ElevenLabs model latency.
How to use the Eleven Multilingual v2 API
Native ElevenLabs API request
Use the native model ID eleven_multilingual_v2 in the JSON body and place the selected voice ID in the endpoint.
https://api.elevenlabs.io/v1/text-to-speech/{voice_id}curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/YOUR_VOICE_ID" \
-H "xi-api-key: $ELEVENLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Your multilingual narration goes here.",
"model_id": "eleven_multilingual_v2",
"voice_settings": {
"stability": 0.5,
"similarity_boost": 0.75,
"style": 0,
"use_speaker_boost": true,
"speed": 1.0
}
}' \
--output narration.mp3Key parameters and limits
| Parámetro | What it controls | Nota práctica |
|---|---|---|
voice_id | Selected ElevenLabs voice | Choose a voice whose accent fits the target language and region. |
model_id | Synthesis model | Utilice eleven_multilingual_v2. |
estabilidad | Variation versus consistency | Lower can vary more; higher can become more restrained. |
similarity_boost | Reference-voice similarity | High values can help identity but do not guarantee it across languages. |
estilo | Style exaggeration | Use carefully when repeatability matters. |
use_speaker_boost | Extra speaker-similarity processing | May increase latency. |
velocidad | Playback pace | Retest numbers and punctuation after changing it. |
previous_text / next_text | Context around a chunk | Useful when splitting long scripts. |
language_code | Language enforcement | Not supported for multilingual_v2 models in the current API reference. |
Importante: keep API keys in a secure environment variable. Never paste a real key into public article code.
Eleven Multilingual v2 API pricing
ElevenLabs’ API pricing page listed Multilingual v2 and v3 at $0.10 per 1,000 characters when checked on August 11, 2026. Taxes, levies, and duties are excluded.
Direct API pricing
At the official rate of $0.10 per 1,000 characters for Multilingual v2 and v3:
ElevenLabs API pricing checked August 11, 2026. Taxes, levies, and duties are excluded.Direct API or GlobalGPT?
| Ruta | Pricing reference | How you work | La mejor relación calidad-precio cuando |
|---|---|---|---|
| ElevenLabs API | $0.10 / 1K characters | Native requests, voices, settings, and metered usage | You need developer controls and product integration |
| GlobalGPT Overall value pick | Basic $5.80/mo Pro $10.80/mo Unlimited $25/mo | Multiple AI models and creative tools in one browser platform, including Audio Generator access | Audio is part of a broader writing, research, image, video, or model-comparison workflow |
That rate makes a 10,000-character maximum-size request $1.00 before taxes. Actual project cost depends on how much text you generate, how often you regenerate, and whether you keep unsuccessful takes. Build your budget from input characters rather than audio minutes.
For creators who do not need to build a native API workflow, GlobalGPT offers a more practical value proposition. When checked on August 11, 2026, its annual plans worked out to $5.80 per month for Basic, $10.80 for Pro, and $25 for Unlimited, combining multiple AI models and creative tools in one subscription. The GlobalGPT Audio Generator includes Eleven Multilingual v2 and Eleven v3 in a browser workspace, so you can generate narration and compare both models without wiring an API request first.
GlobalGPT uses its own Credits system, so this is not a direct per-character price comparison with the ElevenLabs API. But when audio is part of a wider workflow involving writing, research, images, video, or multiple AI models, GlobalGPT is the stronger overall value choice. Puedes compare GlobalGPT plans here.
Eleven Multilingual v2 vs Eleven v3 vs Flash v2.5
Which ElevenLabs model should you choose?
Match the model to the job instead of choosing by release date alone.
Multilingual v2
Choose for: steady narration, established workflows, 29-language coverage, and scripts up to 10,000 characters.
Once v3
Choose for: dramatic delivery, multi-speaker work, and broader 70+ language coverage.
Flash v2.5
Choose for: responsive products, high throughput, 40,000-character requests, and lower API cost.
| Modelo | Languages | Character limit | Official emphasis | Mejor ajuste |
|---|---|---|---|---|
| Multilingual v2 | 29 | 10,000 | Long-form stability and consistent quality | Narration, courses, localization, established voice workflows |
| Once v3 | 70+ | 5,000 | Expressive speech and multi-speaker dialogue | Performance, characters, dramatic delivery, broader language coverage |
| Flash v2.5 | 32 | 40,000 | About 75 ms model latency and lower API cost | Interactive products, throughput, longer requests, cost-sensitive scale |
GlobalGPT also includes Eleven v3 in the same Audio Generator. Run the same short script through both models: use Multilingual v2 for the steady narration baseline, then compare v3 when you want more expressive delivery. Your own script is a better decision tool than a generic demo because it exposes the names, punctuation, pacing, and language switches that matter to your project.
Who should—and should not—use Multilingual v2?
Buen ajuste
- Narrators producing courses, explainers, corporate videos, or localized voiceovers.
- Teams that already have a validated ElevenLabs voice and want a stable model ID.
- Projects needing one of the 29 supported languages and scripts longer than v3’s 5,000-character limit.
- Developers who value native voice controls and chunk-continuity fields over the lowest possible latency.
Look elsewhere when
- You need highly expressive acting, natural multi-speaker dialogue, or a language outside the 29-language list; start with Eleven v3.
- You are building a responsive voice agent or high-throughput service; evaluate Flash v2.5.
- You need phoneme-level markup; Multilingual v2 does not support phoneme tags.
- You require a guaranteed regional accent without testing a matching voice and real script.
- You need a definitive subjective quality judgment before a qualified listener has reviewed your target languages.
Preguntas frecuentes
Is Eleven Multilingual v2 still available in 2026?
Yes. ElevenLabs listed it as a flagship Text to Speech model and did not include it in the deprecated-model table when checked on August 11, 2026.
When was Eleven Multilingual v2 released?
ElevenLabs announced Multilingual v2 on August 22, 2023. The “2026” in this review refers to the review date, not the launch year.
How many languages does Eleven Multilingual v2 support?
It officially supports 29 languages, including English, Spanish, Chinese, Japanese, German, French, Hindi, Korean, Portuguese, Arabic, and Russian.
What is the Eleven Multilingual v2 model ID?
The native ElevenLabs API model ID is eleven_multilingual_v2. The Anywhere test route used the separate hyphenated identifier eleven-multilingual-v2.
What is the character limit?
The official single-request limit is 10,000 characters, which ElevenLabs estimates at about ten minutes of audio.
Is Multilingual v2 better than Eleven v3?
Neither is universally better. Multilingual v2 is the steadier, longer-request choice; v3 offers broader language coverage, more expressive delivery, and multi-speaker capabilities.
Is Multilingual v2 good for long-form narration?
ElevenLabs positions it as its most stable long-form model. Our 1,399-character route test completed successfully, but that single run does not prove full-book or 10,000-character stability.
Can I set a language code in the API?
The current Create speech API reference says language_code is not supported for multilingual_v2 models.
Does Multilingual v2 support phoneme tags?
No. ElevenLabs recommends pronunciation dictionaries with aliases when you need to control the rendering of names or specialized terms.
How much does the API cost?
ElevenLabs listed Multilingual v2 at $0.10 per 1,000 characters on August 11, 2026, excluding taxes, levies, and duties.
Veredicto final
Eleven Multilingual v2 earns a clear place in the 2026 ElevenLabs lineup: it is the established choice for steady multilingual narration, a validated voice workflow, and requests up to 10,000 characters. It is not the automatic winner for every TTS project. Eleven v3 is the stronger starting point for expressive acting and broader languages, while Flash v2.5 is designed for speed and scale.
Our objective test adds useful confidence about route reliability and file integrity: all five formal Anywhere tasks completed on the first attempt, produced valid 44.1 kHz mono MP3s, and stayed within the pre-registered cost boundary. It does not replace human listening, so we are not assigning a subjective naturalness, emotion, or accent score from machine checks alone.
Bring a paragraph from your next narration, course, or localized video to GlobalGPT’s Audio Generator. Start with Eleven Multilingual v2, keep the voice and script consistent, and compare it with Eleven v3 when expressive delivery matters. It is a direct, practical way to choose the model that fits your work.




