{"id":20189,"date":"2026-09-29T09:38:58","date_gmt":"2026-09-29T13:38:58","guid":{"rendered":"https:\/\/wp.glbgpt.com\/?p=20189"},"modified":"2026-09-29T09:38:59","modified_gmt":"2026-09-29T13:38:59","slug":"openai-audio-speech-api","status":"publish","type":"post","link":"https:\/\/wp.glbgpt.com\/pt-br\/hub\/openai-audio-speech-api","title":{"rendered":"API de \u00e1udio e fala OpenAI: c\u00f3digo, vozes, formatos e custos"},"content":{"rendered":"<p class=\"wp-block-paragraph\">O <strong>OpenAI Audio Speech API<\/strong> turns text into spoken audio through <code>POST https:\/\/api.openai.com\/v1\/audio\/speech<\/code>. Send a model, your input text, and a voice; save the response as an MP3 or another supported audio format. For a new integration that needs instructions for tone and delivery, start with <code>gpt-4o-mini-tts<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A working request is only the first step. You also need the right billing unit, a playable output file, and a plan for long scripts. OpenAI\u2019s newer TTS model uses text and audio tokens for pricing; its older TTS models use characters. That difference matters when you estimate a month of narration.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If your job also includes writing the script, making visuals, and assembling a video, <a href=\"https:\/\/www.glbgpt.com\/audio-generator?inviter=hub_audio&amp;login=1\">Espa\u00e7o de trabalho de \u00e1udio do GlobalGPT<\/a> brings speech generation into an affordable multi-model subscription platform. You can work across text, voice, images, and video in one dashboard, then use its API and CLI for repeatable production work. Its public TTS API uses a separate task workflow, covered below.<\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<figure class=\"wp-block-image size-large\"><a href=\"https:\/\/www.glbgpt.com\/audio-generator?inviter=hub_content_audio&amp;login=1\"><img alt=\"\" fetchpriority=\"high\" decoding=\"async\" width=\"1024\" height=\"639\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-1024x639.png\" class=\"wp-image-20191\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-1024x639.png 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-300x187.png 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-768x479.png 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-18x11.png 18w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-1536x958.png 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/audio-2048x1278.png 2048w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/figure>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-3e41869c wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link has-background has-medium-font-size has-custom-font-size wp-element-button\" href=\"https:\/\/www.glbgpt.com\/audio-generator?inviter=hub_content_audio&amp;login=1\" style=\"background:linear-gradient(265deg,rgb(255,245,203) 0%,rgb(182,227,212) 16%,rgb(51,167,181) 100%)\"><strong>Experimente v\u00e1rios modelos de \u00e1udio agora<\/strong><\/a><\/div>\n<\/div>\n<\/div>\n\n\n\n<nav class=\"gg-speech-toc\" aria-label=\"Conte\u00fado do artigo\" style=\"box-sizing:border-box;font:16px\/1.55 system-ui,sans-serif;padding:24px;background:#fff;border:1px solid #c9d9df;border-radius:8px;margin:28px 0;color:#163641!important;max-width:100%;min-width:0;\"><h2 style=\"box-sizing:border-box;font:750 21px\/1.3 system-ui,sans-serif;color:#163641!important;margin:0 0 15px;line-height:1.3;\">Jump to a step<\/h2><ol style=\"box-sizing:border-box;padding-left:22px;margin:0;columns:18rem 2;column-gap:35px;\"><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#what-is-speech-api\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">What does the OpenAI Audio Speech API do?<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#openai-speech-example\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Generate your first audio file with the OpenAI Speech API<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#models-voices-formats\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Choose a model, voice, and output format<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#speech-api-pricing\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">What does the OpenAI Speech API cost?<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#globalgpt-tts-api\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Build a TTS workflow with the GlobalGPT API<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#tts-audio-examples\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Listen to two product-tour examples<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#speech-api-errors\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Fix common speech-generation problems<\/a><\/li><li style=\"box-sizing:border-box;break-inside:avoid;padding:0 0 10px;color:#365462!important;\"><a href=\"#speech-api-faq\" style=\"box-sizing:border-box;color:#096570!important;text-decoration:underline;text-underline-offset:3px;\">Perguntas frequentes<\/a><\/li><\/ol><\/nav>\n\n\n\n<h2 id=\"what-is-speech-api\" class=\"wp-block-heading\">What does the OpenAI Audio Speech API do?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The Speech API is a text-to-speech endpoint: your application supplies written words and receives generated speech. Typical uses include product walkthroughs, accessibility narration, learning material, and short voiceovers. You control the text before it is spoken, which makes it useful when the wording must follow an approved script.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For transcription, the direction is reversed: recorded audio goes in and text comes out. A live voice assistant adds another requirement\u2014listening and responding through a conversation. OpenAI\u2019s <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/audio\">audio documentation<\/a> separates these workflows. Choose the Speech endpoint when you already have the text you want read aloud.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Speech generation supports streaming, so an application can receive audio in chunks. Saving that stream to a file does not automatically play it in a browser or create a conversational agent; your application still handles playback and any conversation logic. OpenAI also requires clear disclosure to listeners that the voice is AI-generated.<\/p>\n\n\n\n<h2 id=\"openai-speech-example\" class=\"wp-block-heading\">Generate your first audio file with the OpenAI Speech API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Run the request on a server or your development machine, with an OpenAI API key stored in the <code>OPENAI_API_KEY<\/code> environment variable. Keep the key out of browser JavaScript. The following examples use the same short product-tour script and save the result as <code>openai-product-tour.mp3<\/code>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Option 1: make a curl request<\/h3>\n\n\n\n<section class=\"gg-code-openai-speech-sh\" style=\"box-sizing:border-box;background:#102c3b;color:#f1f7fb!important;border-radius:7px;padding:20px;margin:25px 0;font:15px\/1.6 system-ui,sans-serif;max-width:100%;min-width:0;\"><h3 style=\"box-sizing:border-box;color:#fff!important;font:700 20px\/1.4 system-ui,sans-serif;margin:0 0 14px;line-height:1.3;\">OpenAI Speech API \u00b7 curl<\/h3><pre style=\"box-sizing:border-box;tab-size:4;white-space:pre-wrap;overflow-x:auto;max-width:100%;margin:0;background:transparent;color:#edf6fc!important;font:13px\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;\"><code style=\"box-sizing:border-box;font:inherit;color:inherit;background:transparent;\">#!\/bin\/sh\n# Set OPENAI_API_KEY in your server environment first.\ncurl --fail-with-body --show-error \\\n  https:\/\/api.openai.com\/v1\/audio\/speech \\\n  -H &quot;Authorization: Bearer $OPENAI_API_KEY&quot; \\\n  -H &quot;Content-Type: application\/json&quot; \\\n  -d &#x27;{\n    &quot;model&quot;: &quot;gpt-4o-mini-tts&quot;,\n    &quot;input&quot;: &quot;Welcome to the Acorn Studio product tour. In three steps, you can turn a short script into an audio guide. First, write the message. Next, choose a voice. Finally, save the audio file and test playback on your phone.&quot;,\n    &quot;voice&quot;: &quot;coral&quot;,\n    &quot;instructions&quot;: &quot;Speak clearly in a warm, conversational tone.&quot;,\n    &quot;response_format&quot;: &quot;mp3&quot;\n  }&#x27; \\\n  --output openai-product-tour.mp3\n<\/code><\/pre><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The response contains audio bytes, so save it to a file rather than attempting to parse it as JSON. The <code>--fail-with-body<\/code> option gives curl a failing exit status for an HTTP error. Check that status before opening the file: an error response can still leave JSON in the file named <code>.mp3<\/code>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Option 2: use the Python SDK<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Install the OpenAI package with <code>python -m pip install openai<\/code>. The SDK reads the API key from the environment. This example writes the response as it arrives, which avoids loading the entire audio response into your own Python buffer.<\/p>\n\n\n\n<section class=\"gg-code-openai-speech-py\" style=\"box-sizing:border-box;background:#102c3b;color:#f1f7fb!important;border-radius:7px;padding:20px;margin:25px 0;font:15px\/1.6 system-ui,sans-serif;max-width:100%;min-width:0;\"><h3 style=\"box-sizing:border-box;color:#fff!important;font:700 20px\/1.4 system-ui,sans-serif;margin:0 0 14px;line-height:1.3;\">OpenAI Speech API \u00b7 Python<\/h3><pre style=\"box-sizing:border-box;tab-size:4;white-space:pre-wrap;overflow-x:auto;max-width:100%;margin:0;background:transparent;color:#edf6fc!important;font:13px\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;\"><code style=\"box-sizing:border-box;font:inherit;color:inherit;background:transparent;\">&quot;&quot;&quot;Install: python -m pip install openai. Set OPENAI_API_KEY on the server.&quot;&quot;&quot;\nfrom pathlib import Path\nfrom openai import OpenAI\n\nclient = OpenAI(max_retries=0)\nscript = (\n    &quot;Welcome to the Acorn Studio product tour. In three steps, you can turn &quot;\n    &quot;a short script into an audio guide. First, write the message. Next, &quot;\n    &quot;choose a voice. Finally, save the audio file and test playback on your phone.&quot;\n)\n\nwith client.audio.speech.with_streaming_response.create(\n    model=&quot;gpt-4o-mini-tts&quot;,\n    voice=&quot;coral&quot;,\n    input=script,\n    instructions=&quot;Speak clearly in a warm, conversational tone.&quot;,\n    response_format=&quot;mp3&quot;,\n) as response:\n    response.stream_to_file(Path(&quot;openai-product-tour.mp3&quot;))\n\nprint(&quot;Saved openai-product-tour.mp3&quot;)\n<\/code><\/pre><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The three essential fields are <code>modelo<\/code>, <code>entrada<\/code>, e <code>voice<\/code>. Here, <code>instructions<\/code> asks for a warm, conversational delivery; it guides the voice rather than adding words to the spoken script. This field works with GPT-4o Mini TTS and is unsupported by <code>tts-1<\/code> e <code>tts-1-hd<\/code>.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/mcp\/3\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936.webp\"><img decoding=\"async\" width=\"1440\" height=\"900\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936.webp\" alt=\"OpenAI text-to-speech documentation with the speech request parameters and a Python quickstart.\" class=\"wp-image-20195\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936.webp 1440w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936-300x188.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936-1024x640.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936-768x480.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-speech-quickstart-20260928_5609ae78b17743b28dcd80ab17ba3936-18x12.webp 18w\" sizes=\"(max-width: 1440px) 100vw, 1440px\" \/><\/a><figcaption class=\"wp-element-caption\">OpenAI\u2019s official text-to-speech quickstart. Select the image to inspect the code at full size.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For a longer document, split the text at sentence or paragraph boundaries, number the segments, and save each output with that number. Keep the same voice and instructions across segments. Leave enough context for natural phrasing, and listen across the joins before distributing the assembled recording.<\/p>\n\n\n\n<h2 id=\"models-voices-formats\" class=\"wp-block-heading\">Choose a model, voice, and output format<\/h2>\n\n\n\n<figure class=\"wp-block-table\" style=\"box-sizing:border-box;margin:28px 0;overflow-x:auto;font:16px\/1.65 system-ui,-apple-system,sans-serif;max-width:100%;min-width:0;color:#17323f;\"><table style=\"box-sizing:border-box;border-collapse:collapse;width:100%;min-width:530px;font-size:16px;line-height:1.6;\"><thead style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Modelo<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Como escolher<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Delivery instructions<\/th><\/tr><\/thead><tbody style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">gpt-4o-mini-tts<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">A starting point for controllable narration and spoken delivery.<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Suportes <code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">instructions<\/code> for tone, style, and pace.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">tts-1<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">OpenAI positions this model for lower-latency speech generation.<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Use its supported voice and speed controls; omit <code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">instructions<\/code>.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">tts-1-hd<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">OpenAI positions this model for higher-quality speech than TTS-1.<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Use its supported voice and speed controls; omit <code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">instructions<\/code>.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">O <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-4o-mini-tts\">GPT-4o Mini TTS model page<\/a> maps the model alias to the <code>gpt-4o-mini-tts-2025-12-15<\/code> snapshot. If reproducible behavior is important to your application, choose a documented snapshot deliberately and record it alongside your saved scripts.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Built-in voices and input limits<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The Speech API reference lists 13 built-in voices for GPT-4o Mini TTS: <strong>alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar<\/strong>. OpenAI recommends <strong>marin<\/strong> ou <strong>cedar<\/strong> for best quality. That is OpenAI\u2019s recommendation; choose your production voice by listening to your own names, numbers, and sentence patterns.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The earlier TTS models support a smaller set: alloy, ash, coral, echo, fable, onyx, nova, sage, and shimmer. Do not assume every voice works with every model. The <a href=\"https:\/\/developers.openai.com\/api\/reference\/resources\/audio\/subresources\/speech\/methods\/create\">Speech API reference<\/a> also accepts a speed between <strong>0.25 and 4.0<\/strong>, com <strong>1.0<\/strong> as the default.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There are two input limits to account for: the endpoint reference limits <code>entrada<\/code> para <strong>4,096 characters<\/strong>, and the Mini TTS model page specifies a maximum of <strong>2,000 input tokens<\/strong>. Characters and tokens are different measurements. Check both constraints for Mini TTS, especially with multilingual text, and split large scripts before submitting them.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Choose the file format for its destination<\/h3>\n\n\n\n<figure class=\"wp-block-table\" style=\"box-sizing:border-box;margin:28px 0;overflow-x:auto;font:16px\/1.65 system-ui,-apple-system,sans-serif;max-width:100%;min-width:0;color:#17323f;\"><table style=\"box-sizing:border-box;border-collapse:collapse;width:100%;min-width:530px;font-size:16px;line-height:1.6;\"><thead style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Formato<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Aplica\u00e7\u00e3o pr\u00e1tica<\/th><\/tr><\/thead><tbody style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">MP3<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Default output; a convenient compressed file for sharing and web playback.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Opus<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Useful for streaming and communication applications that support the codec.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">AAC<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Consider when your playback environment or media workflow calls for AAC.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">FLAC<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Lossless compression for workflows that need to retain the audio signal.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">WAV<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Uncompressed audio in a familiar container for editing and processing.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">PCM<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Raw 24 kHz, 16-bit signed little-endian samples with no file header; your player needs those settings.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Start with MP3 for a download link or a simple audio player. Choose WAV when the next step is editing. Select raw PCM only when your application knows how to interpret it; changing a filename from <code>.pcm<\/code> para <code>.wav<\/code> does not create a WAV header.<\/p>\n\n\n\n<h2 id=\"speech-api-pricing\" class=\"wp-block-heading\">What does the OpenAI Speech API cost?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">As checked on September 28, 2026, OpenAI lists GPT-4o Mini TTS at <strong>$0.60 per million text input tokens<\/strong> mais <strong>$12 per million audio output tokens<\/strong>. TTS-1 costs <strong>$15 per million characters<\/strong>, and TTS-1 HD costs <strong>$30 per million characters<\/strong>. Calculate each model in its own units.<\/p>\n\n\n\n<section class=\"gg-speech-prices\" aria-labelledby=\"ggsp-title\" style=\"box-sizing:border-box;font:16px\/1.55 system-ui,sans-serif;color:#132b39!important;background:#f5faf9;border:1px solid #cbded9;border-radius:8px;padding:24px;margin:28px 0;max-width:100%;min-width:0;\">\n<h3 id=\"ggsp-title\" style=\"box-sizing:border-box;font:700 23px\/1.3 system-ui,sans-serif;color:#132b39!important;margin:0 0 8px;line-height:1.3;\">Speech API pricing: keep the units separate<\/h3>\n<p style=\"box-sizing:border-box;color:#284654!important;margin:8px 0;\">OpenAI bills its newer TTS model by tokens and its two earlier TTS models by characters. GlobalGPT lists its TTS tasks in credits.<\/p>\n<div class=\"ggsp-scroll\" style=\"box-sizing:border-box;overflow-x:auto;\"><table style=\"box-sizing:border-box;width:100%;border-collapse:collapse;min-width:540px;margin:18px 0;\">\n<thead style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><th style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;background:#dceee9;font-weight:700;border:1px solid #c9d7df;\">Route and model<\/th><th style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;background:#dceee9;font-weight:700;border:1px solid #c9d7df;\">Taxa<\/th><th style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;background:#dceee9;font-weight:700;border:1px solid #c9d7df;\">Exemplo<\/th><\/tr><\/thead>\n<tbody style=\"box-sizing:border-box;\">\n<tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">OpenAI \u00b7 GPT-4o Mini TTS<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$0.60 \/ 1M text input tokens<br style=\"box-sizing:border-box;\">$12 \/ 1M audio output tokens<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$0.0606 for 1,000 text input tokens + 5,000 audio output tokens<\/td><\/tr>\n<tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">OpenAI \u00b7 TTS-1<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$15 \/ 1M characters<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$0.015 for 1,000 characters<\/td><\/tr>\n<tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">OpenAI \u00b7 TTS-1 HD<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$30 \/ 1M characters<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">$0.030 for 1,000 characters<\/td><\/tr>\n<tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">GlobalGPT \u00b7 Eleven v3, Eleven Multilingual v2, or Qwen Audio 3.0 TTS Flash<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">330 credits \/ 1,000 characters<\/td><td style=\"box-sizing:border-box;text-align:left;padding:13px 12px;border-bottom:1px solid #cbded9;vertical-align:top;color:#132b39!important;border:1px solid #c9d7df;\">330 credits at the listed 1,000-character rate; request an estimate before submitting<\/td><\/tr>\n<\/tbody><\/table><\/div>\n<p class=\"ggsp-note\" style=\"box-sizing:border-box;color:#284654!important;margin:8px 0;font-size:14px;\">These are pricing examples, not charges from the audio samples below. Tokens, characters, minutes, and credits are different units. The GlobalGPT task estimate supplies the applicable credit amount; the table does not convert credits into dollars.<\/p>\n<small style=\"box-sizing:border-box;display:block;font-size:13px;color:#3c5561!important;\">Rates checked September 28, 2026. Sources: <a href=\"https:\/\/developers.openai.com\/api\/docs\/pricing\" style=\"box-sizing:border-box;color:#075b63!important;text-decoration:underline;\">Pre\u00e7os da API OpenAI<\/a> e <a href=\"https:\/\/www.glbgpt.com\/home\/api\/docs\" style=\"box-sizing:border-box;color:#075b63!important;text-decoration:underline;\">GlobalGPT API documentation<\/a>.<\/small>\n<\/section>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/mcp\/3\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b.webp\"><img decoding=\"async\" width=\"1800\" height=\"1400\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b.webp\" alt=\"OpenAI pricing tables showing GPT-4o Mini TTS input and output token rates and the TTS character rates.\" class=\"wp-image-20194\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b.webp 1800w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b-300x233.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b-1024x796.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b-768x597.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b-1536x1195.webp 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/openai-tts-pricing-20260928_ae89c6abd2a349fda0d60b7766ba6c4b-15x12.webp 15w\" sizes=\"(max-width: 1800px) 100vw, 1800px\" \/><\/a><figcaption class=\"wp-element-caption\">OpenAI API speech pricing, captured September 28, 2026. Token-based and character-based charges use different units.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For example, 100 scripts of 1,000 characters each total 100,000 characters. At the listed rates, that is <strong>$1.50 with TTS-1<\/strong> ou <strong>$3.00 with TTS-1 HD<\/strong>, before other application costs. The same character count is insufficient to calculate Mini TTS exactly: you need the text input and audio output token quantities.<\/p>\n\n\n\n<section class=\"gg-speech-calc\" id=\"gg-speech-calc\" aria-labelledby=\"ggsc-title\" style=\"box-sizing:border-box;font:16px\/1.55 system-ui,sans-serif;color:#17323f!important;background:#fff;border:1px solid #bdd7d5;border-top:5px solid #147d7d;border-radius:8px;padding:24px;margin:28px 0;max-width:100%;min-width:0;\">\n<h3 id=\"ggsc-title\" style=\"box-sizing:border-box;font:700 23px\/1.3 system-ui,sans-serif;color:#17323f!important;margin:0 0 10px;line-height:1.3;\">Estimate your speech-generation cost<\/h3>\n<p style=\"box-sizing:border-box;color:#294b59!important;margin:12px 0;\">Enter the units your chosen route bills. There is no automatic character-to-token or minute-to-token conversion.<\/p>\n<label for=\"ggsc-mode\" style=\"box-sizing:border-box;color:#294b59!important;display:block;font-weight:650;margin:16px 0 5px;\">Model and billing method<\/label>\n<select id=\"ggsc-mode\" style=\"box-sizing:border-box;width:100%;max-width:100%;padding:11px;border:1px solid #789ba3;border-radius:5px;background:#fff;color:#17323f;font:inherit;\"><option value=\"tts1\" style=\"box-sizing:border-box;\">OpenAI TTS-1 \u00b7 characters<\/option><option value=\"hd\" style=\"box-sizing:border-box;\">OpenAI TTS-1 HD \u00b7 characters<\/option><option value=\"mini\" style=\"box-sizing:border-box;\">OpenAI GPT-4o Mini TTS \u00b7 text and audio tokens<\/option><option value=\"glb\" style=\"box-sizing:border-box;\">GlobalGPT TTS \u00b7 character-based credit estimate<\/option><\/select>\n<div class=\"ggsc-grid\" style=\"box-sizing:border-box;display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,240px),1fr));gap:16px;\">\n<div style=\"box-sizing:border-box;\"><label for=\"ggsc-first\" id=\"ggsc-first-label\" style=\"box-sizing:border-box;color:#294b59!important;display:block;font-weight:650;margin:16px 0 5px;\">Input characters<\/label><input id=\"ggsc-first\" type=\"number\" min=\"0\" max=\"1000000000\" step=\"1\" value=\"1000\" inputmode=\"numeric\" style=\"box-sizing:border-box;width:100%;max-width:100%;padding:11px;border:1px solid #789ba3;border-radius:5px;background:#fff;color:#17323f;font:inherit;\"><\/div>\n<div id=\"ggsc-audio-field\" hidden style=\"box-sizing:border-box;\"><label for=\"ggsc-audio\" style=\"box-sizing:border-box;color:#294b59!important;display:block;font-weight:650;margin:16px 0 5px;\">Output audio tokens<\/label><input id=\"ggsc-audio\" type=\"number\" min=\"0\" max=\"1000000000\" step=\"1\" value=\"5000\" inputmode=\"numeric\" style=\"box-sizing:border-box;width:100%;max-width:100%;padding:11px;border:1px solid #789ba3;border-radius:5px;background:#fff;color:#17323f;font:inherit;\"><\/div>\n<\/div>\n<output class=\"ggsc-result\" id=\"ggsc-result\" aria-live=\"polite\" style=\"box-sizing:border-box;display:block;padding:18px;background:#e5f3ef;color:#103d38!important;border-radius:5px;font-weight:750;font-size:22px;margin-top:20px;overflow-wrap:anywhere;\">Estimated cost: $0.015<\/output>\n<p class=\"ggsc-help\" id=\"ggsc-help\" style=\"box-sizing:border-box;color:#294b59!important;font-size:14px;margin:12px 0 0;\">Formula: characters \u00d7 $15 \u00f7 1,000,000. A text total may span several API requests.<\/p>\n<p class=\"ggsc-help\" style=\"box-sizing:border-box;color:#294b59!important;font-size:14px;margin:12px 0 0;\">Rates checked September 28, 2026: <a href=\"https:\/\/developers.openai.com\/api\/docs\/pricing\" style=\"box-sizing:border-box;color:#09616b!important;text-decoration:underline;\">OpenAI<\/a> \u00b7 <a href=\"https:\/\/www.glbgpt.com\/home\/api\/docs\" style=\"box-sizing:border-box;color:#09616b!important;text-decoration:underline;\">GlobalGPT<\/a>. Excludes storage, hosting, taxes, and other parts of your application.<\/p>\n<noscript style=\"box-sizing:border-box;\"><p style=\"box-sizing:border-box;color:#294b59!important;margin:12px 0;\">JavaScript is required for the calculator. TTS-1: characters \u00d7 $15 \/ 1M. TTS-1 HD: characters \u00d7 $30 \/ 1M. GPT-4o Mini TTS: text input tokens \u00d7 $0.60 \/ 1M + audio output tokens \u00d7 $12 \/ 1M. GlobalGPT: characters \u00d7 330 \/ 1,000 credits; use its estimate endpoint for the actual request.<\/p><\/noscript>\n<\/section>\n<script>\n(function(){\n  const root=document.getElementById('gg-speech-calc');\n  if(!root || root.dataset.ready) return;\n  root.dataset.ready='true';\n  const mode=root.querySelector('#ggsc-mode'), first=root.querySelector('#ggsc-first'), audio=root.querySelector('#ggsc-audio');\n  const result=root.querySelector('#ggsc-result'), help=root.querySelector('#ggsc-help');\n  function valid(raw){return raw!=='' && Number.isSafeInteger(Number(raw)) && Number(raw)>=0 && Number(raw)<=1000000000;}\n  function number(value,places){return value.toLocaleString('en-US',{minimumFractionDigits:0,maximumFractionDigits:places});}\n  function update(){\n    const mini=mode.value==='mini';\n    root.querySelector('#ggsc-audio-field').hidden=!mini;\n    root.querySelector('#ggsc-first-label').textContent=mini?'Input text tokens':'Input characters';\n    if(!valid(first.value) || (mini&#038;&#038;!valid(audio.value))){result.textContent='Enter whole numbers from 0 to 1,000,000,000.';help.textContent='Empty values, decimals, and negative counts cannot be used.';return;}\n    const a=Number(first.value), b=Number(audio.value);\n    if(mini){result.textContent='Estimated cost: $'+number((a*0.6+b*12)\/1000000,8);help.textContent='Formula: text input tokens \u00d7 $0.60 \/ 1M + audio output tokens \u00d7 $12 \/ 1M. Use token counts or explicit budget assumptions, not text length.';}\n    else if(mode.value==='glb'){result.textContent='Proportional estimate: '+number(a*330\/1000,3)+' credits';help.textContent='Based on 330 credits \/ 1,000 characters. This does not model rounding or minimum charges. Call POST \/tasks\/estimate for your actual request; credits are not dollars.';}\n    else{const rate=mode.value==='hd'?30:15;result.textContent='Estimated cost: $'+number(a*rate\/1000000,8);help.textContent='Formula: characters \u00d7 $'+rate+' \u00f7 1,000,000. A text total may span several API requests.';}\n  }\n  mode.addEventListener('change',update);first.addEventListener('input',update);audio.addEventListener('input',update);update();\n})();\n<\/script>\n\n\n\n<p class=\"wp-block-paragraph\">For planning, enter actual usage counts or explicit assumptions in the calculator. Treat the Mini TTS audio-token field as a budget input until you have measured usage. Do not turn an approximate per-minute figure into a guaranteed price for every voice, speaking pace, or language.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The GlobalGPT model catalog expresses TTS prices in credits, and its estimate endpoint gives a quote for the request you intend to send. Keep API credit spending separate from choosing a multi-model workspace subscription for everyday writing and media work. This makes the budget easier to follow without assuming the two products share one billing allowance.<\/p>\n\n\n\n<h2 id=\"globalgpt-tts-api\" class=\"wp-block-heading\">Build a TTS workflow with the GlobalGPT API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For a repeatable script-to-voice-to-video workflow, <a href=\"https:\/\/www.glbgpt.com\/audio-generator?inviter=hub_audio&amp;login=1\">create speech in GlobalGPT<\/a> and keep the rest of the project in the same workspace. The <a href=\"https:\/\/www.glbgpt.com\/home\/api\">GlobalGPT API<\/a> exposes media generation as tasks, while the <a href=\"https:\/\/www.glbgpt.com\/home\/cli\">GlobalGPT CLI<\/a> connects that work with your terminal and development tools.<\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<figure class=\"wp-block-image size-large\"><a href=\"https:\/\/www.glbgpt.com\/home\/api?inviter=hub_api&amp;login=1\"><img alt=\"\" loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"702\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-1024x702.png\" class=\"wp-image-20178\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-1024x702.png 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-300x206.png 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-767x526.png 767w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-17x12.png 17w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api-1536x1053.png 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/api.png 1938w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/figure>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-3e41869c wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/www.glbgpt.com\/home\/api?inviter=hub_api&amp;login=1\">Get API Key Now<\/a><\/div>\n<\/div>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Use the public base URL <code>https:\/\/api2.glbgpt.com\/ai-api\/open\/v1<\/code>. Send TTS requests to <code>POST \/tasks<\/code>, query <code>GET \/tasks\/{id}<\/code>, and download <code>output.url<\/code> after the task succeeds. The text field is <code>prompt<\/code>. Implement this task contract directly; it differs from OpenAI\u2019s immediate audio response.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/mcp\/3\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7.webp\"><img loading=\"lazy\" decoding=\"async\" width=\"1440\" height=\"680\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1.webp\" alt=\"GlobalGPT API documentation showing media-task submission, polling, and output retrieval.\" class=\"wp-image-20198\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1.webp 1440w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1-300x142.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1-1024x484.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1-768x363.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-media-tasks-20260928_a012b9ace6664ba99a8e9964f8d24fb7-1-18x9.webp 18w\" sizes=\"(max-width: 1440px) 100vw, 1440px\" \/><\/a><figcaption class=\"wp-element-caption\">GlobalGPT media tasks follow a submit, poll, and retrieve workflow.<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Pick a documented GlobalGPT TTS model<\/h3>\n\n\n\n<figure class=\"wp-block-table\" style=\"box-sizing:border-box;margin:28px 0;overflow-x:auto;font:16px\/1.65 system-ui,-apple-system,sans-serif;max-width:100%;min-width:0;color:#17323f;\"><table style=\"box-sizing:border-box;border-collapse:collapse;width:100%;min-width:530px;font-size:16px;line-height:1.6;\"><thead style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">ID do modelo<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Entrada e sa\u00edda<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Useful distinction<\/th><\/tr><\/thead><tbody style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">eleven-v3<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Up to 5,000 characters; MP3<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Supports expressive audio tags.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">onze-multil\u00edngue-v2<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Up to 10,000 characters; MP3<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">29 languages; audio tags are read as text.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\"><code style=\"box-sizing:border-box;font:0.85em\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;background:#eef3f5;padding:2px 4px;border-radius:3px;color:#18364a;\">qwen-audio-3.0-tts-flash<\/code><\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Up to 4,096 characters; WAV<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">49 voices; voice and language controls use Qwen-specific fields.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">All three entries list <strong>330 cr\u00e9ditos por 1.000 caracteres<\/strong>. Usar <code>POST \/tasks\/estimate<\/code> with the intended request body for the applicable estimate; the documentation says this creates no task and reserves no credits. Public API access uses paid credits.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/mcp\/3\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434.webp\"><img loading=\"lazy\" decoding=\"async\" width=\"1800\" height=\"995\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434.webp\" alt=\"GlobalGPT model documentation for Eleven v3 and Eleven Multilingual v2, including limits, audio format, and credit rates.\" class=\"wp-image-20197\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434.webp 1800w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434-300x166.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434-1024x566.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434-768x425.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434-1536x849.webp 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-eleven-audio-models-20260928_41a26c4c4ce341e59f1160af00aa2434-18x10.webp 18w\" sizes=\"(max-width: 1800px) 100vw, 1800px\" \/><\/a><figcaption class=\"wp-element-caption\">The ElevenLabs entries specify different text limits and audio-tag behavior.<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/static.futureshareai.com\/glb_features\/mcp\/3\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a.webp\"><img loading=\"lazy\" decoding=\"async\" width=\"1440\" height=\"555\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a.webp\" alt=\"GlobalGPT Qwen Audio 3.0 TTS Flash model entry with WAV output, 49 voices, and a 4096-character limit.\" class=\"wp-image-20196\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a.webp 1440w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a-300x116.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a-1024x395.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a-768x296.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/09\/globalgpt-qwen-audio-model-20260928_61b0fe72a16343d69d22c00d2bedf30a-18x7.webp 18w\" sizes=\"(max-width: 1440px) 100vw, 1440px\" \/><\/a><figcaption class=\"wp-element-caption\">Qwen Audio 3.0 TTS Flash has its own model ID and voice parameters.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For the broader creation workflow, see <a href=\"https:\/\/www.glbgpt.com\/hub\/text-to-speech\/\">text-to-speech workflows in GlobalGPT<\/a>. If you are researching the separately named open-source Qwen3-TTS family, read <a href=\"https:\/\/www.glbgpt.com\/hub\/qwen3-tts-review\/\">Qwen3-TTS models and API options<\/a>; do not treat the Qwen-Audio model ID above as an interchangeable Qwen3-TTS deployment.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Estimate, submit once, then poll the saved task<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Instalar <code>requests<\/code> com <code>python -m pip install requests<\/code> e definir <code>GLOBALGPT_API_KEY<\/code> on the server. The example below uses Eleven v3, selects a documented public voice ID, and sets an illustrative ceiling of 500 credits. Change that ceiling to match your own budget before running it.<\/p>\n\n\n\n<section class=\"gg-code-globalgpt-tts-py\" style=\"box-sizing:border-box;background:#102c3b;color:#f1f7fb!important;border-radius:7px;padding:20px;margin:25px 0;font:15px\/1.6 system-ui,sans-serif;max-width:100%;min-width:0;\"><h3 style=\"box-sizing:border-box;color:#fff!important;font:700 20px\/1.4 system-ui,sans-serif;margin:0 0 14px;line-height:1.3;\">GlobalGPT public TTS task API \u00b7 Python<\/h3><pre style=\"box-sizing:border-box;tab-size:4;white-space:pre-wrap;overflow-x:auto;max-width:100%;margin:0;background:transparent;color:#edf6fc!important;font:13px\/1.6 ui-monospace,monospace;overflow-wrap:anywhere;\"><code style=\"box-sizing:border-box;font:inherit;color:inherit;background:transparent;\">&quot;&quot;&quot;Install requests; set GLOBALGPT_API_KEY. Uses GlobalGPT&#x27;s public task API.&quot;&quot;&quot;\nimport json\nimport os\nimport time\nimport uuid\nfrom pathlib import Path\nimport requests\n\nBASE = &quot;https:\/\/api2.glbgpt.com\/ai-api\/open\/v1&quot;\nSTATE = Path(&quot;globalgpt-tts-job.json&quot;)\nOUTPUT = Path(&quot;globalgpt-product-tour.mp3&quot;)\nMAX_CREDITS = 500  # Your per-job ceiling; edit before running.\n\npayload = {\n    &quot;model&quot;: &quot;eleven-v3&quot;,\n    &quot;prompt&quot;: (\n        &quot;Welcome to the Acorn Studio product tour. In three steps, you can turn &quot;\n        &quot;a short script into an audio guide. First, write the message. Next, &quot;\n        &quot;choose a voice. Finally, save the audio file and test playback on your phone.&quot;\n    ),\n    &quot;metadata&quot;: {\n        &quot;voice_id&quot;: &quot;JBFqnCBsd6RMkjVDRZzb&quot;,\n        &quot;stability&quot;: 0.5,\n        &quot;language_code&quot;: &quot;en&quot;,\n    },\n}\n\napi = requests.Session()\napi.headers[&quot;Authorization&quot;] = &quot;Bearer &quot; + os.environ[&quot;GLOBALGPT_API_KEY&quot;]\n\nif STATE.exists():\n    job = json.loads(STATE.read_text())\n    if job[&quot;payload&quot;] != payload:\n        raise RuntimeError(&quot;This job belongs to different input. Use a new state filename.&quot;)\nelse:\n    estimate = api.post(BASE + &quot;\/tasks\/estimate&quot;, json=payload, timeout=30)\n    estimate.raise_for_status()\n    quote = estimate.json()\n    print(&quot;Estimated credits:&quot;, quote[&quot;credits&quot;])\n    if quote[&quot;credits&quot;] &gt; MAX_CREDITS:\n        raise RuntimeError(&quot;Estimate exceeds your per-job ceiling.&quot;)\n    job = {&quot;payload&quot;: payload, &quot;idempotency_key&quot;: str(uuid.uuid4())}\n    STATE.write_text(json.dumps(job, indent=2))\n\nif &quot;task_id&quot; not in job:\n    # The persisted key keeps a transport retry tied to this exact request.\n    submitted = api.post(\n        BASE + &quot;\/tasks&quot;, json=payload,\n        headers={&quot;Idempotency-Key&quot;: job[&quot;idempotency_key&quot;]}, timeout=60,\n    )\n    submitted.raise_for_status()\n    task = submitted.json()\n    job[&quot;task_id&quot;] = task[&quot;id&quot;]\n    STATE.write_text(json.dumps(job, indent=2))\n    if task.get(&quot;ignored_parameters&quot;):\n        print(&quot;Ignored parameters:&quot;, task[&quot;ignored_parameters&quot;])\n\nprint(&quot;Task:&quot;, job[&quot;task_id&quot;])\ndeadline = time.monotonic() + 600\nwhile time.monotonic() &lt; deadline:\n    checked = api.get(BASE + &quot;\/tasks\/&quot; + str(job[&quot;task_id&quot;]), timeout=30)\n    checked.raise_for_status()\n    task = checked.json()\n    if task.get(&quot;ignored_parameters&quot;):\n        print(&quot;Ignored parameters:&quot;, task[&quot;ignored_parameters&quot;])\n    if task[&quot;status&quot;] == &quot;failed&quot;:\n        raise RuntimeError(&quot;Generation failed: &quot; + json.dumps(task.get(&quot;error&quot;, {})))\n    if task[&quot;status&quot;] == &quot;succeeded&quot;:\n        # Use a fresh request: do not send the API key to the media host.\n        audio = requests.get(task[&quot;output&quot;][&quot;url&quot;], timeout=60)\n        audio.raise_for_status()\n        OUTPUT.write_bytes(audio.content)\n        print(&quot;Saved:&quot;, OUTPUT)\n        break\n    time.sleep(10)\nelse:\n    raise TimeoutError(&quot;Task still pending. Rerun to poll the saved task ID.&quot;)\n<\/code><\/pre><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The job file saves the exact request, an idempotency key, and the task ID. If your connection drops, rerun the same job so the code can resume the existing task. To intentionally generate different audio, use a new job filename. Keep the job file private if its script contains confidential information.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Verificar <code>ignored_parameters<\/code> so an unsupported setting does not silently become your assumed configuration. ElevenLabs controls such as <code>voice_id<\/code> differ from Qwen\u2019s <code>voice<\/code> e <code>language_type<\/code> fields. Poll every 5\u201310 seconds, handle a <code>failed<\/code> status, and save successful files promptly: the documented result retention is <strong>30 dias<\/strong>.<\/p>\n\n\n\n<h2 id=\"tts-audio-examples\" class=\"wp-block-heading\">Listen to two product-tour examples<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">These two AI-generated samples use the same 216-character product-tour script with Eleven v3 and Qwen Audio 3.0 TTS Flash. Each is the first output from one request, with default voice settings. They give you two short files to inspect and audition; they are not a controlled comparison of matched voices or an OpenAI voice benchmark.<\/p>\n\n\n\n<aside class=\"gg-speech-method\" style=\"box-sizing:border-box;max-width:100%;font:16px\/1.6 system-ui,sans-serif;background:#fff9e9;border:1px solid #e3d29b;border-radius:8px;margin:25px 0;padding:22px;color:#443713!important;min-width:0;\"><h3 style=\"box-sizing:border-box;font:750 21px\/1.35 system-ui,sans-serif;margin:0 0 10px;color:#443713!important;line-height:1.3;\">How to use these samples<\/h3><p style=\"box-sizing:border-box;margin:8px 0;color:#443713!important;\"><strong style=\"box-sizing:border-box;color:#34290c!important;\">Tarefa:<\/strong> a three-step product tour, generated September 28, 2026. Same English text; one request per model; default voices; no rerolls.<\/p><p style=\"box-sizing:border-box;margin:8px 0;color:#443713!important;\"><strong style=\"box-sizing:border-box;color:#34290c!important;\">Preste aten\u00e7\u00e3o ao seguinte:<\/strong> complete wording, pauses between the steps, and whether the pace suits an onboarding screen. The output formats and file properties are recorded separately below.<\/p><p style=\"box-sizing:border-box;margin:8px 0;color:#443713!important;\"><strong style=\"box-sizing:border-box;color:#34290c!important;\">\u00c2mbito:<\/strong> model-output examples. The API listings above are integration recipes; these recordings do not demonstrate execution of those code listings.<\/p><\/aside>\n\n\n\n<h3 class=\"wp-block-heading\">Eleven v3: product-tour audio<\/h3>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/static.futureshareai.com\/anywhere-test\/task_BzifB3WGdrVGrrzM8WaC84AAuOLoGYcU.mp3\" preload=\"metadata\"><\/audio><figcaption class=\"wp-element-caption\">eleven-v3 \u2014 AI-generated voice, September 28, 2026.<\/figcaption><\/figure>\n\n\n\n<section class=\"gg-speech-a1\" style=\"box-sizing:border-box;font:16px\/1.55 system-ui,sans-serif;color:#17333d!important;border:1px solid #c6d8dc;border-radius:8px;background:#f5f9fb;padding:24px;margin:16px 0 30px;max-width:100%;min-width:0;\"><h3 style=\"box-sizing:border-box;color:#17333d!important;font:700 22px\/1.3 system-ui,sans-serif;margin:0 0 16px;line-height:1.3;\">eleven-v3 \u00b7 product-tour sample<\/h3>\n<dl style=\"box-sizing:border-box;display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,240px),1fr));gap:12px;margin:14px 0;\"><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Sa\u00edda<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">MP3 \u00b7 mono \u00b7 44.1 kHz<\/dd><\/div><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Comprimento<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">About 15.6 seconds<\/dd><\/div><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Gera\u00e7\u00e3o<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">First output \u00b7 1 request<\/dd><\/div><\/dl>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Configura\u00e7\u00f5es:<\/strong> the exact text below, default voice and other settings; no voice override or delivery instruction. The default voice identity was not recorded.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Status and timing:<\/strong> generation succeeded. The task's server timestamps show 13 seconds from submission to completion; this is one request, not a latency benchmark.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">File observation:<\/strong> The MP3 contains a mono stream at 44.1 kHz and is approximately 251 kB. Its frames can be read to the end of the file.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Cost reference:<\/strong> the public catalog rate is 330 credits per 1,000 characters. Use a task estimate for a new request; this rate is not an invoice for the sample.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Practical use:<\/strong> A compact MP3 for auditioning a short product tour. Check the full wording and playback on your target device before using it in a release.<\/p><details style=\"box-sizing:border-box;border-top:1px solid #c6d8dc;padding-top:12px;\"><summary style=\"box-sizing:border-box;color:#125e67!important;font-weight:700;cursor:pointer;\">Read the exact input \u00b7 216 characters<\/summary><blockquote style=\"box-sizing:border-box;margin:14px 0 0;padding:14px;border-left:3px solid #198a86;background:#fff;color:#17333d!important;\">Welcome to the Acorn Studio product tour. In three steps, you can turn a short script into an audio guide. First, write the message. Next, choose a voice. Finally, save the audio file and test playback on your phone.<\/blockquote><\/details><\/section>\n\n\n\n<h3 class=\"wp-block-heading\">Qwen Audio 3.0 TTS Flash: product-tour audio<\/h3>\n\n\n\n<figure class=\"wp-block-audio\"><audio controls src=\"https:\/\/static.futureshareai.com\/anywhere-test\/task_5slI3lgNp1AMOt9RFM8vDsGymH6D6hb5.wav\" preload=\"metadata\"><\/audio><figcaption class=\"wp-element-caption\">qwen-audio-3.0-tts-flash \u2014 AI-generated voice, September 28, 2026.<\/figcaption><\/figure>\n\n\n\n<section class=\"gg-speech-a2\" style=\"box-sizing:border-box;font:16px\/1.55 system-ui,sans-serif;color:#17333d!important;border:1px solid #c6d8dc;border-radius:8px;background:#f5f9fb;padding:24px;margin:16px 0 30px;max-width:100%;min-width:0;\"><h3 style=\"box-sizing:border-box;color:#17333d!important;font:700 22px\/1.3 system-ui,sans-serif;margin:0 0 16px;line-height:1.3;\">qwen-audio-3.0-tts-flash \u00b7 product-tour sample<\/h3>\n<dl style=\"box-sizing:border-box;display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,240px),1fr));gap:12px;margin:14px 0;\"><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Sa\u00edda<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">WAV \u00b7 mono \u00b7 24 kHz<\/dd><\/div><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Comprimento<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">15.52 seconds of PCM data<\/dd><\/div><div style=\"box-sizing:border-box;\"><dt style=\"box-sizing:border-box;color:#294752!important;font-size:13px;\">Gera\u00e7\u00e3o<\/dt><dd style=\"box-sizing:border-box;color:#294752!important;font-weight:700;margin:3px 0 0;\">First output \u00b7 1 request<\/dd><\/div><\/dl>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Configura\u00e7\u00f5es:<\/strong> the exact text below, default voice and other settings; no voice override or delivery instruction. The default voice identity was not recorded.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Status and timing:<\/strong> generation succeeded. The task's server timestamps show 17 seconds from submission to completion; this is one request, not a latency benchmark.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">File observation:<\/strong> The WAV contains mono, 16-bit PCM at 24 kHz and is approximately 745 kB. Its header declares more data than the file contains; the original output is preserved here.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Cost reference:<\/strong> the public catalog rate is 330 credits per 1,000 characters. Use a task estimate for a new request; this rate is not an invoice for the sample.<\/p>\n<p style=\"box-sizing:border-box;color:#294752!important;margin:12px 0;\"><strong style=\"box-sizing:border-box;\">Practical use:<\/strong> A WAV sample for evaluating a narration workflow. Resolve the header-length inconsistency in your production pipeline and check complete playback before distribution.<\/p><details style=\"box-sizing:border-box;border-top:1px solid #c6d8dc;padding-top:12px;\"><summary style=\"box-sizing:border-box;color:#125e67!important;font-weight:700;cursor:pointer;\">Read the exact input \u00b7 216 characters<\/summary><blockquote style=\"box-sizing:border-box;margin:14px 0 0;padding:14px;border-left:3px solid #198a86;background:#fff;color:#17333d!important;\">Welcome to the Acorn Studio product tour. In three steps, you can turn a short script into an audio guide. First, write the message. Next, choose a voice. Finally, save the audio file and test playback on your phone.<\/blockquote><\/details><\/section>\n\n\n\n<p class=\"wp-block-paragraph\">The Qwen sample is about three times the size of the Eleven sample at a similar duration, reflecting the different WAV and MP3 formats. Its WAV header also declares more data than the file physically contains. That is a file-packaging issue worth checking in your target player; it does not by itself establish anything about the quality of the spoken voice.<\/p>\n\n\n\n<h2 id=\"speech-api-errors\" class=\"wp-block-heading\">Fix common speech-generation problems<\/h2>\n\n\n\n<figure class=\"wp-block-table\" style=\"box-sizing:border-box;margin:28px 0;overflow-x:auto;font:16px\/1.65 system-ui,-apple-system,sans-serif;max-width:100%;min-width:0;color:#17323f;\"><table style=\"box-sizing:border-box;border-collapse:collapse;width:100%;min-width:530px;font-size:16px;line-height:1.6;\"><thead style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">Sintoma<\/th><th style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;background:#eaf4f3;color:#173e43;\">O que verificar primeiro<\/th><\/tr><\/thead><tbody style=\"box-sizing:border-box;\"><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">The saved MP3 will not play<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Check the HTTP status and response type. An API error may have been saved with an audio extension.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">The request rejects a voice or instruction<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Match the voice and optional fields to the selected model. Older OpenAI TTS models do not accept instructions.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">A long script is rejected<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Check the route\u2019s character limit and the Mini TTS token limit, then split at natural boundaries.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">A 401 or 429 response appears<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Check the key\/account for 401. For 429, distinguish rate limits from insufficient quota before retrying.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">A GlobalGPT request remains pending<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Poll the saved task ID and inspect status; a timeout is not a reason to submit a duplicate job.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">The voice does not match the requested settings<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Check model-specific metadata and ignored_parameters before assuming the setting was applied.<\/td><\/tr><tr style=\"box-sizing:border-box;\"><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">An audio duration or seek bar looks wrong<\/td><td style=\"box-sizing:border-box;padding:15px;text-align:left;border:1px solid #d3e1e4;vertical-align:top;color:#294752;\">Compare file metadata with the actual media payload and test the complete file in your target player.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">For an OpenAI error, start with its <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/error-codes\">error-code guidance<\/a>. For a GlobalGPT task, save the task ID and returned error details. During development, keep transport failures, billing failures, and an unsatisfactory voice take separate: each calls for a different fix.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before shipping narration, listen to the whole recording on the device your audience will use. Check the opening, names and numbers, transitions, and final sentence. Include an AI-voice disclosure near playback, and keep a copy of the exact text and model settings so an edited script can be regenerated deliberately.<\/p>\n\n\n\n<h2 id=\"speech-api-faq\" class=\"wp-block-heading\">Perguntas frequentes<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">What is the OpenAI Audio Speech API endpoint?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Use POST https:\/\/api.openai.com\/v1\/audio\/speech. The required fields are model, input, and voice. The response is generated audio; MP3 is the default format.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Is the OpenAI Speech API free?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The Speech API has usage-based pricing. Budget using the current API rates rather than assuming a free allowance: GPT-4o Mini TTS is billed by text input and audio output tokens, while TTS-1 and TTS-1 HD are billed by characters.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can I control the voice and speaking style?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Choose a voice supported by your model. GPT-4o Mini TTS also accepts instructions for delivery, such as a warm or calm tone. The instructions field does not work with TTS-1 or TTS-1 HD.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Does the Speech API support streaming?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. It can stream generated audio in chunks. Your application still needs to handle playback; streaming text-to-speech alone does not provide a complete live voice conversation.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can I use GlobalGPT by changing the OpenAI base URL?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Use GlobalGPT\u2019s documented media-task workflow for TTS: submit to POST \/tasks, poll GET \/tasks\/{id}, and download output.url after success. Its documented TTS catalog includes ElevenLabs and Qwen models with their own model IDs and parameters.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Which audio format should I start with?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Start with MP3 for a simple download or web player. WAV is useful when the next step is editing. Raw PCM needs a player configured for its sample rate and bit depth because it has no file header.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Start with one short script, one voice, and one saved file. Once playback and costs are predictable, add chunking, retries, and the rest of your media workflow. Keep the chosen API\u2019s request format, billing units, and output handling together in your implementation.<\/p>\n\n\n\n<section class=\"gg-speech-cta\" style=\"box-sizing:border-box;font:16px\/1.6 system-ui,sans-serif;background:#123d46;border:1px solid #2c6570;border-radius:8px;padding:26px;margin:28px 0;color:#fff!important;max-width:100%;min-width:0;\">\n<h3 style=\"box-sizing:border-box;font:700 24px\/1.3 system-ui,sans-serif;margin:0 0 10px;color:#fff!important;line-height:1.3;\">Give your next script a voice<\/h3>\n<p style=\"box-sizing:border-box;margin:0 0 20px;color:#edf7f5!important;\">Draft a short product tour, create a voiceover with ElevenLabs or Qwen, and continue with images and video in GlobalGPT's multi-model workspace.<\/p>\n<a class=\"ggsc-button\" href=\"https:\/\/www.glbgpt.com\/audio-generator?inviter=hub_audio&amp;login=1\" style=\"box-sizing:border-box;display:inline-block;background:#d6f26f;color:#16351f!important;font-weight:750;padding:11px 18px;border-radius:5px;text-decoration:none;max-width:100%;\">Create voice audio with GlobalGPT<\/a>\n<\/section>","protected":false},"excerpt":{"rendered":"<p>The OpenAI Audio Speech API turns text into spoken audio through POST https:\/\/api.openai.com\/v1\/audio\/speech. Send a model, your input text, and a voice; save the response as an MP3 or another supported audio format. For a new integration that needs instructions for tone and delivery, start with gpt-4o-mini-tts. A working request is only the first step. [&hellip;]<\/p>","protected":false},"author":7,"featured_media":20190,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_seopress_robots_primary_cat":"","_seopress_titles_title":"OpenAI Audio Speech API: Examples, Voices & Pricing","_seopress_titles_desc":"Use the OpenAI Audio Speech API with Python and curl. Compare voices, formats, and costs, then build a GlobalGPT TTS task workflow with audio examples.","_seopress_robots_index":"","footnotes":""},"categories":[7],"tags":[],"class_list":["post-20189","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-chat"],"acf":[],"_links":{"self":[{"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/posts\/20189","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/comments?post=20189"}],"version-history":[{"count":2,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/posts\/20189\/revisions"}],"predecessor-version":[{"id":20199,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/posts\/20189\/revisions\/20199"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/media\/20190"}],"wp:attachment":[{"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/media?parent=20189"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/categories?post=20189"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/pt-br\/wp-json\/wp\/v2\/tags?post=20189"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}