{"id":20169,"date":"2026-09-30T03:22:22","date_gmt":"2026-09-30T07:22:22","guid":{"rendered":"https:\/\/wp.glbgpt.com\/?p=20169"},"modified":"2026-09-30T03:22:23","modified_gmt":"2026-09-30T07:22:23","slug":"openai-audio-api","status":"publish","type":"post","link":"https:\/\/wp.glbgpt.com\/es\/hub\/openai-audio-api","title":{"rendered":"API de audio OpenAI: caracter\u00edsticas, precios y c\u00f3mo utilizarla"},"content":{"rendered":"<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">AUDIO API GUIDE \u00b7 UPDATED SEPTEMBER 29, 2026<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Need to turn a recording into text, add a voiceover to a video, or let customers talk to an AI assistant? The <strong>OpenAI Audio API<\/strong> gives developers several ways to handle speech, but each job uses a different model, request, and billing unit.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Start with the task you want to finish. Below, you\u2019ll find the main endpoints, current model choices, Python examples, and pricing calculations. If your app also needs other providers\u2019 voices, chat models, images, or video, the <a href=\"https:\/\/www.glbgpt.com\/home\/api\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">API GlobalGPT<\/a> offers another route: one API entry point with audio tasks alongside those services.<\/p>\n\n\n\n<div class=\"wp-block-group has-border-color has-text-color has-background is-layout-flow wp-block-group-is-layout-flow\" style=\"border-color:rgb(171, 216, 197);border-style:solid;border-width:1px;border-radius:10px;color:rgb(23, 32, 42);background-color:rgb(237, 248, 243);margin-top:24px;margin-right:0px;margin-bottom:24px;margin-left:0px;padding-top:22px;padding-right:24px;padding-bottom:22px;padding-left:24px\">\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>Respuesta r\u00e1pida:<\/strong> Use OpenAI\u2019s <code>\/v1\/audio\/transcriptions<\/code> for recorded speech, <code>\/v1\/audio\/speech<\/code> for generated speech, and <code>\/v1\/audio\/translations<\/code> for translating recordings into English text. For an ongoing spoken conversation, choose a live voice session. GlobalGPT also provides an audio task API, with its own model catalog and request format.<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-group has-border-color has-text-color has-background is-layout-flow wp-block-group-is-layout-flow\" style=\"border-color:rgb(216, 225, 237);border-style:solid;border-width:1px;border-radius:10px;color:rgb(23, 32, 42);background-color:rgb(243, 246, 250);margin-top:24px;margin-right:0px;margin-bottom:24px;margin-left:0px;padding-top:22px;padding-right:24px;padding-bottom:22px;padding-left:24px\">\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>En esta p\u00e1gina<\/strong><\/p>\n\n\n\n<ol style=\"margin-top:0px;margin-right:0px;margin-bottom:20px;margin-left:0px;padding-left:24px;line-height:1.75\" class=\"wp-block-list\">\n<li><a href=\"#what-is-openai-audio-api\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">What is the OpenAI Audio API?<\/a><\/li>\n\n\n\n<li><a href=\"#choose-an-audio-endpoint\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Choose an endpoint and model<\/a><\/li>\n\n\n\n<li><a href=\"#openai-audio-api-examples\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Get started with Python examples<\/a><\/li>\n\n\n\n<li><a href=\"#openai-audio-api-pricing\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Understand Audio API pricing<\/a><\/li>\n\n\n\n<li><a href=\"#globalgpt-audio-api\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Use the GlobalGPT API for audio and more<\/a><\/li>\n\n\n\n<li><a href=\"#audio-api-limits\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Avoid common setup mistakes<\/a><\/li>\n\n\n\n<li><a href=\"#choose-your-api\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Elige la ruta adecuada<\/a><\/li>\n\n\n\n<li><a href=\"#faq\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Preguntas frecuentes<\/a><\/li>\n<\/ol>\n<\/div>\n\n\n\n<h2 id=\"what-is-openai-audio-api\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">What is the OpenAI Audio API?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">The OpenAI Audio API is a set of developer interfaces for processing and generating speech. Your application sends text or audio to OpenAI, selects a model, and receives spoken audio or text in return. It can power a transcription app, narrated article, subtitle tool, or the speech stages of a larger assistant.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">OpenAI <a href=\"https:\/\/developers.openai.com\/api\/reference\/resources\/audio\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Audio API reference<\/a> separates speech generation, transcription, and translation. These are useful building blocks: transcription produces words from a recording; a separate text model can turn those words into a summary, and a speech model can read the summary aloud.<\/p>\n\n\n\n<figure data-module=\"audio-task-map\" style=\";background:#f0f7f5!important;color:#173b32!important;border:1px solid #c1dbd1!important;border-radius:12px;padding:24px!important;margin:28px 0!important;\"><h3 style=\";font-size:22px;line-height:1.35;color:#213f37;margin:28px 0 14px;font-size:21px!important;line-height:1.35!important;color:#173b32!important;margin:0 0 14px!important;\">Three requests, three different outputs<\/h3><div class=\"flow-grid\" style=\"display:flex;flex-wrap:wrap;gap:14px;align-items:stretch;display:flex!important;flex-wrap:wrap!important;gap:14px!important;margin-top:20px!important;\">\n<div style=\"flex:1 1 210px;border:1px solid #c7dbe7;border-radius:10px;background:#ffffff;padding:18px;flex:1 1 180px!important;min-width:0!important;background:#ffffff!important;border:1px solid #cbded7!important;border-radius:8px!important;padding:18px!important;\"><p style=\";margin:0 0 16px;line-height:1.75;\"><strong>Recording \u2192 transcript<\/strong><\/p><p style=\";margin:0 0 16px;line-height:1.75;\">Interview, meeting, lecture, or voice memo<\/p><p style=\";margin:0 0 16px;line-height:1.75;\"><code>\/audio\/transcriptions<\/code><\/p><\/div>\n<div style=\"flex:1 1 210px;border:1px solid #c7dbe7;border-radius:10px;background:#ffffff;padding:18px;flex:1 1 180px!important;min-width:0!important;background:#ffffff!important;border:1px solid #cbded7!important;border-radius:8px!important;padding:18px!important;\"><p style=\";margin:0 0 16px;line-height:1.75;\"><strong>Script \u2192 spoken audio<\/strong><\/p><p style=\";margin:0 0 16px;line-height:1.75;\">Narration, accessibility audio, or an app response<\/p><p style=\";margin:0 0 16px;line-height:1.75;\"><code>\/audio\/speech<\/code><\/p><\/div>\n<div style=\"flex:1 1 210px;border:1px solid #c7dbe7;border-radius:10px;background:#ffffff;padding:18px;flex:1 1 180px!important;min-width:0!important;background:#ffffff!important;border:1px solid #cbded7!important;border-radius:8px!important;padding:18px!important;\"><p style=\";margin:0 0 16px;line-height:1.75;\"><strong>Foreign-language recording \u2192 English text<\/strong><\/p><p style=\";margin:0 0 16px;line-height:1.75;\">An English translation of recorded speech<\/p><p style=\";margin:0 0 16px;line-height:1.75;\"><code>\/audio\/translations<\/code><\/p><\/div>\n<\/div><figcaption style=\";font-size:13px!important;line-height:1.6!important;color:#50665f!important;margin-top:16px!important;\">OpenAI file-based audio tasks. A live two-way conversation requires a separate session-based voice integration.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">ChatGPT is an application you can use directly; the API lets you add models to your own application. If you only want to speak with ChatGPT, start with <a href=\"https:\/\/www.glbgpt.com\/hub\/chatgpt-voice-rollout\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">ChatGPT Voice access and limits<\/a>. An API project requires its own setup and usage budget.<\/p>\n\n\n\n<h2 id=\"choose-an-audio-endpoint\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">Which audio endpoint and model should you use?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Choose around the output your application needs. A caption file, a transcript with speaker labels, and a live conversation are different deliverables, even when all three start with someone speaking.<\/p>\n\n\n\n<div class=\"table-wrap\" data-module=\"endpoint-table\" style=\"overflow-x:auto;max-width:100%;margin:24px 0 30px;\"><table style=\"width:100%;min-width:620px;border-collapse:separate!important;border-spacing:0!important;border:1px solid #dbe3ec!important;border-radius:9px;overflow:hidden;background:#ffffff!important;color:#17202a!important;font-size:15px!important;line-height:1.6!important;\"><thead><tr><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Tu tarea<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">OpenAI route<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Model or starting point<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Key detail<\/th><\/tr><\/thead><tbody>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Transcribe a completed recording<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>\/v1\/audio\/transcriptions<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-transcribe<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">OpenAI\u2019s recommended starting point for file transcription; preserves the spoken language.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Generate narration from text<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>\/v1\/audio\/speech<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-4o-mini-tts<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Choose a voice and give delivery instructions, such as a calm tone or slower pace.<\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Translate a recording into English<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>\/v1\/audio\/translations<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>susurro-1<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Returns English text. This model has a scheduled retirement; see the migration note below.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Get speaker-labeled segments<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>\/v1\/audio\/transcriptions<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-4o-transcribe-diarize<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Solicitud <code>diarized_json<\/code>. This is also a retiring model.<\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Display live captions<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Realtime transcription session<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-live-transcribe<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Processes incoming audio without an assistant speaking back.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Build an ongoing voice conversation<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">GPT-Live session<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-live-1<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Can listen while speaking and delegate work to a backend agent.<\/td><\/tr>\n<\/tbody><\/table><\/div>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">La corriente <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/speech-to-text\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Gu\u00eda de transcripci\u00f3n de archivos<\/a> recomienda <code>gpt-transcribe<\/code> for ordinary recordings. For a podcast transcript, begin there. If you need word timestamps or subtitle formats, check the model-specific options before choosing a model; support varies.<\/p>\n\n\n\n<div class=\"wp-block-group has-border-color has-text-color has-background is-layout-flow wp-block-group-is-layout-flow\" style=\"border-color:rgb(185, 207, 238);border-style:solid;border-width:1px;border-radius:10px;color:rgb(23, 32, 42);background-color:rgb(238, 245, 255);margin-top:24px;margin-right:0px;margin-bottom:24px;margin-left:0px;padding-top:22px;padding-right:24px;padding-bottom:22px;padding-left:24px\">\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>Migration date to plan around:<\/strong> Listas de OpenAI <code>susurro-1<\/code>, <code>gpt-4o-transcribe<\/code>, <code>gpt-4o-mini-transcribe<\/code>, y <code>gpt-4o-transcribe-diarize<\/code> for API retirement on <strong>February 26, 2027<\/strong>. Existing translation, subtitle, and diarization integrations should follow the <a href=\"https:\/\/developers.openai.com\/api\/docs\/deprecations#2026-08-26-transcription-models\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">official deprecation guidance<\/a> rather than assuming a new model accepts every old parameter.<\/p>\n<\/div>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Streaming speech is different from a live conversation<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Streaming text-to-speech starts playing a response before the entire audio file finishes generating. It does not, by itself, let your app listen to a user\u2019s interruption or manage an ongoing call. Similarly, streaming a transcript from a completed file does not establish a live microphone session.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For a new conversational voice app, OpenAI\u2019s <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/audio\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">audio and voice guide<\/a> points to GPT-Live. Its voice model handles the conversation while a backend agent does research or uses tools. The Realtime API remains a separate option when you need its session and tool model.<\/p>\n\n\n\n<h2 id=\"openai-audio-api-examples\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">How to use the OpenAI Audio API<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For the examples below, create an OpenAI API project, configure API billing, and keep your key in the <code>OPENAI_API_KEY<\/code> environment variable on your server or local development machine. Install the Python SDK with <code>pip install openai<\/code>. Do not put a permanent API key into a public web page or mobile app bundle.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Start with a short file or script so you can inspect the response before adding the request to your application.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">1. Transcribe a recording<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Save a recording as <code>meeting.mp3<\/code>, then run this script from the same folder. The returned <code>texto<\/code> is the transcript, not a meeting summary or action-item list.<\/p>\n\n\n\n<pre class=\"wp-block-code has-text-color has-background\" style=\"border-radius:10px;color:rgb(225, 237, 244);background-color:rgb(18, 33, 45);margin-top:22px;margin-right:0px;margin-bottom:22px;margin-left:0px;padding-top:20px;padding-right:20px;padding-bottom:20px;padding-left:20px;font-size:14px;line-height:1.65\"><code>from openai import OpenAI\n\nclient = OpenAI()\n\nwith open(\"meeting.mp3\", \"rb\") as audio_file:\n    transcript = client.audio.transcriptions.create(\n        model=\"gpt-transcribe\",\n        file=audio_file,\n    )\n\nprint(transcript.text)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For product names or industry vocabulary, provide relevant context instead of expecting the model to guess spellings. The current transcription guide supports <code>consulte<\/code>, plus model-specific keyword and language hints. Test those options on a short sample from the recordings your app actually handles.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">If your source is a recorded demo or interview video, our guide to <a href=\"https:\/\/www.glbgpt.com\/hub\/can-chatgpt-transcribe-videos-heres-what-you-need-to-know\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">transcribing audio from a video<\/a> explains media preparation and the steps between getting a transcript and producing a useful summary.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">2. Turn a script into spoken audio<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Choose a supported voice and describe the delivery. This example writes an MP3 file to your machine; it does not play the file automatically.<\/p>\n\n\n\n<pre class=\"wp-block-code has-text-color has-background\" style=\"border-radius:10px;color:rgb(225, 237, 244);background-color:rgb(18, 33, 45);margin-top:22px;margin-right:0px;margin-bottom:22px;margin-left:0px;padding-top:20px;padding-right:20px;padding-bottom:20px;padding-left:20px;font-size:14px;line-height:1.65\"><code>from pathlib import Path\nfrom openai import OpenAI\n\nclient = OpenAI()\n\nwith client.audio.speech.with_streaming_response.create(\n    model=\"gpt-4o-mini-tts\",\n    voice=\"coral\",\n    input=\"Your meeting summary is ready. Here are the next steps.\",\n    instructions=\"Use a calm, clear tone and a measured pace.\",\n) as response:\n    response.stream_to_file(Path(\"narration.mp3\"))<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">En <a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/text-to-speech\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">text-to-speech guide<\/a> documents controls for tone, speed, intonation, and other delivery choices. MP3 is the default output; other supported formats include WAV, PCM, Opus, AAC, and FLAC. Voice choices depend on the model. OpenAI also requires a clear disclosure that the voice is AI-generated.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For occasional narration, you may prefer a browser tool over writing code. Our guide to <a href=\"https:\/\/www.glbgpt.com\/hub\/text-to-speech\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">text-to-speech tools and workflows<\/a> covers that route, including what to check before exporting a voiceover.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">3. Translate recorded speech into English text<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">The translation endpoint uses <code>susurro-1<\/code> and outputs English text. For an existing integration, the request looks like this. Account for the February 2027 retirement before building a new long-term dependency on it.<\/p>\n\n\n\n<pre class=\"wp-block-code has-text-color has-background\" style=\"border-radius:10px;color:rgb(225, 237, 244);background-color:rgb(18, 33, 45);margin-top:22px;margin-right:0px;margin-bottom:22px;margin-left:0px;padding-top:20px;padding-right:20px;padding-bottom:20px;padding-left:20px;font-size:14px;line-height:1.65\"><code>from openai import OpenAI\n\nclient = OpenAI()\n\nwith open(\"interview.mp3\", \"rb\") as audio_file:\n    translation = client.audio.translations.create(\n        model=\"whisper-1\",\n        file=audio_file,\n    )\n\nprint(translation.text)<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">This request does not produce dubbed English audio. Dubbing requires additional steps, including speech generation and, depending on the project, timing and video alignment. For translation into a different target language, use a translation approach that explicitly supports that language.<\/p>\n\n\n\n<h2 id=\"openai-audio-api-pricing\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">How much does the OpenAI Audio API cost?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">There is no single Audio API price. File transcription, speech generation, and live sessions use different billing units. The figures below come from <a href=\"https:\/\/developers.openai.com\/api\/docs\/pricing\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Precios de la API de OpenAI<\/a>, checked on September 29, 2026. They are OpenAI charges, separate from GlobalGPT prices.<\/p>\n\n\n\n<div class=\"table-wrap\" data-module=\"pricing-table\" style=\"overflow-x:auto;max-width:100%;margin:24px 0 30px;\"><table style=\"width:100%;min-width:620px;border-collapse:separate!important;border-spacing:0!important;border:1px solid #dbe3ec!important;border-radius:9px;overflow:hidden;background:#ffffff!important;color:#17202a!important;font-size:15px!important;line-height:1.6!important;\"><thead><tr><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Model or service<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Published billing information<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Budget example<\/th><\/tr><\/thead><tbody>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-transcribe<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Estimated cost: <strong>$0,0045 por minuto<\/strong><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">1,000 minutes \u2248 <strong>$4.50<\/strong><\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-live-transcribe<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Estimated cost: <strong>$0.017 per minute<\/strong><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">1,000 minutes \u2248 <strong>$17.00<\/strong><\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>susurro-1<\/code> \u00b7 retiring transcription\/translation model<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Estimated cost: <strong>$0.006 per minute<\/strong><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">60 minutes \u2248 <strong>$0.36<\/strong>; plan migration before February 26, 2027<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-4o-mini-tts<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><strong>$0,60 \/ 1 mill\u00f3n de tokens de entrada de texto<\/strong> m\u00e1s <strong>$12 \/ 1M fichas de salida de audio<\/strong><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">100,000 input tokens + 100,000 output audio tokens = <strong>$1.26<\/strong><\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>tts-1<\/code><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><strong>$15 \/ 1M caracteres<\/strong><\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">100,000 characters = <strong>$1.50<\/strong><\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>tts-1-hd<\/code><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><strong>$30 \/ 1M caracteres<\/strong><\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">100,000 characters = <strong>$3.00<\/strong><\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><code>gpt-live-1<\/code> voice sessions<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\"><strong>$0,05 por minuto<\/strong>, billed per second<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">60 minutes = <strong>$3.00<\/strong>, plus backend model and tool usage<\/td><\/tr>\n<\/tbody><\/table><\/div>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">The transcription rows use OpenAI\u2019s published <em>estimated cost<\/em> figures. The TTS token example is an arithmetic illustration based on two specified usage totals; it is not a prediction that a script will produce equal input and output token counts. Characters, text tokens, and audio tokens are different units.<\/p>\n\n\n\n<figure data-module=\"transcription-budget\" style=\";background:#f0f7f5!important;color:#173b32!important;border:1px solid #c1dbd1!important;border-radius:12px;padding:24px!important;margin:28px 0!important;\"><h3 style=\";font-size:22px;line-height:1.35;color:#213f37;margin:28px 0 14px;font-size:21px!important;line-height:1.35!important;color:#173b32!important;margin:0 0 14px!important;\">A 1,000-minute transcription budget<\/h3><p style=\";margin:0 0 16px;line-height:1.75;\">About 16 hours and 40 minutes of audio, using OpenAI\u2019s published per-minute estimates.<\/p>\n<div class=\"price-row\" style=\";margin:18px 0!important;\"><p class=\"price-label\" style=\";margin:0 0 16px;line-height:1.75;display:flex!important;flex-wrap:wrap!important;justify-content:space-between!important;gap:10px!important;font-size:15px!important;margin-bottom:7px!important;\"><strong>Recorded files \u00b7 gpt-transcribe \u00b7 \u2248 $4.50<\/strong><\/p><div class=\"price-track\" style=\";height:14px!important;background:#dce9e4!important;border-radius:20px!important;overflow:hidden!important;\"><span style=\"width:26.47%!important;display:block!important;height:100%!important;background:#197d66!important;border-radius:20px!important;\"><\/span><\/div><\/div>\n<div class=\"price-row\" style=\";margin:18px 0!important;\"><p class=\"price-label\" style=\";margin:0 0 16px;line-height:1.75;display:flex!important;flex-wrap:wrap!important;justify-content:space-between!important;gap:10px!important;font-size:15px!important;margin-bottom:7px!important;\"><strong>Live transcription \u00b7 gpt-live-transcribe \u00b7 \u2248 $17.00<\/strong><\/p><div class=\"price-track\" style=\";height:14px!important;background:#dce9e4!important;border-radius:20px!important;overflow:hidden!important;\"><span style=\"width:100%!important;display:block!important;height:100%!important;background:#197d66!important;border-radius:20px!important;\"><\/span><\/div><\/div>\n<figcaption style=\";font-size:13px!important;line-height:1.6!important;color:#50665f!important;margin-top:16px!important;\">Different input modes, not an accuracy ranking. Calculations exclude additional text-model processing, storage, app hosting, and retries.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">To estimate a real project, count the work after transcription too. A meeting assistant may need a transcript, a summary, action items, and spoken playback. Each model call can contribute to the bill. A voice agent may also use search or other tools while the conversation continues.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For GPT-Live specifically, the <a href=\"https:\/\/developers.openai.com\/api\/docs\/models\/gpt-live-1\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">model pricing page<\/a> separates session duration from backend model and tool charges. A one-hour call costing $3 for the voice session is not necessarily a $3 total application cost.<\/p>\n\n\n\n<h2 id=\"globalgpt-audio-api\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">GlobalGPT also has an API for audio, chat, images, and video<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">If audio is one part of a broader product, GlobalGPT offers a useful option. Its <a href=\"https:\/\/www.glbgpt.com\/home\/api\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">API overview<\/a> documents a common entry point for multiple model categories: chat requests use <code>\/v1\/chat\/completions<\/code> o <code>\/v1\/respuestas<\/code>, while image, video, and audio generation use <code>\/v1\/tareas<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For example, a content app might draft a script with a chat model, create a narration track, generate an illustration, and produce a short video. When <a href=\"https:\/\/www.glbgpt.com\/hub\/all-in-one-ai-models\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">building a multi-model workflow<\/a>, give each model a clear job and pass reviewed results between steps.<\/p>\n\n\n\n<div class=\"table-wrap\" data-module=\"provider-scope\" style=\"overflow-x:auto;max-width:100%;margin:24px 0 30px;\"><table style=\"width:100%;min-width:620px;border-collapse:separate!important;border-spacing:0!important;border:1px solid #dbe3ec!important;border-radius:9px;overflow:hidden;background:#ffffff!important;color:#17202a!important;font-size:15px!important;line-height:1.6!important;\"><thead><tr><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Pregunta<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">API directa OpenAI<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">API GlobalGPT<\/th><\/tr><\/thead><tbody>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Which audio models are relevant?<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">OpenAI models, including <code>gpt-transcribe<\/code> y <code>gpt-4o-mini-tts<\/code>.<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Check the account\u2019s API model catalog. The platform\u2019s audio tools include Eleven, Qwen, Seed, and Scribe entries; availability and parameters must be checked per API model.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">How are audio requests sent?<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Task-specific <code>\/v1\/audio\/\u2026<\/code> endpoints; live voice uses a session API.<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The API overview routes audio generation through <code>\/v1\/tareas<\/code>.<\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">How do you check pricing?<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">OpenAI\u2019s model and API pricing pages.<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The GlobalGPT API overview says <code>GET \/v1\/models<\/code> lists model IDs with prices.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">When is it a useful choice?<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">You need a particular OpenAI audio model or its documented voice-session features.<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">You want audio alongside multiple providers\u2019 chat, image, and video models.<\/td><\/tr>\n<\/tbody><\/table><\/div>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">GlobalGPT\u2019s public <a href=\"https:\/\/www.glbgpt.com\/en\/models\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">audio model catalog<\/a> includes Eleven v3, Eleven Multilingual v2, Qwen Audio 3.0, and Seed Audio 1.0 for speech-related creation, plus Scribe and AI Note Taker for transcription tasks. Check the account\u2019s API catalog for the specific models available to your integration.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For longer scripts, the <a href=\"https:\/\/www.glbgpt.com\/hub\/elevenlabs-multilingual-v2-review\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Eleven Multilingual v2 review<\/a> discusses narration, language support, and how it differs from other ElevenLabs models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For creative audio that includes speech and ambience, <a href=\"https:\/\/www.glbgpt.com\/hub\/seed-audio-1-0-review\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">Seed Audio\u2019s voice and sound effects guide<\/a> explores a broader task than reading text aloud. Choose around the kind of audio you need to deliver.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">How to start with the GlobalGPT API<\/h3>\n\n\n\n<ol style=\"margin-top:0px;margin-right:0px;margin-bottom:20px;margin-left:0px;padding-left:24px;line-height:1.75\" class=\"wp-block-list\">\n<li>Open the GlobalGPT API area and create a key under <strong>API \u2192 API Keys \u2192 Create key<\/strong>. Save it securely when it is displayed.<\/li>\n\n\n\n<li>Read the model catalog to find the model ID and price for the task you need.<\/li>\n\n\n\n<li>Use the selected model\u2019s documented parameters. Follow the audio task request format for <code>\/v1\/tareas<\/code>, rather than pasting an OpenAI <code>\/audio\/speech<\/code> payload unchanged.<\/li>\n\n\n\n<li>Start with a short sample. Check the returned audio, completion status, and billed usage before processing a larger batch.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">The documented API base URL is <code>https:\/\/api2.glbgpt.com\/ai-api\/open\/v1<\/code>. This read-only catalog request is a starting point after you store your key in <code>GLOBALGPT_API_KEY<\/code>:<\/p>\n\n\n\n<pre class=\"wp-block-code has-text-color has-background\" style=\"border-radius:10px;color:rgb(225, 237, 244);background-color:rgb(18, 33, 45);margin-top:22px;margin-right:0px;margin-bottom:22px;margin-left:0px;padding-top:20px;padding-right:20px;padding-bottom:20px;padding-left:20px;font-size:14px;line-height:1.65\"><code>curl \"https:\/\/api2.glbgpt.com\/ai-api\/open\/v1\/models\" \\\n  -H \"Authorization: Bearer $GLOBALGPT_API_KEY\"<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Use the price attached to the model you actually select. A GlobalGPT website subscription price, a credit balance, and an OpenAI per-minute rate are not interchangeable estimates for the same request.<\/p>\n\n\n\n<div class=\"wp-block-group has-border-color has-text-color has-background is-layout-flow wp-block-group-is-layout-flow\" style=\"border-color:rgb(185, 207, 238);border-style:solid;border-width:1px;border-radius:10px;color:rgb(23, 32, 42);background-color:rgb(238, 245, 255);margin-top:24px;margin-right:0px;margin-bottom:24px;margin-left:0px;padding-top:22px;padding-right:24px;padding-bottom:22px;padding-left:24px\">\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>Integration boundary:<\/strong> GlobalGPT provides its own audio API. Its public documentation does not establish support for OpenAI\u2019s native audio endpoints, GPT-Live sessions, or every OpenAI speech model. If your existing app depends on one of those features, verify that specific requirement before switching providers.<\/p>\n<\/div>\n\n\n\n<h2 id=\"audio-api-limits\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">Common Audio API mistakes and how to avoid them<\/h2>\n\n\n\n<div class=\"table-wrap\" data-module=\"troubleshooting-table\" style=\"overflow-x:auto;max-width:100%;margin:24px 0 30px;\"><table style=\"width:100%;min-width:620px;border-collapse:separate!important;border-spacing:0!important;border:1px solid #dbe3ec!important;border-radius:9px;overflow:hidden;background:#ffffff!important;color:#17202a!important;font-size:15px!important;line-height:1.6!important;\"><thead><tr><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Problema<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Qu\u00e9 hay que comprobar<\/th><th scope=\"col\" style=\"background:#17364d!important;color:#ffffff!important;font-weight:700!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Pr\u00f3ximo paso pr\u00e1ctico<\/th><\/tr><\/thead><tbody>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">A recording will not upload<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">File size and actual encoding.<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">For OpenAI file transcription, keep each upload within 25 MB and use MP3, MP4, MPEG, MPGA, M4A, WAV, or WebM. Compress or split larger files at natural pauses.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The response has no speaker labels<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Whether the chosen model and response format support diarization.<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">For an existing diarization integration, request <code>diarized_json<\/code>; plan for the model\u2019s announced retirement.<\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The translation is in English, but you wanted another language<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The endpoint\u2019s target-language support.<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">The file translation endpoint returns English text. Use a separate translation step or a service that supports your target language.<\/td><\/tr>\n<tr><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Audio plays, but the assistant cannot handle an interruption<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Whether you built streaming playback or a conversational session.<\/td><td style=\"background:#f4f7fb!important;color:#17202a!important;border:0!important;border-bottom:1px solid #dbe3ec!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Use a live voice architecture when your app must listen while it speaks.<\/td><\/tr>\n<tr><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">An OpenAI example fails against another provider<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Base URL, endpoint, model ID, and request fields.<\/td><td style=\"background:#ffffff!important;color:#17202a!important;border:0!important;border-bottom:0!important;padding:14px!important;text-align:left!important;vertical-align:top!important;\">Follow the provider\u2019s own audio documentation. Similar SDK setup does not guarantee identical audio routes.<\/td><\/tr>\n<\/tbody><\/table><\/div>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Before choosing on voice quality, test a short sample that includes your real vocabulary: product names, numbers, abbreviations, and the languages your users speak. For transcription, inspect the words and speaker assignments. For narration, listen for pronunciation, pauses, pace, and whether the intended tone survives a longer paragraph.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Use recordings you have permission to process, keep keys on a trusted server, and review the data terms for the service receiving your audio. For a customer-facing product, make recording and AI-voice disclosures part of the interface.<\/p>\n\n\n\n<h2 id=\"choose-your-api\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">\u00bfQu\u00e9 ruta de la API deber\u00edas elegir?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>Choose OpenAI directly<\/strong> when your app needs a specific OpenAI speech model, a documented transcription option, or a native live voice session. Start with the model suited to the task and design around its actual request and billing rules.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\"><strong>Toma como ejemplo GlobalGPT<\/strong> when your product combines audio with models from several providers, or when you want to explore voice generation alongside chat, images, and video. Begin with the API catalog, select an available model, and validate a small task before scaling up.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Music, narration, and transcription call for different model choices. Our comparison on <a href=\"https:\/\/www.glbgpt.com\/hub\/eleven-music-v2-vs-qwen-audio-3-0-vs-seed-audio-1-0\/\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;\">choosing an audio model by task<\/a> helps separate song generation, spoken narration, and more complex audio scenes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">For a single voiceover or meeting transcript, a web tool may be enough. An API becomes useful when your own application must trigger the task repeatedly, pass results to another step, or serve the result to your users.<\/p>\n\n\n\n<h2 id=\"faq\" class=\"wp-block-heading has-text-color\" style=\"color:rgb(23, 46, 64);margin-top:46px;margin-right:0px;margin-bottom:18px;margin-left:0px;font-size:28px;line-height:1.3\">Preguntas frecuentes<\/h2>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Is the OpenAI Audio API one model?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">No. It is a set of audio interfaces. Speech generation, file transcription, and file translation use different endpoints and model choices. GPT-Live and Realtime add separate session-based options for live audio applications.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Can the OpenAI Audio API convert text to speech?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Yes. The speech endpoint accepts text, a model, and a voice, then returns audio. With gpt-4o-mini-tts, you can also provide instructions for delivery, such as tone and pace. MP3 is the default output format.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Which model should I use for a new transcription app?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">OpenAI\u2019s current file transcription guide recommends gpt-transcribe for general recorded speech. Check separately for specialized requirements such as word timestamps, subtitle formats, or speaker labels, because support differs by model.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Is Whisper being retired?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">OpenAI lists its hosted whisper-1 API model for retirement on February 26, 2027, alongside three older GPT-4o transcription models. This is an API model retirement, not a statement that all Whisper software or all OpenAI audio services are shutting down.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Does audio translation return spoken audio?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">No. The file translation endpoint returns English text from recorded speech. To produce translated speech, you need an additional speech-generation step or a dedicated speech-translation service.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Can I use the Audio API for a live voice assistant?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">You can combine transcription, a text model, and speech generation into a voice assistant, but live turn handling needs additional design. OpenAI recommends GPT-Live as the starting point for new conversational voice applications; it also documents a separate Realtime API.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Does GlobalGPT have an audio API?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">Yes. GlobalGPT\u2019s API overview documents audio generation through its tasks endpoint, alongside image and video generation. Its models endpoint lists model IDs and prices. Check the chosen model\u2019s availability and parameters before sending a request.<\/p>\n\n\n\n<h3 class=\"wp-block-heading has-text-color\" style=\"color:rgb(33, 63, 55);margin-top:28px;margin-right:0px;margin-bottom:14px;margin-left:0px;font-size:22px;line-height:1.35\">Is GlobalGPT\u2019s audio API identical to OpenAI\u2019s?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\" style=\"margin-top:0px;margin-right:0px;margin-bottom:16px;margin-left:0px;line-height:1.75\">They document different audio routes: OpenAI provides dedicated audio endpoints, while GlobalGPT routes audio generation through its tasks API. Use the selected provider\u2019s model IDs and request fields; OpenAI endpoint compatibility and live voice support must be verified separately.<\/p>\n\n\n\n<div class=\"cta\" data-module=\"globalgpt-cta\" style=\"background:#183b31!important;color:#ffffff!important;padding:28px!important;border-radius:12px!important;margin:30px 0!important;\"><h3 style=\"color:#ffffff!important;font-size:24px!important;line-height:1.35!important;margin:0 0 14px!important;\">Build audio into a broader AI app<\/h3><p style=\"color:#ffffff!important;line-height:1.75!important;margin:0 0 16px!important;\">Explore GlobalGPT\u2019s API for audio, chat, image, and video tasks. Check the model catalog and pricing, then start with the task your users need most.<\/p><p style=\"color:#ffffff!important;line-height:1.75!important;margin:0 0 16px!important;\"><a class=\"cta-button\" href=\"https:\/\/www.glbgpt.com\/home\/api\" style=\"display:inline-block!important;white-space:normal!important;background:#a1ebcc!important;color:#133b2d!important;padding:12px 18px!important;border-radius:7px!important;font-weight:700!important;text-decoration:none!important;\">Explore the GlobalGPT API \u2192<\/a><\/p><p style=\"color:#ffffff!important;line-height:1.75!important;margin:0 0 16px!important;\"><a href=\"https:\/\/www.glbgpt.com\/home?inviter=hub_popup&amp;login=1\" style=\";color:#08755e;text-decoration:underline;text-underline-offset:3px;color:#b4efd6!important;\">Open your GlobalGPT workspace<\/a><\/p><\/div>\n\n\n\n<script type=\"application\/ld+json\">{\n    \"@context\": \"https:\\\/\\\/schema.org\",\n    \"@type\": \"FAQPage\",\n    \"mainEntity\": [\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Is the OpenAI Audio API one model?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"No. It is a set of audio interfaces. Speech generation, file transcription, and file translation use different endpoints and model choices. GPT-Live and Realtime add separate session-based options for live audio applications.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can the OpenAI Audio API convert text to speech?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. The speech endpoint accepts text, a model, and a voice, then returns audio. With gpt-4o-mini-tts, you can also provide instructions for delivery, such as tone and pace. MP3 is the default output format.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Which model should I use for a new transcription app?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"OpenAI\\u2019s current file transcription guide recommends gpt-transcribe for general recorded speech. Check separately for specialized requirements such as word timestamps, subtitle formats, or speaker labels, because support differs by model.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Is Whisper being retired?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"OpenAI lists its hosted whisper-1 API model for retirement on February 26, 2027, alongside three older GPT-4o transcription models. This is an API model retirement, not a statement that all Whisper software or all OpenAI audio services are shutting down.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Does audio translation return spoken audio?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"No. The file translation endpoint returns English text from recorded speech. To produce translated speech, you need an additional speech-generation step or a dedicated speech-translation service.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Can I use the Audio API for a live voice assistant?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"You can combine transcription, a text model, and speech generation into a voice assistant, but live turn handling needs additional design. OpenAI recommends GPT-Live as the starting point for new conversational voice applications; it also documents a separate Realtime API.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Does GlobalGPT have an audio API?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"Yes. GlobalGPT\\u2019s API overview documents audio generation through its tasks endpoint, alongside image and video generation. Its models endpoint lists model IDs and prices. Check the chosen model\\u2019s availability and parameters before sending a request.\"\n            }\n        },\n        {\n            \"@type\": \"Question\",\n            \"name\": \"Is GlobalGPT\\u2019s audio API identical to OpenAI\\u2019s?\",\n            \"acceptedAnswer\": {\n                \"@type\": \"Answer\",\n                \"text\": \"They document different audio routes: OpenAI provides dedicated audio endpoints, while GlobalGPT routes audio generation through its tasks API. Use the selected provider\\u2019s model IDs and request fields; OpenAI endpoint compatibility and live voice support must be verified separately.\"\n            }\n        }\n    ]\n}<\/script>","protected":false},"excerpt":{"rendered":"<p>Understand the OpenAI Audio API&#8217;s speech, transcription, and translation routes, with current model choices, pricing examples, and Python requests. Explore how GlobalGPT&#8217;s own API adds audio tasks alongside chat, image, and video models.<\/p>","protected":false},"author":13,"featured_media":20180,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_seopress_robots_primary_cat":"","_seopress_titles_title":"OpenAI Audio API: Pricing, Examples & GlobalGPT Guide","_seopress_titles_desc":"Need transcription or AI voiceovers? Explore OpenAI Audio API models, prices, and code examples, plus GlobalGPT's API for audio and other AI tasks.","_seopress_robots_index":"","footnotes":""},"categories":[109],"tags":[],"class_list":["post-20169","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-research"],"acf":[],"_links":{"self":[{"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/posts\/20169","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/users\/13"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/comments?post=20169"}],"version-history":[{"count":2,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/posts\/20169\/revisions"}],"predecessor-version":[{"id":20179,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/posts\/20169\/revisions\/20179"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/media\/20180"}],"wp:attachment":[{"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/media?parent=20169"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/categories?post=20169"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/es\/wp-json\/wp\/v2\/tags?post=20169"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}