AUDIO API GUIDE · UPDATED SEPTEMBER 29, 2026
Need to turn a recording into text, add a voiceover to a video, or let customers talk to an AI assistant? The OpenAI Audio API gives developers several ways to handle speech, but each job uses a different model, request, and billing unit.
Start with the task you want to finish. Below, you’ll find the main endpoints, current model choices, Python examples, and pricing calculations. If your app also needs other providers’ voices, chat models, images, or video, the GlobalGPT API offers another route: one API entry point with audio tasks alongside those services.
Snel antwoord: Use OpenAI’s /v1/audio/transcriptions for recorded speech, /v1/audio/speech for generated speech, and /v1/audio/translations for translating recordings into English text. For an ongoing spoken conversation, choose a live voice session. GlobalGPT also provides an audio task API, with its own model catalog and request format.
Op deze pagina
What is the OpenAI Audio API?
The OpenAI Audio API is a set of developer interfaces for processing and generating speech. Your application sends text or audio to OpenAI, selects a model, and receives spoken audio or text in return. It can power a transcription app, narrated article, subtitle tool, or the speech stages of a larger assistant.
OpenAI's Audio API reference separates speech generation, transcription, and translation. These are useful building blocks: transcription produces words from a recording; a separate text model can turn those words into a summary, and a speech model can read the summary aloud.
Three requests, three different outputs
Recording → transcript
Interview, meeting, lecture, or voice memo
/audio/transcriptions
Script → spoken audio
Narration, accessibility audio, or an app response
/audio/speech
Foreign-language recording → English text
An English translation of recorded speech
/audio/translations
ChatGPT is an application you can use directly; the API lets you add models to your own application. If you only want to speak with ChatGPT, start with ChatGPT Voice access and limits. An API project requires its own setup and usage budget.
Which audio endpoint and model should you use?
Choose around the output your application needs. A caption file, a transcript with speaker labels, and a live conversation are different deliverables, even when all three start with someone speaking.
De huidige handleiding voor het transcriberen van bestanden beveelt aan gpt-transcribe for ordinary recordings. For a podcast transcript, begin there. If you need word timestamps or subtitle formats, check the model-specific options before choosing a model; support varies.
Migration date to plan around: OpenAI-lijsten whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, en gpt-4o-transcribe-diarize for API retirement on February 26, 2027. Existing translation, subtitle, and diarization integrations should follow the official deprecation guidance rather than assuming a new model accepts every old parameter.
Streaming speech is different from a live conversation
Streaming text-to-speech starts playing a response before the entire audio file finishes generating. It does not, by itself, let your app listen to a user’s interruption or manage an ongoing call. Similarly, streaming a transcript from a completed file does not establish a live microphone session.
For a new conversational voice app, OpenAI’s audio and voice guide points to GPT-Live. Its voice model handles the conversation while a backend agent does research or uses tools. The Realtime API remains a separate option when you need its session and tool model.
How to use the OpenAI Audio API
For the examples below, create an OpenAI API project, configure API billing, and keep your key in the OPENAI_API_KEY environment variable on your server or local development machine. Install the Python SDK with pip install openai. Do not put a permanent API key into a public web page or mobile app bundle.
Start with a short file or script so you can inspect the response before adding the request to your application.
1. Transcribe a recording
Save a recording as meeting.mp3, then run this script from the same folder. The returned tekst is the transcript, not a meeting summary or action-item list.
from openai import OpenAI
client = OpenAI()
with open("meeting.mp3", "rb") as audio_file:
transcript = client.audio.transcriptions.create(
model="gpt-transcribe",
file=audio_file,
)
print(transcript.text)
For product names or industry vocabulary, provide relevant context instead of expecting the model to guess spellings. The current transcription guide supports prompt, plus model-specific keyword and language hints. Test those options on a short sample from the recordings your app actually handles.
If your source is a recorded demo or interview video, our guide to transcribing audio from a video explains media preparation and the steps between getting a transcript and producing a useful summary.
2. Turn a script into spoken audio
Choose a supported voice and describe the delivery. This example writes an MP3 file to your machine; it does not play the file automatically.
from pathlib import Path
from openai import OpenAI
client = OpenAI()
with client.audio.speech.with_streaming_response.create(
model="gpt-4o-mini-tts",
voice="coral",
input="Your meeting summary is ready. Here are the next steps.",
instructions="Use a calm, clear tone and a measured pace.",
) as response:
response.stream_to_file(Path("narration.mp3"))
De text-to-speech guide documents controls for tone, speed, intonation, and other delivery choices. MP3 is the default output; other supported formats include WAV, PCM, Opus, AAC, and FLAC. Voice choices depend on the model. OpenAI also requires a clear disclosure that the voice is AI-generated.
For occasional narration, you may prefer a browser tool over writing code. Our guide to text-to-speech tools and workflows covers that route, including what to check before exporting a voiceover.
3. Translate recorded speech into English text
The translation endpoint uses whisper-1 and outputs English text. For an existing integration, the request looks like this. Account for the February 2027 retirement before building a new long-term dependency on it.
from openai import OpenAI
client = OpenAI()
with open("interview.mp3", "rb") as audio_file:
translation = client.audio.translations.create(
model="whisper-1",
file=audio_file,
)
print(translation.text)
This request does not produce dubbed English audio. Dubbing requires additional steps, including speech generation and, depending on the project, timing and video alignment. For translation into a different target language, use a translation approach that explicitly supports that language.
How much does the OpenAI Audio API cost?
There is no single Audio API price. File transcription, speech generation, and live sessions use different billing units. The figures below come from De API-tarieven van OpenAI, checked on September 29, 2026. They are OpenAI charges, separate from GlobalGPT prices.
The transcription rows use OpenAI’s published estimated cost figures. The TTS token example is an arithmetic illustration based on two specified usage totals; it is not a prediction that a script will produce equal input and output token counts. Characters, text tokens, and audio tokens are different units.
A 1,000-minute transcription budget
About 16 hours and 40 minutes of audio, using OpenAI’s published per-minute estimates.
Recorded files · gpt-transcribe · ≈ $4.50
Live transcription · gpt-live-transcribe · ≈ $17.00
To estimate a real project, count the work after transcription too. A meeting assistant may need a transcript, a summary, action items, and spoken playback. Each model call can contribute to the bill. A voice agent may also use search or other tools while the conversation continues.
For GPT-Live specifically, the model pricing page separates session duration from backend model and tool charges. A one-hour call costing $3 for the voice session is not necessarily a $3 total application cost.
GlobalGPT also has an API for audio, chat, images, and video
If audio is one part of a broader product, GlobalGPT offers a useful option. Its API overview documents a common entry point for multiple model categories: chat requests use /v1/chat/voltooiingen of /v1/antwoorden, while image, video, and audio generation use /v1/taken.
For example, a content app might draft a script with a chat model, create a narration track, generate an illustration, and produce a short video. When building a multi-model workflow, give each model a clear job and pass reviewed results between steps.
GlobalGPT’s public audio model catalog includes Eleven v3, Eleven Multilingual v2, Qwen Audio 3.0, and Seed Audio 1.0 for speech-related creation, plus Scribe and AI Note Taker for transcription tasks. Check the account’s API catalog for the specific models available to your integration.
For longer scripts, the Eleven Multilingual v2 review discusses narration, language support, and how it differs from other ElevenLabs models.
For creative audio that includes speech and ambience, Seed Audio’s voice and sound effects guide explores a broader task than reading text aloud. Choose around the kind of audio you need to deliver.
How to start with the GlobalGPT API
- Open the GlobalGPT API area and create a key under API → API Keys → Create key. Save it securely when it is displayed.
- Read the model catalog to find the model ID and price for the task you need.
- Use the selected model’s documented parameters. Follow the audio task request format for
/v1/taken, rather than pasting an OpenAI/audio/speechpayload unchanged. - Start with a short sample. Check the returned audio, completion status, and billed usage before processing a larger batch.
The documented API base URL is https://api2.glbgpt.com/ai-api/open/v1. This read-only catalog request is a starting point after you store your key in GLOBALGPT_API_KEY:
curl "https://api2.glbgpt.com/ai-api/open/v1/models" \
-H "Authorization: Bearer $GLOBALGPT_API_KEY"
Use the price attached to the model you actually select. A GlobalGPT website subscription price, a credit balance, and an OpenAI per-minute rate are not interchangeable estimates for the same request.
Integration boundary: GlobalGPT provides its own audio API. Its public documentation does not establish support for OpenAI’s native audio endpoints, GPT-Live sessions, or every OpenAI speech model. If your existing app depends on one of those features, verify that specific requirement before switching providers.
Common Audio API mistakes and how to avoid them
Before choosing on voice quality, test a short sample that includes your real vocabulary: product names, numbers, abbreviations, and the languages your users speak. For transcription, inspect the words and speaker assignments. For narration, listen for pronunciation, pauses, pace, and whether the intended tone survives a longer paragraph.
Use recordings you have permission to process, keep keys on a trusted server, and review the data terms for the service receiving your audio. For a customer-facing product, make recording and AI-voice disclosures part of the interface.
Welke API-route moet je kiezen?
Choose OpenAI directly when your app needs a specific OpenAI speech model, a documented transcription option, or a native live voice session. Start with the model suited to the task and design around its actual request and billing rules.
Neem GlobalGPT eens in overweging when your product combines audio with models from several providers, or when you want to explore voice generation alongside chat, images, and video. Begin with the API catalog, select an available model, and validate a small task before scaling up.
Music, narration, and transcription call for different model choices. Our comparison on choosing an audio model by task helps separate song generation, spoken narration, and more complex audio scenes.
For a single voiceover or meeting transcript, a web tool may be enough. An API becomes useful when your own application must trigger the task repeatedly, pass results to another step, or serve the result to your users.
Veelgestelde vragen
Is the OpenAI Audio API one model?
No. It is a set of audio interfaces. Speech generation, file transcription, and file translation use different endpoints and model choices. GPT-Live and Realtime add separate session-based options for live audio applications.
Can the OpenAI Audio API convert text to speech?
Yes. The speech endpoint accepts text, a model, and a voice, then returns audio. With gpt-4o-mini-tts, you can also provide instructions for delivery, such as tone and pace. MP3 is the default output format.
Which model should I use for a new transcription app?
OpenAI’s current file transcription guide recommends gpt-transcribe for general recorded speech. Check separately for specialized requirements such as word timestamps, subtitle formats, or speaker labels, because support differs by model.
Is Whisper being retired?
OpenAI lists its hosted whisper-1 API model for retirement on February 26, 2027, alongside three older GPT-4o transcription models. This is an API model retirement, not a statement that all Whisper software or all OpenAI audio services are shutting down.
Does audio translation return spoken audio?
No. The file translation endpoint returns English text from recorded speech. To produce translated speech, you need an additional speech-generation step or a dedicated speech-translation service.
Can I use the Audio API for a live voice assistant?
You can combine transcription, a text model, and speech generation into a voice assistant, but live turn handling needs additional design. OpenAI recommends GPT-Live as the starting point for new conversational voice applications; it also documents a separate Realtime API.
Does GlobalGPT have an audio API?
Yes. GlobalGPT’s API overview documents audio generation through its tasks endpoint, alongside image and video generation. Its models endpoint lists model IDs and prices. Check the chosen model’s availability and parameters before sending a request.
Is GlobalGPT’s audio API identical to OpenAI’s?
They document different audio routes: OpenAI provides dedicated audio endpoints, while GlobalGPT routes audio generation through its tasks API. Use the selected provider’s model IDs and request fields; OpenAI endpoint compatibility and live voice support must be verified separately.
Build audio into a broader AI app
Explore GlobalGPT’s API for audio, chat, image, and video tasks. Check the model catalog and pricing, then start with the task your users need most.



