OpenAI 오디오 API: 기능, 가격 및 사용 방법

오디오 파형, 전사본 및 API 연결을 묘사한 편집용 일러스트레이션
Understand the OpenAI Audio API's speech, transcription, and translation routes, with current model choices, pricing examples, and Python requests. Explore how GlobalGPT's own API adds audio tasks alongside chat, image, and video models.

AUDIO API GUIDE · UPDATED SEPTEMBER 29, 2026

Need to turn a recording into text, add a voiceover to a video, or let customers talk to an AI assistant? The OpenAI Audio API gives developers several ways to handle speech, but each job uses a different model, request, and billing unit.

Start with the task you want to finish. Below, you’ll find the main endpoints, current model choices, Python examples, and pricing calculations. If your app also needs other providers’ voices, chat models, images, or video, the GlobalGPT API offers another route: one API entry point with audio tasks alongside those services.

간단한 답변: Use OpenAI’s /v1/audio/transcriptions for recorded speech, /v1/audio/speech for generated speech, and /v1/audio/translations for translating recordings into English text. For an ongoing spoken conversation, choose a live voice session. GlobalGPT also provides an audio task API, with its own model catalog and request format.

What is the OpenAI Audio API?

The OpenAI Audio API is a set of developer interfaces for processing and generating speech. Your application sends text or audio to OpenAI, selects a model, and receives spoken audio or text in return. It can power a transcription app, narrated article, subtitle tool, or the speech stages of a larger assistant.

OpenAI의 Audio API reference separates speech generation, transcription, and translation. These are useful building blocks: transcription produces words from a recording; a separate text model can turn those words into a summary, and a speech model can read the summary aloud.

Three requests, three different outputs

Recording → transcript

Interview, meeting, lecture, or voice memo

/audio/transcriptions

Script → spoken audio

Narration, accessibility audio, or an app response

/audio/speech

Foreign-language recording → English text

An English translation of recorded speech

/audio/translations

OpenAI file-based audio tasks. A live two-way conversation requires a separate session-based voice integration.

ChatGPT is an application you can use directly; the API lets you add models to your own application. If you only want to speak with ChatGPT, start with ChatGPT Voice access and limits. An API project requires its own setup and usage budget.

Which audio endpoint and model should you use?

Choose around the output your application needs. A caption file, a transcript with speaker labels, and a live conversation are different deliverables, even when all three start with someone speaking.

여러분의 과제OpenAI routeModel or starting pointKey detail
Transcribe a completed recording/v1/audio/transcriptionsgpt-transcribeOpenAI’s recommended starting point for file transcription; preserves the spoken language.
Generate narration from text/v1/audio/speechgpt-4o-mini-ttsChoose a voice and give delivery instructions, such as a calm tone or slower pace.
Translate a recording into English/v1/audio/translationswhisper-1Returns English text. This model has a scheduled retirement; see the migration note below.
Get speaker-labeled segments/v1/audio/transcriptionsgpt-4o-transcribe-diarize요청 diarized_json. This is also a retiring model.
Display live captionsRealtime transcription sessiongpt-live-transcribeProcesses incoming audio without an assistant speaking back.
Build an ongoing voice conversationGPT-Live sessiongpt-live-1Can listen while speaking and delegate work to a backend agent.

현재 파일 전사 가이드 권장합니다 gpt-transcribe for ordinary recordings. For a podcast transcript, begin there. If you need word timestamps or subtitle formats, check the model-specific options before choosing a model; support varies.

Migration date to plan around: OpenAI 목록 whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, 및 gpt-4o-transcribe-diarize for API retirement on February 26, 2027. Existing translation, subtitle, and diarization integrations should follow the official deprecation guidance rather than assuming a new model accepts every old parameter.

Streaming speech is different from a live conversation

Streaming text-to-speech starts playing a response before the entire audio file finishes generating. It does not, by itself, let your app listen to a user’s interruption or manage an ongoing call. Similarly, streaming a transcript from a completed file does not establish a live microphone session.

For a new conversational voice app, OpenAI’s audio and voice guide points to GPT-Live. Its voice model handles the conversation while a backend agent does research or uses tools. The Realtime API remains a separate option when you need its session and tool model.

How to use the OpenAI Audio API

For the examples below, create an OpenAI API project, configure API billing, and keep your key in the OPENAI_API_KEY environment variable on your server or local development machine. Install the Python SDK with pip install openai. Do not put a permanent API key into a public web page or mobile app bundle.

Start with a short file or script so you can inspect the response before adding the request to your application.

1. Transcribe a recording

Save a recording as meeting.mp3, then run this script from the same folder. The returned 텍스트 is the transcript, not a meeting summary or action-item list.

from openai import OpenAI

client = OpenAI()

with open("meeting.mp3", "rb") as audio_file:
    transcript = client.audio.transcriptions.create(
        model="gpt-transcribe",
        file=audio_file,
    )

print(transcript.text)

For product names or industry vocabulary, provide relevant context instead of expecting the model to guess spellings. The current transcription guide supports 프롬프트, plus model-specific keyword and language hints. Test those options on a short sample from the recordings your app actually handles.

If your source is a recorded demo or interview video, our guide to transcribing audio from a video explains media preparation and the steps between getting a transcript and producing a useful summary.

2. Turn a script into spoken audio

Choose a supported voice and describe the delivery. This example writes an MP3 file to your machine; it does not play the file automatically.

from pathlib import Path
from openai import OpenAI

client = OpenAI()

with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="coral",
    input="Your meeting summary is ready. Here are the next steps.",
    instructions="Use a calm, clear tone and a measured pace.",
) as response:
    response.stream_to_file(Path("narration.mp3"))

그리고 text-to-speech guide documents controls for tone, speed, intonation, and other delivery choices. MP3 is the default output; other supported formats include WAV, PCM, Opus, AAC, and FLAC. Voice choices depend on the model. OpenAI also requires a clear disclosure that the voice is AI-generated.

For occasional narration, you may prefer a browser tool over writing code. Our guide to text-to-speech tools and workflows covers that route, including what to check before exporting a voiceover.

3. Translate recorded speech into English text

The translation endpoint uses whisper-1 and outputs English text. For an existing integration, the request looks like this. Account for the February 2027 retirement before building a new long-term dependency on it.

from openai import OpenAI

client = OpenAI()

with open("interview.mp3", "rb") as audio_file:
    translation = client.audio.translations.create(
        model="whisper-1",
        file=audio_file,
    )

print(translation.text)

This request does not produce dubbed English audio. Dubbing requires additional steps, including speech generation and, depending on the project, timing and video alignment. For translation into a different target language, use a translation approach that explicitly supports that language.

How much does the OpenAI Audio API cost?

There is no single Audio API price. File transcription, speech generation, and live sessions use different billing units. The figures below come from OpenAI의 API 요금, checked on September 29, 2026. They are OpenAI charges, separate from GlobalGPT prices.

Model or servicePublished billing informationBudget example
gpt-transcribeEstimated cost: 분당 $0.00451,000 minutes ≈ $4.50
gpt-live-transcribeEstimated cost: $0.017 per minute1,000 minutes ≈ $17.00
whisper-1 · retiring transcription/translation modelEstimated cost: $0.006 per minute60 minutes ≈ $0.36; plan migration before February 26, 2027
gpt-4o-mini-tts$0.60 / 1M text input tokens 플러스 $12 / 1M audio output tokens100,000 input tokens + 100,000 output audio tokens = $1.26
tts-1$15 / 1M characters100,000 characters = $1.50
tts-1-hd$30 / 1M characters100,000 characters = $3.00
gpt-live-1 voice sessions분당 $0.05, billed per second60 minutes = $3.00, plus backend model and tool usage

The transcription rows use OpenAI’s published estimated cost figures. The TTS token example is an arithmetic illustration based on two specified usage totals; it is not a prediction that a script will produce equal input and output token counts. Characters, text tokens, and audio tokens are different units.

A 1,000-minute transcription budget

About 16 hours and 40 minutes of audio, using OpenAI’s published per-minute estimates.

Recorded files · gpt-transcribe · ≈ $4.50

Live transcription · gpt-live-transcribe · ≈ $17.00

Different input modes, not an accuracy ranking. Calculations exclude additional text-model processing, storage, app hosting, and retries.

To estimate a real project, count the work after transcription too. A meeting assistant may need a transcript, a summary, action items, and spoken playback. Each model call can contribute to the bill. A voice agent may also use search or other tools while the conversation continues.

For GPT-Live specifically, the model pricing page separates session duration from backend model and tool charges. A one-hour call costing $3 for the voice session is not necessarily a $3 total application cost.

GlobalGPT also has an API for audio, chat, images, and video

If audio is one part of a broader product, GlobalGPT offers a useful option. Its API overview documents a common entry point for multiple model categories: chat requests use /v1/chat/complaints 또는 /v1/응답, while image, video, and audio generation use /v1/tasks.

For example, a content app might draft a script with a chat model, create a narration track, generate an illustration, and produce a short video. When building a multi-model workflow, give each model a clear job and pass reviewed results between steps.

질문OpenAI 직접 APIGlobalGPT API
Which audio models are relevant?OpenAI models, including gpt-transcribe 그리고 gpt-4o-mini-tts.Check the account’s API model catalog. The platform’s audio tools include Eleven, Qwen, Seed, and Scribe entries; availability and parameters must be checked per API model.
How are audio requests sent?Task-specific /v1/audio/… endpoints; live voice uses a session API.The API overview routes audio generation through /v1/tasks.
How do you check pricing?OpenAI’s model and API pricing pages.The GlobalGPT API overview says GET /v1/models lists model IDs with prices.
When is it a useful choice?You need a particular OpenAI audio model or its documented voice-session features.You want audio alongside multiple providers’ chat, image, and video models.

GlobalGPT’s public audio model catalog includes Eleven v3, Eleven Multilingual v2, Qwen Audio 3.0, and Seed Audio 1.0 for speech-related creation, plus Scribe and AI Note Taker for transcription tasks. Check the account’s API catalog for the specific models available to your integration.

For longer scripts, the Eleven Multilingual v2 review discusses narration, language support, and how it differs from other ElevenLabs models.

For creative audio that includes speech and ambience, Seed Audio’s voice and sound effects guide explores a broader task than reading text aloud. Choose around the kind of audio you need to deliver.

How to start with the GlobalGPT API

  1. Open the GlobalGPT API area and create a key under API → API Keys → Create key. Save it securely when it is displayed.
  2. Read the model catalog to find the model ID and price for the task you need.
  3. Use the selected model’s documented parameters. Follow the audio task request format for /v1/tasks, rather than pasting an OpenAI /audio/speech payload unchanged.
  4. Start with a short sample. Check the returned audio, completion status, and billed usage before processing a larger batch.

The documented API base URL is https://api2.glbgpt.com/ai-api/open/v1. This read-only catalog request is a starting point after you store your key in GLOBALGPT_API_KEY:

curl "https://api2.glbgpt.com/ai-api/open/v1/models" \
  -H "Authorization: Bearer $GLOBALGPT_API_KEY"

Use the price attached to the model you actually select. A GlobalGPT website subscription price, a credit balance, and an OpenAI per-minute rate are not interchangeable estimates for the same request.

Integration boundary: GlobalGPT provides its own audio API. Its public documentation does not establish support for OpenAI’s native audio endpoints, GPT-Live sessions, or every OpenAI speech model. If your existing app depends on one of those features, verify that specific requirement before switching providers.

Common Audio API mistakes and how to avoid them

문제확인할 사항실질적인 다음 단계
A recording will not uploadFile size and actual encoding.For OpenAI file transcription, keep each upload within 25 MB and use MP3, MP4, MPEG, MPGA, M4A, WAV, or WebM. Compress or split larger files at natural pauses.
The response has no speaker labelsWhether the chosen model and response format support diarization.For an existing diarization integration, request diarized_json; plan for the model’s announced retirement.
The translation is in English, but you wanted another languageThe endpoint’s target-language support.The file translation endpoint returns English text. Use a separate translation step or a service that supports your target language.
Audio plays, but the assistant cannot handle an interruptionWhether you built streaming playback or a conversational session.Use a live voice architecture when your app must listen while it speaks.
An OpenAI example fails against another providerBase URL, endpoint, model ID, and request fields.Follow the provider’s own audio documentation. Similar SDK setup does not guarantee identical audio routes.

Before choosing on voice quality, test a short sample that includes your real vocabulary: product names, numbers, abbreviations, and the languages your users speak. For transcription, inspect the words and speaker assignments. For narration, listen for pronunciation, pauses, pace, and whether the intended tone survives a longer paragraph.

Use recordings you have permission to process, keep keys on a trusted server, and review the data terms for the service receiving your audio. For a customer-facing product, make recording and AI-voice disclosures part of the interface.

어떤 API 경로를 선택해야 할까요?

Choose OpenAI directly when your app needs a specific OpenAI speech model, a documented transcription option, or a native live voice session. Start with the model suited to the task and design around its actual request and billing rules.

GlobalGPT를 고려해 보세요 when your product combines audio with models from several providers, or when you want to explore voice generation alongside chat, images, and video. Begin with the API catalog, select an available model, and validate a small task before scaling up.

Music, narration, and transcription call for different model choices. Our comparison on choosing an audio model by task helps separate song generation, spoken narration, and more complex audio scenes.

For a single voiceover or meeting transcript, a web tool may be enough. An API becomes useful when your own application must trigger the task repeatedly, pass results to another step, or serve the result to your users.

자주 묻는 질문

Is the OpenAI Audio API one model?

No. It is a set of audio interfaces. Speech generation, file transcription, and file translation use different endpoints and model choices. GPT-Live and Realtime add separate session-based options for live audio applications.

Can the OpenAI Audio API convert text to speech?

Yes. The speech endpoint accepts text, a model, and a voice, then returns audio. With gpt-4o-mini-tts, you can also provide instructions for delivery, such as tone and pace. MP3 is the default output format.

Which model should I use for a new transcription app?

OpenAI’s current file transcription guide recommends gpt-transcribe for general recorded speech. Check separately for specialized requirements such as word timestamps, subtitle formats, or speaker labels, because support differs by model.

Is Whisper being retired?

OpenAI lists its hosted whisper-1 API model for retirement on February 26, 2027, alongside three older GPT-4o transcription models. This is an API model retirement, not a statement that all Whisper software or all OpenAI audio services are shutting down.

Does audio translation return spoken audio?

No. The file translation endpoint returns English text from recorded speech. To produce translated speech, you need an additional speech-generation step or a dedicated speech-translation service.

Can I use the Audio API for a live voice assistant?

You can combine transcription, a text model, and speech generation into a voice assistant, but live turn handling needs additional design. OpenAI recommends GPT-Live as the starting point for new conversational voice applications; it also documents a separate Realtime API.

Does GlobalGPT have an audio API?

Yes. GlobalGPT’s API overview documents audio generation through its tasks endpoint, alongside image and video generation. Its models endpoint lists model IDs and prices. Check the chosen model’s availability and parameters before sending a request.

Is GlobalGPT’s audio API identical to OpenAI’s?

They document different audio routes: OpenAI provides dedicated audio endpoints, while GlobalGPT routes audio generation through its tasks API. Use the selected provider’s model IDs and request fields; OpenAI endpoint compatibility and live voice support must be verified separately.

Build audio into a broader AI app

Explore GlobalGPT’s API for audio, chat, image, and video tasks. Check the model catalog and pricing, then start with the task your users need most.

Explore the GlobalGPT API →

Open your GlobalGPT workspace

게시물을 공유하세요:

관련 게시물