OpenAI Realtime API Review
Frontier speech-to-speech model API (gpt-realtime family) for low-latency voice agents over WebRTC, WebSocket, and SIP.
See pricing
If you sign up through links on this page, vsref may earn a commission; programs exist on both sides of most comparisons, and commissions never change verdicts. How we make money
What we know about OpenAI Realtime API
OpenAI Realtime API is a voice ai platform: frontier speech-to-speech model API (gpt-realtime family) for low-latency voice agents over WebRTC, WebSocket, and SIP. This profile tracks every OpenAI Realtime API fact we have verified, each linked to a primary source and dated.
OpenAI Realtime API does not publish a public per-minute rate we have been able to verify, so the price figure here stays blank until we can confirm one. We would rather show nothing than a guessed number; if you have a current OpenAI Realtime API quote, it is the fastest way for us to close that gap.
On capabilities, OpenAI Realtime API covers native sip trunking, interruption handling (barge-in), api-first (full lifecycle via api), and self-serve signup, and does not offer llm/tts costs passed through at cost?, annual contract required for best pricing?, and bring-your-own llm. Each of those is verified against OpenAI Realtime API's own docs or dashboard, not marketing copy.
In total we track 10 verified facts for OpenAI Realtime API today, each linking the primary source it came from so you can check our work - and vendor claims we have not measured ourselves are labeled as such on the OpenAI Realtime API fact sheet below.
Reviewed by vsref Editorialfacts verified Jul 30, 2026Methodology →
Fact sheet
Every row independently verifiedConsidering a switch? Best OpenAI Realtime API alternatives →
TTS in this stack
A voice agent's voice quality and per-minute cost come from its text-to-speech engine. See how dedicated engines compare on price, latency, and cloning rights: ElevenLabs, Cartesia, OpenAI TTS, or the full text-to-speech comparison →
STT in this stack
A voice agent hears through its speech-to-text engine, and transcription accuracy and streaming latency are priced per audio minute. Compare the engines builders pair with OpenAI Realtime API: Deepgram, AssemblyAI, GPT-4o Transcribe, or the full speech-to-text comparison →