Voice has long been the most natural way humans communicate — but building software that truly understands, reasons over, and acts on spoken language in real time has remained an unsolved engineering challenge. OpenAI's latest announcement changes that. The company has released three new real-time audio models via its API — GPT‑Realtime‑2, GPT‑Realtime‑Translate, and GPT‑Realtime‑Whisper — representing what developers and researchers are calling the most significant real-time AI voice release the company has shipped. This is the generation leap that moves voice from simple call-and-response toward genuine intelligence.
GPT‑Realtime‑2: The First Voice Model with GPT-5-Class Reasoning
At the centre of the release is GPT‑Realtime‑2 — OpenAI's most advanced voice reasoning model to date and the first to bring GPT-5-class reasoning to live spoken conversations. Unlike traditional voice assistants that operate through separate sequential stages — speech recognition, language understanding, and speech synthesis — this model processes speech in a continuous stream, interpreting and responding without noticeable delay.
The model operates with a 128,000-token context window — a fourfold increase over its predecessor's 32,000-token limit — and supports adjustable reasoning effort levels for different use cases. On the Big Bench Audio Intelligence benchmark, it scored 96.6% at the highest reasoning effort setting, representing a 15.2% improvement over GPT‑Realtime‑1.5.
Critically, GPT‑Realtime‑2 can use tools and trigger actions during an ongoing conversation — retrieving data, performing operations, and executing workflows in connected systems without pausing or breaking conversational flow. Developers can also enable parallel tool calls, short preambles to signal processing, and improved error recovery — making the model genuinely production-ready for complex enterprise applications.
GPT‑Realtime‑Translate: Breaking Language Barriers in Real Time
The second model, GPT‑Realtime‑Translate, is built entirely for live speech translation — processing spoken input continuously and generating translations in real time without requiring speakers to pause or complete full sentences. It supports over 70 input languages and approximately 13 output languages, maintaining the pace of natural conversation throughout.
In OpenAI's own evaluations across Hindi, Tamil, and Telugu, GPT‑Realtime‑Translate delivered 12.5% lower Word Error Rates than any other model tested, along with lower fallback rates, higher task completion, and latency that sustained natural conversational rhythm. Priced at approximately $0.034 per minute of audio processing — a usage-based model that makes it commercially accessible for a broad range of applications — the model is already being explored by Deutsche Telekom for more natural cross-language customer interactions at scale.
"Together, the models we are launching move realtime audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds."— OpenAI, Official Announcement
GPT‑Realtime‑Whisper: Streaming Speech-to-Text at Production Speed
The third model, GPT‑Realtime‑Whisper, evolves OpenAI's widely adopted Whisper speech recognition technology into a fully real-time streaming transcription system. Where the original Whisper was designed for post-recording analysis, this new version transcribes speech continuously as it is spoken — enabling live products to feel faster, more responsive, and more natural.
Priced at $0.017 per minute, GPT‑Realtime‑Whisper is optimised for a broad range of production use cases: meeting captions that appear in the moment, notes and summaries generated before a conversation ends, voice agents that need to understand users continuously, and faster follow-up workflows across customer support, healthcare, sales, and recruiting. The model makes live speech directly usable inside business workflows as it happens — not after the fact.
Real-World Applications Across Industries
The breadth of production applications enabled by these three models spans virtually every sector where human communication happens at scale:
- →Healthcare — Live medical documentation where a clinician dictates notes during a patient encounter and structured records are generated as the conversation happens, without any post-processing step
- →Property & Commerce — Zillow's voice-powered home search assistant handles spoken filters, pulls live listings, and books tours without a single tap
- →Broadcasting & Accessibility — Real-time captions for live events, broadcasts, classrooms, and meetings using GPT‑Realtime‑Whisper's streaming transcription
- →Education — Adaptive tutors that listen to a student's spoken answer, reason about the quality of that answer, ask clarifying follow-ups, and deliver feedback — all within a single continuous session
- →Telecoms & Customer Support — Deutsche Telekom is exploring GPT‑Realtime‑Translate to deliver more natural cross-language interactions at enterprise scale in its customer service operations
Safety, Enterprise Controls & Customisation
Alongside the new capabilities, OpenAI has integrated active content classifiers to halt harmful content in real time, alongside developer tools for adding additional safeguards appropriate to specific use cases. The platform supports EU Data Residency and adheres to OpenAI's enterprise privacy commitments — addressing a critical requirement for regulated industry adoption.
A notable new capability in the text-to-speech layer allows developers to instruct the model on how to speak — for example, "talk like a sympathetic customer service agent" — unlocking a new level of voice personality customisation for branded applications. The OpenAI Agents SDK has also been updated with a dedicated voice agent module, and GPT‑Realtime‑2 can be wired directly into multi-agent workflows.
"Voice is becoming one of the most natural ways for people to use software. It lets someone ask for help while driving, change a travel plan while walking through an airport, get support in their preferred language, or move through a task without stopping to type."— OpenAI, Official Announcement
