OpenAI Releases Three New Realtime Voice Models for the API With GPT-5-Class Reasoning | Free Download

OpenAI has introduced Three new real-time audio models For its API: gpt-realtime-2, gpt-realtime-translate, And gpt-realtime-whisper. These models are now accessible in Realtime APIs and Playgrounds, allowing developers to incorporate them into existing applications via codecs.

The new tools expand voice functionality from basic turn-based interactions to include real-time reasoning, multi-language translation, and live streaming transcription.

OpenAI’s new realtime audio models: GPT-Realtime-2, Translate and Whisper

gpt-realtime-2 OpenAI’s first live voice model with reasoning capabilities comparable to GPT-5. It is designed to handle complex requests, call tools, and recover from interruptions during ongoing conversations. The main update to gpt-realtime-1.5 includes an adjustable logic effort with settings for minimum, low, medium, high, and very high, with low included by default.

Its context window is expanded 32,000 to 128,000 tokensSupports long workflows. The model can call multiple tools in parallel, providing audible status updates such as “checking your calendar” or “looking at it now”. It also includes preambles that allow short phrases like “let me check this” to be said before completing a request.

Improvements have been made to the understanding of domain-specific terminology, including proper nouns and health care terminology. Additionally, the model offers more controllable tone and delivery.

gpt-realtime-translate Provides live translation from over 70 input languages ​​into 13 output languages, keeping pace with the speaker. It is intended for use in cross-border customer support, live events, education platforms, and creator tools serving global audiences. Deutsche Telekom is testing the model for multilingual customer support, while Vimeo is experimenting with translating product education videos in real time as they play.

gpt-realtime-whisper is a streaming speech-to-text model designed for low-latency transcription. It transcribes audio as it is spoken, making it suitable for applications such as live captioning, meeting notes that update during conversations, voice assistants that require continuous understanding, and post-call workflows in areas such as customer support, healthcare, and sales.

Pricing, security, and compliance for OpenAI’s Realtime Audio API

Pricing details include several options:

gpt-realtime-2The cost is $32 per million audio input tokens, $0.40 per million cached input tokens, and $64 per million audio output tokens.

gpt-realtime-translate Charges are $0.034 per minute.

gpt-realtime-whisper Cost $0.017 per minute.

The Realtime API features proactive classifiers that can block conversations that violate OpenAI’s content policies. Developers can increase security by adding additional guardrails using the Agent SDK. The API also supports EU data residency for applications located in the EU and complies with OpenAI’s enterprise privacy standards.

According to OpenAI’s usage policies, developers are required to inform users when interacting with AI, unless the context clearly indicates so.

Add Ghacks as a favorite source on Google

Source:Ghacks

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top