OpenAI Unveils Three Real-Time Speech Models with GPT-5-Level Reasoning

OpenAI, the AI giant, has once again expanded the frontiers of voice technology with the official launch of three new real-time voice models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. These models are now integrated into the Realtime API for developers, targeting common voice interaction challenges like high latency, lack of natural interruption handling, and multilingual support issues.
The standout feature of this release is GPT-Realtime-2, described as the most intelligent AI voice model to date and the first voice tool with GPT-5-level reasoning. Unlike conventional voice assistants, it enables highly natural, fluid conversations while simultaneously performing real-time logical reasoning, seamlessly calling external tools, and accurately detecting and responding to user interruptions or corrections. This breakthrough signals that future voice assistants will evolve beyond simple command executors into real-time collaborative partners capable of managing multi-step complex tasks.
Pricing for GPT-Realtime-2 is set at $32 per million tokens for audio input (approximately RMB 218) and $64 per million tokens for output (approximately RMB 436). Cached input costs are significantly lower, at just $0.4 per million tokens.
Beyond the core reasoning model, the two other functional models offer distinct capabilities. GPT-Realtime-Translate delivers strong translation performance, supporting real-time conversion across 70 input languages and 13 output languages. Its translation speed nearly matches the speaker's pace, making it ideal for high-demand real-time communication scenarios like international meetings. GPT-Realtime-Whisper focuses on ultra-low latency streaming transcription, delivering a 'voice follows the person' experience that significantly reduces waiting times for meeting notes and live subtitles. These two models use more flexible billing, charged per minute at $0.034 and $0.017 respectively.
Industry analysts view OpenAI's latest moves as a shift in AI voice interaction from 'simple response' to 'deep real-time understanding,' further cementing the company's technological leadership in the intelligent era.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500

OpenAI, the AI giant, has once again expanded the frontiers of voice technology with the official launch of three new real-time voice models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. These models are now integrated into the Realtime API for developers, targeting common voice interaction challenges like high latency, lack of natural interruption handling, and multilingual support issues.
The standout feature of this release is GPT-Realtime-2, described as the most intelligent AI voice model to date and the first voice tool with GPT-5-level reasoning. Unlike conventional voice assistants, it enables highly natural, fluid conversations while simultaneously performing real-time logical reasoning, seamlessly calling external tools, and accurately detecting and responding to user interruptions or corrections. This breakthrough signals that future voice assistants will evolve beyond simple command executors into real-time collaborative partners capable of managing multi-step complex tasks.
Pricing for GPT-Realtime-2 is set at $32 per million tokens for audio input (approximately RMB 218) and $64 per million tokens for output (approximately RMB 436). Cached input costs are significantly lower, at just $0.4 per million tokens.
Beyond the core reasoning model, the two other functional models offer distinct capabilities. GPT-Realtime-Translate delivers strong translation performance, supporting real-time conversion across 70 input languages and 13 output languages. Its translation speed nearly matches the speaker's pace, making it ideal for high-demand real-time communication scenarios like international meetings. GPT-Realtime-Whisper focuses on ultra-low latency streaming transcription, delivering a 'voice follows the person' experience that significantly reduces waiting times for meeting notes and live subtitles. These two models use more flexible billing, charged per minute at $0.034 and $0.017 respectively.
Industry analysts view OpenAI's latest moves as a shift in AI voice interaction from 'simple response' to 'deep real-time understanding,' further cementing the company's technological leadership in the intelligent era.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






