OpenAI unveils advanced voice models for more natural real-time conversations

OpenAI has launched new conversational models, GPT-Live-1 and GPT-Live-1 mini, designed to sound more natural and manage turn-taking more effectively. These full-duplex models can speak and listen simultaneously, allowing users to interrupt naturally and enabling features such as live translation.
The company is also rolling out GPT-Live-1 mini as the default replacement for the current Advanced Voice Mode in ChatGPT. Paid-tier users will have access to the larger GPT-Live-1 model. Previously, the system relied on a combination of a speech-to-text model, a large language model for generating responses, and a text-to-speech model to produce the final output.
During a press briefing, OpenAI stated that the new models address issues like interrupting users while they are speaking and lacking sufficient intelligence to answer questions. The updated models will route queries to the latest text models, such as GPT-5.5, for search, reasoning, or agentic tasks while maintaining the conversation flow.
OpenAI also demonstrated that the model can remain silent for extended periods, absorbing conversational context until it is addressed. Additionally, since the new voice mode integrates with newer GPT models, it can present information visually. Other startups, such as Monogram—which raised $40 million in seed funding from DST and Lux Capital—are also focusing on visual responses to make assistants more interactive.
The company noted that the new voice mode in ChatGPT is built for longer conversations. At the briefing, ChatGPT Voice’s product lead, Atty Eleti, mentioned holding 30- to 40-minute conversations with the voice feature during walks.
OpenAI believes voice could become the primary interface for complex computing tasks. Reports suggest the company may launch AI-capable earbuds this year, though no details on hardware products were provided.
“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work,” Eleti said.
OpenAI has been enhancing voice-based features over the past few years to make ChatGPT’s voice mode sound more natural. The company reports that over 150 million people use Voice and Dictation features to talk to ChatGPT.
Competitors are also working to make their assistants more expressive.
Both Apple and Amazon have updated their assistants to be more conversational, with improved context handling. Startups like Sesame, co-founded by Oculus co-founder Brendan Iribe and Ankit Kumar, have launched AI assistants that engage in more natural conversation while completing background tasks.
OpenAI is following a similar path, aiming to let users interact with its assistant hands-free for longer periods. Despite claiming the new voice mode sounds more natural, the company emphasized that it is not designed as an AI companion. It noted that the new models include safeguards to provide age-appropriate responses to teens and offer resources if conversations touch on topics like self-harm.
The new voice mode still has room for improvement. During a demo of the live translation feature for Hindi, the assistant spoke with a heavy American accent and used unnatural, slightly bookish Hindi. The company said the mode is optimized for “most spoken languages” but did not specify which ones.
Related article
Warner Music acquires AI attribution startup Sureel AI
Warner Music Group (WMG) confirmed on Wednesday that it is acquiring Sureel AI, an artificial intelligence attribution startup. Sureel’s proprietary technology generates “AI DNA” for musical tracks, deconstructing them into constituent elements to tr
Amazon introduces Alexa for Shopping while pushing Rufus to the background
Amazon has launched Alexa for Shopping, merging its Rufus shopping chatbot with Alexa+ across the app, website, and Echo Show devices.The assistant answers product queries, compares items, tracks prices, and supports shopping reminders. It also handl
Microsoft, Azure and AI Tech Combat California Wildfire Risks
Microsoft invests in AI-driven wildfire detection, with Juan Lavista Ferres, CVP and Chief Data Scientist, discussing strategies to mitigate environmental damage.According to NASA, climate change impacts everyone on Earth, manifesting as rising tempe
Related Special Topic Recommendations
Comments (0)
0/500

OpenAI has launched new conversational models, GPT-Live-1 and GPT-Live-1 mini, designed to sound more natural and manage turn-taking more effectively. These full-duplex models can speak and listen simultaneously, allowing users to interrupt naturally and enabling features such as live translation.
The company is also rolling out GPT-Live-1 mini as the default replacement for the current Advanced Voice Mode in ChatGPT. Paid-tier users will have access to the larger GPT-Live-1 model. Previously, the system relied on a combination of a speech-to-text model, a large language model for generating responses, and a text-to-speech model to produce the final output.
During a press briefing, OpenAI stated that the new models address issues like interrupting users while they are speaking and lacking sufficient intelligence to answer questions. The updated models will route queries to the latest text models, such as GPT-5.5, for search, reasoning, or agentic tasks while maintaining the conversation flow.
OpenAI also demonstrated that the model can remain silent for extended periods, absorbing conversational context until it is addressed. Additionally, since the new voice mode integrates with newer GPT models, it can present information visually. Other startups, such as Monogram—which raised $40 million in seed funding from DST and Lux Capital—are also focusing on visual responses to make assistants more interactive.
The company noted that the new voice mode in ChatGPT is built for longer conversations. At the briefing, ChatGPT Voice’s product lead, Atty Eleti, mentioned holding 30- to 40-minute conversations with the voice feature during walks.
OpenAI believes voice could become the primary interface for complex computing tasks. Reports suggest the company may launch AI-capable earbuds this year, though no details on hardware products were provided.
“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work,” Eleti said.
OpenAI has been enhancing voice-based features over the past few years to make ChatGPT’s voice mode sound more natural. The company reports that over 150 million people use Voice and Dictation features to talk to ChatGPT.
Competitors are also working to make their assistants more expressive.
Both Apple and Amazon have updated their assistants to be more conversational, with improved context handling. Startups like Sesame, co-founded by Oculus co-founder Brendan Iribe and Ankit Kumar, have launched AI assistants that engage in more natural conversation while completing background tasks.
OpenAI is following a similar path, aiming to let users interact with its assistant hands-free for longer periods. Despite claiming the new voice mode sounds more natural, the company emphasized that it is not designed as an AI companion. It noted that the new models include safeguards to provide age-appropriate responses to teens and offer resources if conversations touch on topics like self-harm.
The new voice mode still has room for improvement. During a demo of the live translation feature for Hindi, the assistant spoke with a heavy American accent and used unnatural, slightly bookish Hindi. The company said the mode is optimized for “most spoken languages” but did not specify which ones.
Warner Music acquires AI attribution startup Sureel AI
Warner Music Group (WMG) confirmed on Wednesday that it is acquiring Sureel AI, an artificial intelligence attribution startup. Sureel’s proprietary technology generates “AI DNA” for musical tracks, deconstructing them into constituent elements to tr
Microsoft, Azure and AI Tech Combat California Wildfire Risks
Microsoft invests in AI-driven wildfire detection, with Juan Lavista Ferres, CVP and Chief Data Scientist, discussing strategies to mitigate environmental damage.According to NASA, climate change impacts everyone on Earth, manifesting as rising tempe





Home






