Home
Microsoft Unveils First Self-Developed Full-Duplex AI Speech Model Capable of Simultaneous Listening and Speaking Across 16 Languages
According to Technology media TestingCatalog, Microsoft is currently testing its inaugural native real-time voice model, MAI Realtime. Supporting 16 languages and two distinct voice styles, this system enables simultaneous listening and speaking, allowing for automatic language detection and mid-conversation switching. This advancement eliminates the need for users to wait for the AI to finish speaking before responding, creating a conversation experience that closely mirrors natural human interaction.

MAI Realtime employs a bidirectional, full-duplex interaction approach, effectively breaking the traditional voice assistant pattern where users speak and the model answers sequentially. By utilizing endpoint detection technology to identify speech starts, pauses, and ends, the system can immediately adjust its output when users interject, achieving synchronized listening and responding. This represents a significant milestone for Microsoft's MAI voice model family, transitioning from one-way to two-way communication. Previous models, such as MAI-Voice-2 and MAI-Transcribe-1.5, were limited to one-way tasks like speech synthesis or recognition.
Two voice styles offer greater naturalism, though singing capabilities are not yet included
Regarding language coverage, MAI Realtime supports 16 languages, including Chinese, English, Japanese, Korean, French, German, and Arabic. Users can either manually specify the conversation language or enable auto-detection, allowing the model to automatically identify and switch languages during the dialogue. In terms of voice options, the test version provides two voices, Victoria and Grant, which are reported to sound more natural than Copilot's current voice mode.
Related article
Apple unveils iOS 27 with local AI and Google partnership boosting Siri
The Information recently highlighted Apple’s strategic approach to integrating artificial intelligence into the upcoming OS 27. By leveraging Google’s Gemini model to train a more efficient, lightweight AI, Apple aims to deliver robust local edge AI
Perplexity Computer Unveils Hybrid Inference to Balance Privacy and Efficiency
Perplexity, an artificial intelligence firm, has revealed plans to launch a significant enhancement to its Perplexity Computer agent this summer. This update introduces a "hybrid agent reasoning" mode that dynamically allocates workloads between loca
NexCOBOT tackles physical AI market barriers and growth
NexCOBOT, showcased at the Robotics Summit & Expo, offers a diverse range of controllers. Source: NexCOBOTHumanoid robotics and physical AI firms are not only securing billions in funding; they are also being acquired by major technology corporations
Related Special Topic Recommendations
Comments (0)
0/500
According to Technology media TestingCatalog, Microsoft is currently testing its inaugural native real-time voice model, MAI Realtime. Supporting 16 languages and two distinct voice styles, this system enables simultaneous listening and speaking, allowing for automatic language detection and mid-conversation switching. This advancement eliminates the need for users to wait for the AI to finish speaking before responding, creating a conversation experience that closely mirrors natural human interaction.

MAI Realtime employs a bidirectional, full-duplex interaction approach, effectively breaking the traditional voice assistant pattern where users speak and the model answers sequentially. By utilizing endpoint detection technology to identify speech starts, pauses, and ends, the system can immediately adjust its output when users interject, achieving synchronized listening and responding. This represents a significant milestone for Microsoft's MAI voice model family, transitioning from one-way to two-way communication. Previous models, such as MAI-Voice-2 and MAI-Transcribe-1.5, were limited to one-way tasks like speech synthesis or recognition.
Two voice styles offer greater naturalism, though singing capabilities are not yet included
Regarding language coverage, MAI Realtime supports 16 languages, including Chinese, English, Japanese, Korean, French, German, and Arabic. Users can either manually specify the conversation language or enable auto-detection, allowing the model to automatically identify and switch languages during the dialogue. In terms of voice options, the test version provides two voices, Victoria and Grant, which are reported to sound more natural than Copilot's current voice mode.
Apple unveils iOS 27 with local AI and Google partnership boosting Siri
The Information recently highlighted Apple’s strategic approach to integrating artificial intelligence into the upcoming OS 27. By leveraging Google’s Gemini model to train a more efficient, lightweight AI, Apple aims to deliver robust local edge AI
Perplexity Computer Unveils Hybrid Inference to Balance Privacy and Efficiency
Perplexity, an artificial intelligence firm, has revealed plans to launch a significant enhancement to its Perplexity Computer agent this summer. This update introduces a "hybrid agent reasoning" mode that dynamically allocates workloads between loca
NexCOBOT tackles physical AI market barriers and growth
NexCOBOT, showcased at the Robotics Summit & Expo, offers a diverse range of controllers. Source: NexCOBOTHumanoid robotics and physical AI firms are not only securing billions in funding; they are also being acquired by major technology corporations











