StepZen Unveils StepAudio 3 Series Voice Model, Securing Multiple Global Firsts

On September 15, StepXingchen officially released the new StepAudio3 series of voice large models. This launch features five distinct products: StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music. Several of these models have secured the top global position in the authoritative Artificial Analysis evaluation list. These five models are now fully accessible on the StepXingchen open platform, addressing five core application scenarios: realistic voice generation, comprehensive audio content creation, real-time voice interaction, complex scenario voice understanding, and music production.
For real-time interaction, StepAudio3Realtime is engineered to deliver full-duplex native conversations. It achieved a global ranking of first place with a comprehensive score of 98.9% in the Artificial Analysis Conversational Dynamics ranking, and also took the top spot in the Artificial Analysis Speech Reasoning ranking with an accuracy of 99.7%. The model accurately captures conversational rhythm, handles user interruptions and continuous feedback, and supports the comprehension of semantics, tone, emotion, paralanguage, and ambient sounds. Furthermore, it enables parallel reasoning and voice generation, along with asynchronous tool calls (Tool Call) and long tasks, ensuring uninterrupted voice conversations.
Regarding speech recognition, StepAudio3ASR combines high-precision transcription with the knowledge reasoning capabilities of large language models, surpassing traditional audio-only transcription methods. It supports diverse scenarios including Chinese, English, dialects, mixed language, long audio, and specialized fields. The model excels in medical, legal, financial, automotive, and programming contexts, and effectively manages complex input environments such as low volume, fast speech rates, unclear pronunciation, singing, and background music. Its non-streaming speech recognition word error rate (WER) stands at just 1.7%, tying for first place globally.
In voice generation, StepAudio3TTS is a realistic model designed for real-time interaction. It aligns perfectly with natural human speech patterns regarding voice base, intonation, rhythm, and pauses. The model reproduces paralinguistic expressions like laughter, hesitation, stammering, repetition, and correction, and can autonomously adjust emotional tones based on semantic content. Utilizing a streaming generation architecture, the model synchronizes generation and playback.
For comprehensive audio content generation, StepAudio3Gen streamlines the traditional long production, editing, and mixing process. Users can generate vocals, sound effects, ambient sounds, and background music simultaneously using natural language descriptions and reference audio. The model allows fine control over paralinguistic features such as character voice, speaking style, emotion, dialects, and laughter. It also enables users to specify the timing and sequence of dialogue, sound effects, and background music, directly influencing the entire content's sound design and time sequencing.
In music creation, StepAudio3Music supports both zero-shot generation and multi-round interactive creation based on ABC notation. Users can input lyrics, a cappella, reference songs, or notation, and use natural language to control genre, vocals, melody, rhythm, instruments, and emotions. The model deeply understands song structure, cross-paragraph melodic development, and energy shifts. Particularly in a cappella accompaniment scenarios, a single a cappella track can automatically generate harmony, rhythm, instruments, and complete arrangements, facilitating true music production.
Related article
xAI Co-Founder Steps Down as Only Five Original Team Members Remain
xAI, Elon Musk’s artificial intelligence venture, has undergone another round of significant staff turnover. Co-founder Toby Pohlen has officially announced his exit. Sharing his thoughts on X, Pohlen reflected on his three-year tenure, noting he han
Kimi K3 Set for Launch This Month With 2.5 Trillion Parameters
The large model landscape has intensified in the second half of 2026. Alongside DeepSeek V4, Moonshot AI’s Kimi K3 is confirmed for release this month.While specific dates remain undisclosed, insiders report Kimi K3 boasts up to 2.5 trillion paramete
Jia Zhangke\'s \'Dunhuang Mama\' Registered: Solo Mother Falls for AI, Travels China
In July 2026, the National Film Administration approved the script outline for Jia Zhangke’s upcoming film, "Dunhuang Mama." The synopsis depicts a contemporary story of a solitary mother in Dunhuang who gradually turns to artificial intelligence to
Related Special Topic Recommendations
Comments (0)
0/500

On September 15, StepXingchen officially released the new StepAudio3 series of voice large models. This launch features five distinct products: StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music. Several of these models have secured the top global position in the authoritative Artificial Analysis evaluation list. These five models are now fully accessible on the StepXingchen open platform, addressing five core application scenarios: realistic voice generation, comprehensive audio content creation, real-time voice interaction, complex scenario voice understanding, and music production.
For real-time interaction, StepAudio3Realtime is engineered to deliver full-duplex native conversations. It achieved a global ranking of first place with a comprehensive score of 98.9% in the Artificial Analysis Conversational Dynamics ranking, and also took the top spot in the Artificial Analysis Speech Reasoning ranking with an accuracy of 99.7%. The model accurately captures conversational rhythm, handles user interruptions and continuous feedback, and supports the comprehension of semantics, tone, emotion, paralanguage, and ambient sounds. Furthermore, it enables parallel reasoning and voice generation, along with asynchronous tool calls (Tool Call) and long tasks, ensuring uninterrupted voice conversations.
Regarding speech recognition, StepAudio3ASR combines high-precision transcription with the knowledge reasoning capabilities of large language models, surpassing traditional audio-only transcription methods. It supports diverse scenarios including Chinese, English, dialects, mixed language, long audio, and specialized fields. The model excels in medical, legal, financial, automotive, and programming contexts, and effectively manages complex input environments such as low volume, fast speech rates, unclear pronunciation, singing, and background music. Its non-streaming speech recognition word error rate (WER) stands at just 1.7%, tying for first place globally.
In voice generation, StepAudio3TTS is a realistic model designed for real-time interaction. It aligns perfectly with natural human speech patterns regarding voice base, intonation, rhythm, and pauses. The model reproduces paralinguistic expressions like laughter, hesitation, stammering, repetition, and correction, and can autonomously adjust emotional tones based on semantic content. Utilizing a streaming generation architecture, the model synchronizes generation and playback.
For comprehensive audio content generation, StepAudio3Gen streamlines the traditional long production, editing, and mixing process. Users can generate vocals, sound effects, ambient sounds, and background music simultaneously using natural language descriptions and reference audio. The model allows fine control over paralinguistic features such as character voice, speaking style, emotion, dialects, and laughter. It also enables users to specify the timing and sequence of dialogue, sound effects, and background music, directly influencing the entire content's sound design and time sequencing.
In music creation, StepAudio3Music supports both zero-shot generation and multi-round interactive creation based on ABC notation. Users can input lyrics, a cappella, reference songs, or notation, and use natural language to control genre, vocals, melody, rhythm, instruments, and emotions. The model deeply understands song structure, cross-paragraph melodic development, and energy shifts. Particularly in a cappella accompaniment scenarios, a single a cappella track can automatically generate harmony, rhythm, instruments, and complete arrangements, facilitating true music production.
xAI Co-Founder Steps Down as Only Five Original Team Members Remain
xAI, Elon Musk’s artificial intelligence venture, has undergone another round of significant staff turnover. Co-founder Toby Pohlen has officially announced his exit. Sharing his thoughts on X, Pohlen reflected on his three-year tenure, noting he han
Kimi K3 Set for Launch This Month With 2.5 Trillion Parameters
The large model landscape has intensified in the second half of 2026. Alongside DeepSeek V4, Moonshot AI’s Kimi K3 is confirmed for release this month.While specific dates remain undisclosed, insiders report Kimi K3 boasts up to 2.5 trillion paramete
Jia Zhangke\'s \'Dunhuang Mama\' Registered: Solo Mother Falls for AI, Travels China
In July 2026, the National Film Administration approved the script outline for Jia Zhangke’s upcoming film, "Dunhuang Mama." The synopsis depicts a contemporary story of a solitary mother in Dunhuang who gradually turns to artificial intelligence to





Home






