Home
iFlytek Xinghuo Unveils 0.65B Encoder Speech Model with 30B MoE, Built on Fully Domestic Computing Power
iFLYTEK has launched Spark-Audio-1.0-Preview, a speech foundation large model built entirely on domestic computing infrastructure.
Previously, audio processing with large models relied on a cascading method: converting speech to text first, then passing it to the model for analysis. This approach suffers from two major flaws: information loss, as text fails to capture tone, emotion, and background sounds, leading to incomplete context; and fragmented workflows, where independent modules for transcription, speaker ID, and translation run in series, causing error accumulation, latency, and a poor user experience.

Beyond transcription: Direct speech comprehension
This speech foundation model moves beyond simple transcription to directly interpret audio. It can distinguish human voices, ambient noise, and music, while grasping semantics, emotions, and contextual cues within the sound.
The initial Spark-Audio-1.0-Preview release features a 0.65B audio encoder (Dense structure) and a 30B-A3B large language model (MoE structure). Trained on 13 million hours of audio and extensive text data using a fully domestic computing cluster, the model accepts both text and audio inputs. It supports tasks including speech transcription, multilingual translation, multi-dialect recognition, ambient sound identification, speaker recognition, sentiment analysis, and complex audio Q&A, enabling machines to transition from "hearing clearly" to "understanding." Similar to iFLYTEK's Spark large model "1+N" system, users can fine-tune this speech foundation model to create specialized systems for speech recognition, real-time translation, and interactive speech applications.
Supports 99 languages and 202 dialects, excelling in complex scenarios
According to official data, Spark-Audio-1.0-Preview performs comparably or better than competitors across various speech evaluation tasks, particularly in challenging real-world scenarios like high noise levels and soft speaking. Its performance is also closely aligned with larger, closed-source speech foundation models.
The model supports recognition of 99 languages and 202 dialects. It demonstrates superior average performance in multilingual and multi-dialect speech recognition compared to the similarly sized Qwen3.5-omni-flash (35B-A3B), and matches the performance of the significantly larger Qwen3.5-omni-plus. In evaluations such as Fleurs, Kespeech, and LibriSpeech for Chinese-English, multilingual, and multi-dialect speech recognition, it outperforms Gemini-3.1 Pro and achieves SOTA on the Fleurs Chinese test set.
iFLYTEK acknowledges that while this version maintains strong performance in general knowledge, math, and coding tasks, there is room for improvement in instruction following, dialogue, and audio understanding. Future updates will focus on strengthening these areas. Spark-Audio-1.0-Preview is now available for public experience, with API access to be released on the iFLYTEK Open Platform soon.
Related article
Zoho Helps Businesses Adopt AI Agents
The market is rapidly evolving toward AI-driven agentsZoho highlights that AI is advancing exponentially, making it crucial for businesses to adapt and rethink their operationsThe excitement surrounding AI shows no signs of slowing down. Keeping up w
Suno Unveils Major Update With Advanced Track Separation and Car Support
Suno has rolled out significant updates to its AI music generation platform, enhancing both web and mobile experiences to streamline creation and boost usability. The platform is advancing toward a more accessible and efficient stage for AI-driven mu
Cohere’s Joëlle Pineau on Sovereign AI and Enterprise Security
IDC research commissioned by Cohere explores sovereign AI, with Chief AI Officer Joëlle Pineau outlining essential strategies for enterprise cloud security and data governance.According to IDC’s report, The State of Sovereign AI Adoption in 2026, com
Related Special Topic Recommendations
Comments (0)
0/500
iFLYTEK has launched Spark-Audio-1.0-Preview, a speech foundation large model built entirely on domestic computing infrastructure.
Previously, audio processing with large models relied on a cascading method: converting speech to text first, then passing it to the model for analysis. This approach suffers from two major flaws: information loss, as text fails to capture tone, emotion, and background sounds, leading to incomplete context; and fragmented workflows, where independent modules for transcription, speaker ID, and translation run in series, causing error accumulation, latency, and a poor user experience.

Beyond transcription: Direct speech comprehension
This speech foundation model moves beyond simple transcription to directly interpret audio. It can distinguish human voices, ambient noise, and music, while grasping semantics, emotions, and contextual cues within the sound.
The initial Spark-Audio-1.0-Preview release features a 0.65B audio encoder (Dense structure) and a 30B-A3B large language model (MoE structure). Trained on 13 million hours of audio and extensive text data using a fully domestic computing cluster, the model accepts both text and audio inputs. It supports tasks including speech transcription, multilingual translation, multi-dialect recognition, ambient sound identification, speaker recognition, sentiment analysis, and complex audio Q&A, enabling machines to transition from "hearing clearly" to "understanding." Similar to iFLYTEK's Spark large model "1+N" system, users can fine-tune this speech foundation model to create specialized systems for speech recognition, real-time translation, and interactive speech applications.
Supports 99 languages and 202 dialects, excelling in complex scenarios
According to official data, Spark-Audio-1.0-Preview performs comparably or better than competitors across various speech evaluation tasks, particularly in challenging real-world scenarios like high noise levels and soft speaking. Its performance is also closely aligned with larger, closed-source speech foundation models.
The model supports recognition of 99 languages and 202 dialects. It demonstrates superior average performance in multilingual and multi-dialect speech recognition compared to the similarly sized Qwen3.5-omni-flash (35B-A3B), and matches the performance of the significantly larger Qwen3.5-omni-plus. In evaluations such as Fleurs, Kespeech, and LibriSpeech for Chinese-English, multilingual, and multi-dialect speech recognition, it outperforms Gemini-3.1 Pro and achieves SOTA on the Fleurs Chinese test set.
iFLYTEK acknowledges that while this version maintains strong performance in general knowledge, math, and coding tasks, there is room for improvement in instruction following, dialogue, and audio understanding. Future updates will focus on strengthening these areas. Spark-Audio-1.0-Preview is now available for public experience, with API access to be released on the iFLYTEK Open Platform soon.
Zoho Helps Businesses Adopt AI Agents
The market is rapidly evolving toward AI-driven agentsZoho highlights that AI is advancing exponentially, making it crucial for businesses to adapt and rethink their operationsThe excitement surrounding AI shows no signs of slowing down. Keeping up w
Suno Unveils Major Update With Advanced Track Separation and Car Support
Suno has rolled out significant updates to its AI music generation platform, enhancing both web and mobile experiences to streamline creation and boost usability. The platform is advancing toward a more accessible and efficient stage for AI-driven mu
Cohere’s Joëlle Pineau on Sovereign AI and Enterprise Security
IDC research commissioned by Cohere explores sovereign AI, with Chief AI Officer Joëlle Pineau outlining essential strategies for enterprise cloud security and data governance.According to IDC’s report, The State of Sovereign AI Adoption in 2026, com











