option
Home
Flash News
Content
GaryPerez
GaryPerez
September 16, 2026

iFLYTEK released Spark-Audio-1.0-Preview, a speech foundation model trained on fully domestic computing power. Unlike traditional cascading systems that convert speech to text, this model directly understands audio, capturing semantics, emotions, and ambient sounds. Built with a 0.65B audio encoder and 30B-A3B LLM, it supports 99 languages and 202 dialects. It outperforms competitors in noisy scenarios and achieves SOTA in Chinese-English recognition. The model enables tasks like sentiment analysis and complex Q&A, marking a shift from transcription to true audio understanding. API access will follow via the iFLYTEK Open Platform.

iFLYTEK released Spark-Audio-1.0-Preview, a speech foundation model trained on fully domestic computing power. Unlike traditional cascading systems that convert speech to text, this model directly understands audio, capturing semantics, emotions, and ambient sounds. Built with a 0.65B audio encoder and 30B-A3B LLM, it supports 99 languages and 202 dialects. It outperforms competitors in noisy scenarios and achieves SOTA in Chinese-English recognition. The model enables tasks like sentiment analysis and complex Q&A, marking a shift from transcription to true audio understanding. API access will follow via the iFLYTEK Open Platform.
Comments (0)
0/300
OR