Tongyi Lab Debuts Fun-CosyVoice3.5 and Fun-AudioGen-VD Speech Models
Today, Tongyi Lab officially unveiled two FreeStyle-enabled voice generation models: Fun-CosyVoice3.5 and Fun-AudioGen-VD. This launch signifies a paradigm shift in speech synthesis, moving from reliance on preset tags to a new framework based on natural language instructions. It achieves a deeply interactive experience, enabling users to "generate speech freely with a single sentence."


Regarding technical architecture and functional upgrades, Fun-CosyVoice3.5 emphasizes multilingual voice cloning and nuanced expression, now adding support for four new languages, including Thai and Indonesian. By integrating DiffRO and GRPO reinforcement learning technologies, the model achieves substantial improvements in prosody and audio quality similarity. Its error rate for rare characters has decreased from 15.2% to 5.3%, and initial packet delay has been reduced by 35%. Complementing this, Fun-AudioGen-VD focuses on sound design and scenario modeling. It supports precise, instruction-based control over gender, emotion, and spatial acoustics, enabling the simulation of complex, integrated scenarios—from a "crazy villain" to a "noisy café" ambiance.
From an industry trend perspective, Tongyi Lab 's initiative elevates speech generation from a simple conversion tool to a full-fledged creation tool. This descriptive and programmable digital expression capability directly empowers sectors like film, gaming, and AI avatars. It reduces content creation costs while significantly expanding the semantic richness of human-computer interaction.
API: https://help.aliyun.com/zh/model-studio/text-to-speech?spm=a2c4g.11186623.help-menu-2400256.d_0_3_2_0.d5536a31V2tEJP
Documentation: https://help.aliyun.com/zh/model-studio/cosyvoice-clone-api?spm=a2c4g.11186623.help-menu-search-2400256.d_2
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
Today,


Regarding technical architecture and functional upgrades,
From an industry trend perspective,
API: https://help.aliyun.com/zh/model-studio/text-to-speech?spm=a2c4g.11186623.help-menu-2400256.d_0_3_2_0.d5536a31V2tEJP
Documentation: https://help.aliyun.com/zh/model-studio/cosyvoice-clone-api?spm=a2c4g.11186623.help-menu-search-2400256.d_2
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






