Alibaba Tongyi unveils voice model with 'FreeStyle' natural language control
Today, Alibaba Tongyi Lab's Speech Team introduced two groundbreaking voice generation models: Fun-CosyVoice3.5 and Fun-AudioGen-VD. The standout feature of these models is their support for "FreeStyle" commands. Instead of complex parameter adjustments, users can precisely control vocal expression styles or build intricate audio scenes from scratch using simple natural language descriptions.

Each model serves distinct purposes:
Fun-CosyVoice3.5: Multilingual Replication and Fine-Grained Control
This enhanced version of CosyVoice achieves core breakthroughs in understanding speech expression nuances.
Command-Driven Generation: Users can input instructions like "speak more confidently" or "slow down with emotional variation" for real-time vocal adjustments.
Language Expansion: Added support for Thai, Indonesian, Portuguese, and Vietnamese maintains industry-leading performance in transcription accuracy (WER) and voice similarity across 13 languages.
Rare Character Optimization: Specialized training reduced error rates for uncommon characters from 15.2% to 5.3%.
Performance Boost: First packet latency decreased by 35%, significantly enhancing real-time interaction fluidity.
Fun-AudioGen-VD: Comprehensive Sound Design
This model acts as an "audio director," generating integrated audio combining "characters + environments."
Voice Customization: Specify gender, age, accent, and detailed characteristics like "hoarse, deep, or low-pitched" voices.
Emotion and Role Play: Simulates roles including customer service agents, broadcasters, and children, even conveying complex states like "outward calm with internal tension."
Immersive Environments: Adds background sounds (battlefield chaos, café murmurs) and spatial effects (cathedral reverb, underwater acoustics) for full spatial simulation.
Tongyi Lab notes these models will democratize high-quality voice creation, offering powerful AI support for podcasting, game development, and film post-production.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
Today, Alibaba Tongyi Lab's Speech Team introduced two groundbreaking voice generation models: Fun-CosyVoice3.5 and Fun-AudioGen-VD. The standout feature of these models is their support for "FreeStyle" commands. Instead of complex parameter adjustments, users can precisely control vocal expression styles or build intricate audio scenes from scratch using simple natural language descriptions.

Each model serves distinct purposes:
Fun-CosyVoice3.5: Multilingual Replication and Fine-Grained Control
This enhanced version of CosyVoice achieves core breakthroughs in understanding speech expression nuances.
Command-Driven Generation: Users can input instructions like "speak more confidently" or "slow down with emotional variation" for real-time vocal adjustments.
Language Expansion: Added support for Thai, Indonesian, Portuguese, and Vietnamese maintains industry-leading performance in transcription accuracy (WER) and voice similarity across 13 languages.
Rare Character Optimization: Specialized training reduced error rates for uncommon characters from 15.2% to 5.3%.
Performance Boost: First packet latency decreased by 35%, significantly enhancing real-time interaction fluidity.
Fun-AudioGen-VD: Comprehensive Sound Design
This model acts as an "audio director," generating integrated audio combining "characters + environments."
Voice Customization: Specify gender, age, accent, and detailed characteristics like "hoarse, deep, or low-pitched" voices.
Emotion and Role Play: Simulates roles including customer service agents, broadcasters, and children, even conveying complex states like "outward calm with internal tension."
Immersive Environments: Adds background sounds (battlefield chaos, café murmurs) and spatial effects (cathedral reverb, underwater acoustics) for full spatial simulation.
Tongyi Lab notes these models will democratize high-quality voice creation, offering powerful AI support for podcasting, game development, and film post-production.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






