Tongyi Unveils First Film-Level Voice AI Model: Emotionally Intelligent Speech Achieved

After AIGC revolutionized image and text generation, the final frontier in film and television—voice acting—is now being breached by Alibaba's Tongyi Lab. On March 16, Tongyi Lab officially launched and open-sourced the world's first multimodal large model for cinematic, multi-scenario voice acting, Fun-CineForge.
For years, AI voice synthesis has been plagued by "robotic" and "announcer-style" tones. In film and television, capturing emotional depth, ambient sound mixing, and lip sync remained significant hurdles. Fun-CineForge was created specifically to overcome these challenges.
This model introduces a groundbreaking "data + model" integrated design. Alongside the model, Tongyi Lab provided a method for constructing high-quality datasets. This enables AI to move beyond mere text reading to deeply understand complex cinematic contexts, replicating subtle emotional nuances and spatial audio effects.
As the newest member of the Alibaba Tongyi family, the open-source Fun-CineForge is a game-changer. It offers video creators a "cinema-grade" post-production tool and, through accessible technology, allows indie creators and mid-budget productions to achieve high-quality, multilingual dubbing at minimal cost.
From the earlier Qwen3-Omni to the current Fun-CineForge , the Tongyi series is rapidly completing the multimodal puzzle. As AI learns to "perform like a human," the entire landscape of film translation and post-production could be reshaped. The model and its dataset construction plan are now available on major open-source platforms, signaling that the era of "cinema-grade AI" is arriving sooner than anticipated.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500

After AIGC revolutionized image and text generation, the final frontier in film and television—voice acting—is now being breached by Alibaba's Tongyi Lab. On March 16,
For years, AI voice synthesis has been plagued by "robotic" and "announcer-style" tones. In film and television, capturing emotional depth, ambient sound mixing, and lip sync remained significant hurdles.
This model introduces a groundbreaking "data + model" integrated design. Alongside the model,
As the newest member of the Alibaba Tongyi family, the open-source
From the earlier
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






