Home
Alibaba's Aliyun Unveils Fun-CineForge: Open-Sourcing Movie-Grade Dubbing Model and Dataset
Recently, the Fun-CineForge project, developed by the speech team at Alibaba Tongyi Lab in collaboration with the University of Science and Technology of China, has been officially open-sourced. This initiative tackles core challenges in film and television dubbing—such as lip synchronization, voice style transfer, and emotional expression—by introducing a comprehensive end-to-end production workflow and large model solutions.

Core Breakthroughs: Solving the "Out-of-Sync" Problem in Film Dubbing
Traditional AI dubbing often struggles with issues like mismatched lip movements, robotic emotional delivery, and difficulty adapting to complex cinematic scenes involving dialogue and multi-person acoustics. Fun-CineForge achieves a significant breakthrough through two key innovations:
MLLM Dubbing Model: Moving beyond simple lip-area audio-video alignment, it employs a multimodal large language model (MLLM) architecture capable of deeply understanding a character's identity and emotional nuances within a scene.
CineDub Large-Scale Dataset: The project created the first richly annotated Chinese TV show dubbing dataset via an automated pipeline, covering diverse scenarios like monologues, narration, dialogue, and multi-speaker interactions.
Project Updates and Open Source Roadmap
The project has seen frequent recent updates, indicating a high level of engineering maturity:
January to March 2026: Released sample datasets and demonstration demos for both Chinese (CineDub-CN) and English (CineDub-EN).
March 16, 2026: Officially open-sourced the inference code and model weights (Checkpoints), allowing developers to access these resources on GitHub.
Dataset Access: Several classic series datasets are now available for research, including the Chinese series "Dream of the Red Chamber" and the English series "Downton Abbey."
Technical Application: From "Dialogue" to "Performance"
Official demos show the model delivering impressive results in remaking classic series like "Romance of the Three Kingdoms." By inputting specific "emotional clues," the model can accurately capture a character's emotional shift—from fear to defiance—achieving high-fidelity voice cloning and natural lip sync.
The launch of Fun-CineForge signals a shift in film and TV AI dubbing from basic "text-to-speech" to an "automated post-production" tool with artistic comprehension. This advancement is poised to significantly reduce production costs for dubbed film and television content.
Project: https://funcineforge.github.io/
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
Recently, the Fun-CineForge project, developed by the speech team at Alibaba Tongyi Lab in collaboration with the University of Science and Technology of China, has been officially open-sourced. This initiative tackles core challenges in film and television dubbing—such as lip synchronization, voice style transfer, and emotional expression—by introducing a comprehensive end-to-end production workflow and large model solutions.

Core Breakthroughs: Solving the "Out-of-Sync" Problem in Film Dubbing
Traditional AI dubbing often struggles with issues like mismatched lip movements, robotic emotional delivery, and difficulty adapting to complex cinematic scenes involving dialogue and multi-person acoustics. Fun-CineForge achieves a significant breakthrough through two key innovations:
MLLM Dubbing Model: Moving beyond simple lip-area audio-video alignment, it employs a multimodal large language model (MLLM) architecture capable of deeply understanding a character's identity and emotional nuances within a scene.
CineDub Large-Scale Dataset: The project created the first richly annotated Chinese TV show dubbing dataset via an automated pipeline, covering diverse scenarios like monologues, narration, dialogue, and multi-speaker interactions.
Project Updates and Open Source Roadmap
The project has seen frequent recent updates, indicating a high level of engineering maturity:
January to March 2026: Released sample datasets and demonstration demos for both Chinese (CineDub-CN) and English (CineDub-EN).
March 16, 2026: Officially open-sourced the inference code and model weights (Checkpoints), allowing developers to access these resources on GitHub.
Dataset Access: Several classic series datasets are now available for research, including the Chinese series "Dream of the Red Chamber" and the English series "Downton Abbey."
Technical Application: From "Dialogue" to "Performance"
Official demos show the model delivering impressive results in remaking classic series like "Romance of the Three Kingdoms." By inputting specific "emotional clues," the model can accurately capture a character's emotional shift—from fear to defiance—achieving high-fidelity voice cloning and natural lip sync.
The launch of Fun-CineForge signals a shift in film and TV AI dubbing from basic "text-to-speech" to an "automated post-production" tool with artistic comprehension. This advancement is poised to significantly reduce production costs for dubbed film and television content.
Project: https://funcineforge.github.io/
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











