Tongyi Lab Unveils PrismAudio for Bidirectional Audio-Visual Separation
In an era where AI video generation is rapidly expanding, the disconnect between visuals and audio—often manifesting as silence or mismatched sound—remains a significant hurdle to viewer immersion. Addressing this challenge, Alibaba’s Tongyi Lab has unveiled PrismAudio, a novel video-to-audio framework. This research, accepted by the premier AI conference ICLR 2026, focuses on automatically synchronizing video content with precise ambient sound effects.

Reasoning Before Generation: Mastering Audio with "Chain of Thought"
Conventional audio generation models typically operate on an "intuitive" basis, frequently producing disjointed results—such as a horse’s hoofbeats sounding like bird chirps, or audio lagging behind the visual action. PrismAudio distinguishes itself by employing a "plan first, then generate" approach.
Decomposition Chain of Thought: Prior to generating audio, the model analyzes the video frame by frame: identifying scene elements, determining sound onset timing, assessing audio characteristics (e.g., crispness or depth), and locating the sound source spatially.
Multi-Dimensional Evaluation: To ensure high-quality output, the development team integrated reinforcement learning, utilizing four "virtual evaluators" to score results across semantic consistency, temporal synchronization, aesthetic quality, and spatial accuracy. This multi-faceted feedback loop resolves the common issue where previous models optimized for one aspect while neglecting others.
Lightweight and Efficient: Processing 9-Second Videos in 0.6 Seconds
Beyond accuracy, PrismAudio delivers exceptional speed. Leveraging its proprietary Fast-GRPO training algorithm, the model achieves a substantial performance boost while maintaining operational efficiency:
Compact Architecture, High Impact: With only 518 million parameters, the model is significantly smaller than comparable systems, which often require tens of billions of parameters.
Rapid Response: Generating high-quality audio for a 9-second video clip takes merely 0.63 seconds, approaching real-time performance.
Industry Perspective: The Rise of Authentic Ambient Audio
PrismAudio’s launch not only offers a robust automation tool for film post-production and short-form video creation but also inspires new approaches to multi-object generation. As AI improves its ability to balance audio texture and spatial positioning, future video production will increasingly achieve true "what you see is what you hear" fidelity.
Paper link: arXiv:2511.18833
Open source link: https://prismaudio-project.github.io/
Related article
Neros Technologies Secures $250M to Deploy Defense Drones by 2026
Neros Archer: An FPV drone engineered for modular payloads and resilient communication links. | Source: NerosNeros Inc. has closed a $250 million Series C funding round, elevating the defense drone contractor’s valuation to $2.5 billion.“This latest
Yushu Technology, A-share's First Humanoid Robot Stock, Sets August 19 Listing With ~61 Billion Yuan Valuation
Scheduled for listing on the Sci-Tech Innovation Board on August 19, Yue Shu Technology will debut at 150.80 yuan per share, establishing a post-issue market capitalization of approximately 60.993 billion yuan. With a price-to-earnings ratio of 219.2
Ali Zhenwu M890 Ultra Node Now Compatible with Qwen3.8, Available on BaiLian Platform for Inference
Alibaba Cloud confirmed on July 23 that its Zhenwu M890 super node has successfully integrated with the latest flagship large model, Qwen3.8, making it available for model inference services on the Alibaba Cloud BaiLian platform. This milestone estab
Related Special Topic Recommendations
Comments (0)
0/500
In an era where AI video generation is rapidly expanding, the disconnect between visuals and audio—often manifesting as silence or mismatched sound—remains a significant hurdle to viewer immersion. Addressing this challenge, Alibaba’s Tongyi Lab has unveiled PrismAudio, a novel video-to-audio framework. This research, accepted by the premier AI conference ICLR 2026, focuses on automatically synchronizing video content with precise ambient sound effects.

Reasoning Before Generation: Mastering Audio with "Chain of Thought"
Conventional audio generation models typically operate on an "intuitive" basis, frequently producing disjointed results—such as a horse’s hoofbeats sounding like bird chirps, or audio lagging behind the visual action. PrismAudio distinguishes itself by employing a "plan first, then generate" approach.
Decomposition Chain of Thought: Prior to generating audio, the model analyzes the video frame by frame: identifying scene elements, determining sound onset timing, assessing audio characteristics (e.g., crispness or depth), and locating the sound source spatially.
Multi-Dimensional Evaluation: To ensure high-quality output, the development team integrated reinforcement learning, utilizing four "virtual evaluators" to score results across semantic consistency, temporal synchronization, aesthetic quality, and spatial accuracy. This multi-faceted feedback loop resolves the common issue where previous models optimized for one aspect while neglecting others.
Lightweight and Efficient: Processing 9-Second Videos in 0.6 Seconds
Beyond accuracy, PrismAudio delivers exceptional speed. Leveraging its proprietary Fast-GRPO training algorithm, the model achieves a substantial performance boost while maintaining operational efficiency:
Compact Architecture, High Impact: With only 518 million parameters, the model is significantly smaller than comparable systems, which often require tens of billions of parameters.
Rapid Response: Generating high-quality audio for a 9-second video clip takes merely 0.63 seconds, approaching real-time performance.
Industry Perspective: The Rise of Authentic Ambient Audio
PrismAudio’s launch not only offers a robust automation tool for film post-production and short-form video creation but also inspires new approaches to multi-object generation. As AI improves its ability to balance audio texture and spatial positioning, future video production will increasingly achieve true "what you see is what you hear" fidelity.
Paper link: arXiv:2511.18833
Open source link: https://prismaudio-project.github.io/
Yushu Technology, A-share's First Humanoid Robot Stock, Sets August 19 Listing With ~61 Billion Yuan Valuation
Scheduled for listing on the Sci-Tech Innovation Board on August 19, Yue Shu Technology will debut at 150.80 yuan per share, establishing a post-issue market capitalization of approximately 60.993 billion yuan. With a price-to-earnings ratio of 219.2
Ali Zhenwu M890 Ultra Node Now Compatible with Qwen3.8, Available on BaiLian Platform for Inference
Alibaba Cloud confirmed on July 23 that its Zhenwu M890 super node has successfully integrated with the latest flagship large model, Qwen3.8, making it available for model inference services on the Alibaba Cloud BaiLian platform. This milestone estab





Home






