ByteDance Open Sources Bernini Framework for Unified Video Generation and Editing
ByteDance's commercialization technology team has officially released an open-source video generation and editing framework named Bernini. At its core, the framework adopts a "understand first, then generate" collaborative mechanism, designed to address common industry challenges like image instability and frame flickering that arise from traditional models failing to interpret complex instructions accurately.
In internal evaluations at ByteDance, Bernini has achieved top-tier results. Its inference code and the second-stage model, Bernini-R, are now publicly available, with the fully featured version scheduled for a complete open-source release in the near future.

Decoupling Semantics from Rendering
Bernini's workflow introduces a novel separation between "semantic planning" and "visual rendering." The process begins with a multimodal large model planner that thoroughly analyzes input materials to produce a "semantic sketch," which is then translated into stable, coherent video frames by the renderer.
This clear division of responsibilities makes the framework particularly useful for controllable editing. Users can adjust scene attributes such as weather, season, or visual style through simple commands, and also gain precise control over camera angles, focus, and subject movements.
Rich Visual Reference Capabilities
Beyond traditional text-based control, Bernini supports using images and videos as visual references, significantly improving creative consistency. In video editing, it can accurately place specific materials or posters into target regions without boundary artifacts or perspective distortion.
For generating new videos, the model handles single-image and multi-angle reference inputs, and can evolve keyframes into continuous sequences. To avoid confusion when connecting multiple visual segments, the team introduced a dedicated positional encoding mechanism that clearly distinguishes reference materials from output targets.
Project page: https://bernini-ai.github.io/
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
ByteDance's commercialization technology team has officially released an open-source video generation and editing framework named Bernini. At its core, the framework adopts a "understand first, then generate" collaborative mechanism, designed to address common industry challenges like image instability and frame flickering that arise from traditional models failing to interpret complex instructions accurately.
In internal evaluations at ByteDance, Bernini has achieved top-tier results. Its inference code and the second-stage model, Bernini-R, are now publicly available, with the fully featured version scheduled for a complete open-source release in the near future.

Decoupling Semantics from Rendering
Bernini's workflow introduces a novel separation between "semantic planning" and "visual rendering." The process begins with a multimodal large model planner that thoroughly analyzes input materials to produce a "semantic sketch," which is then translated into stable, coherent video frames by the renderer.
This clear division of responsibilities makes the framework particularly useful for controllable editing. Users can adjust scene attributes such as weather, season, or visual style through simple commands, and also gain precise control over camera angles, focus, and subject movements.
Rich Visual Reference Capabilities
Beyond traditional text-based control, Bernini supports using images and videos as visual references, significantly improving creative consistency. In video editing, it can accurately place specific materials or posters into target regions without boundary artifacts or perspective distortion.
For generating new videos, the model handles single-image and multi-angle reference inputs, and can evolve keyframes into continuous sequences. To avoid confusion when connecting multiple visual segments, the team introduced a dedicated positional encoding mechanism that clearly distinguishes reference materials from output targets.
Project page: https://bernini-ai.github.io/
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






