Home
DeepMind CEO Demis Hassabis Announces Future Integration of Google's Gemini and Veo AI Models

In a recent episode of the podcast Possible, co-hosted by LinkedIn co-founder Reid Hoffman, Google DeepMind CEO Demis Hassabis shared some exciting news about Google's plans. He revealed that Google is looking to merge its Gemini AI models with the Veo video-generating models. This fusion aims to enhance Gemini's grasp of the physical world, making it more adept at understanding real-life dynamics.
Hassabis emphasized that from the get-go, Gemini was designed to be multimodal. "We've always built Gemini, our foundation model, to be multimodal from the beginning," he explained. The motivation behind this approach? A vision for a universal digital assistant that can truly assist in everyday life. "An assistant that … actually helps you in the real world," Hassabis elaborated.
The AI industry is steadily progressing toward what you might call "omni" models—those capable of handling and synthesizing various types of media. Google's latest Gemini iterations, for instance, can produce not just text but also audio and images. Meanwhile, OpenAI's ChatGPT default model can whip up images on the spot, including delightful Studio Ghibli-style art. Amazon isn't far behind, with plans to roll out an "any-to-any" model later this year.
These omni models demand a hefty amount of training data—think images, videos, audio, and text. Hassabis hinted that Veo's video data primarily comes from YouTube, a treasure trove owned by Google. "Basically, by watching YouTube videos — a lot of YouTube videos — [Veo 2] can figure out, you know, the physics of the world," he noted.
Google had previously mentioned to TechCrunch that its models "may be" trained on "some" YouTube content, aligning with agreements made with YouTube creators. It's worth noting that last year, Google expanded its terms of service, partly to access more data for training its AI models.
Related article
Gemini offers free personalized AI image generation to U.S. users
On Monday, Google revealed that the Gemini app is extending its personalized image generation capabilities, powered by Nano Banana, to a wider audience. As of today, all eligible users in the United States can use this feature at no cost, removing th
New Gemini functionalities set to be integrated into Google TV soon
On Wednesday, Google revealed a series of new AI-driven features set to arrive on Google TV, including a dedicated section for short-form videos that will place YouTube Shorts directly on the home screen.At the heart of these updates are enhanced cap
New Products Unveiled at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and Enhanced Gemini Capabilities
During its Made by Google 2026 event held on Wednesday, Google unveiled a range of new products, including the Pixel 11 series, the Pixel Watch 5, and a tracking device designed to compete with Apple’s AirTag. The company also highlighted several new
Related Special Topic Recommendations
Comments (2)
0/500
The integration of Gemini and Veo sounds promising! Could this be the key to generating truly coherent multimodal content, or are we just stitching together different black boxes? The computational cost for such combined models might be enormous though. A fascinating glimpse into the future roadmap of Google's AI.

In a recent episode of the podcast Possible, co-hosted by LinkedIn co-founder Reid Hoffman, Google DeepMind CEO Demis Hassabis shared some exciting news about Google's plans. He revealed that Google is looking to merge its Gemini AI models with the Veo video-generating models. This fusion aims to enhance Gemini's grasp of the physical world, making it more adept at understanding real-life dynamics.
Hassabis emphasized that from the get-go, Gemini was designed to be multimodal. "We've always built Gemini, our foundation model, to be multimodal from the beginning," he explained. The motivation behind this approach? A vision for a universal digital assistant that can truly assist in everyday life. "An assistant that … actually helps you in the real world," Hassabis elaborated.
The AI industry is steadily progressing toward what you might call "omni" models—those capable of handling and synthesizing various types of media. Google's latest Gemini iterations, for instance, can produce not just text but also audio and images. Meanwhile, OpenAI's ChatGPT default model can whip up images on the spot, including delightful Studio Ghibli-style art. Amazon isn't far behind, with plans to roll out an "any-to-any" model later this year.
These omni models demand a hefty amount of training data—think images, videos, audio, and text. Hassabis hinted that Veo's video data primarily comes from YouTube, a treasure trove owned by Google. "Basically, by watching YouTube videos — a lot of YouTube videos — [Veo 2] can figure out, you know, the physics of the world," he noted.
Google had previously mentioned to TechCrunch that its models "may be" trained on "some" YouTube content, aligning with agreements made with YouTube creators. It's worth noting that last year, Google expanded its terms of service, partly to access more data for training its AI models.
Gemini offers free personalized AI image generation to U.S. users
On Monday, Google revealed that the Gemini app is extending its personalized image generation capabilities, powered by Nano Banana, to a wider audience. As of today, all eligible users in the United States can use this feature at no cost, removing th
New Gemini functionalities set to be integrated into Google TV soon
On Wednesday, Google revealed a series of new AI-driven features set to arrive on Google TV, including a dedicated section for short-form videos that will place YouTube Shorts directly on the home screen.At the heart of these updates are enhanced cap
New Products Unveiled at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and Enhanced Gemini Capabilities
During its Made by Google 2026 event held on Wednesday, Google unveiled a range of new products, including the Pixel 11 series, the Pixel Watch 5, and a tracking device designed to compete with Apple’s AirTag. The company also highlighted several new
The integration of Gemini and Veo sounds promising! Could this be the key to generating truly coherent multimodal content, or are we just stitching together different black boxes? The computational cost for such combined models might be enormous though. A fascinating glimpse into the future roadmap of Google's AI.











