Home
Cosmos3, First Fully Open-Source Multimodal Physical AI Model, Debuts; NVIDIA Co-Founds Cosmos Alliance
Significant progress has been made in physical artificial intelligence. On June 1, NVIDIA officially launched Cosmos3, an open-world foundational large model for physical AI. As the first fully open-source and multimodal physical AI large model in the world, it uses an innovative hybrid Transformer architecture that combines visual reasoning, world generation, and action prediction into a single system. This could dramatically shorten the physical AI training and evaluation cycle from months to just days.
Cosmos3 offers a fresh solution to the long-standing industry challenge of "difficulty generalizing in real-world scenarios with limited data and fragmented simulation frameworks." The model was trained on a massive physical AI dataset containing billions of text, images, videos, audio, and motion trajectories. It can naturally understand and generate cross-modal content, achieving industry-leading accuracy in physical simulation.

On the technical side, Cosmos3 innovatively combines a reasoning Transformer with a generative Transformer. The model first deeply analyzes object interaction rules, motion states, and spatiotemporal relationships, then accurately generates video and predicts action trajectories. This design gives it strong multimodal understanding of images and text, the ability to simulate and predict physical environments, and action strategy capabilities to help robots complete specific tasks. In several mainstream physical AI benchmarks, including Artificial Analysis, Physics-IQ, and RoboLab, Cosmos3 ranks among the top open-source models.
To accommodate different development stages, NVIDIA has released multiple versions: Cosmos3Super, focused on secondary training of robot and autonomous driving models with extreme precision, and Cosmos3Nano, which can perform high-quality video parsing and action reasoning in seconds. Both are now officially available. The Cosmos3Edge version, designed for real-time inference on edge devices, is also on the release roadmap.
Alongside the large model launch, NVIDIA co-founded the "NVIDIA Cosmos Coalition" with leading world model research teams and AI developers worldwide, including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI. NVIDIA founder and CEO Jensen Huang stated that with continuous breakthroughs in multimodal reasoning and world models, the transformative era of physical AI has arrived. The release of these cutting-edge open-source models will help developers around the world achieve technological leaps and create the next generation of intelligent systems capable of perceiving, reasoning, and acting in the real world.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
Significant progress has been made in physical artificial intelligence. On June 1, NVIDIA officially launched Cosmos3, an open-world foundational large model for physical AI. As the first fully open-source and multimodal physical AI large model in the world, it uses an innovative hybrid Transformer architecture that combines visual reasoning, world generation, and action prediction into a single system. This could dramatically shorten the physical AI training and evaluation cycle from months to just days.
Cosmos3 offers a fresh solution to the long-standing industry challenge of "difficulty generalizing in real-world scenarios with limited data and fragmented simulation frameworks." The model was trained on a massive physical AI dataset containing billions of text, images, videos, audio, and motion trajectories. It can naturally understand and generate cross-modal content, achieving industry-leading accuracy in physical simulation.

On the technical side, Cosmos3 innovatively combines a reasoning Transformer with a generative Transformer. The model first deeply analyzes object interaction rules, motion states, and spatiotemporal relationships, then accurately generates video and predicts action trajectories. This design gives it strong multimodal understanding of images and text, the ability to simulate and predict physical environments, and action strategy capabilities to help robots complete specific tasks. In several mainstream physical AI benchmarks, including Artificial Analysis, Physics-IQ, and RoboLab, Cosmos3 ranks among the top open-source models.
To accommodate different development stages, NVIDIA has released multiple versions: Cosmos3Super, focused on secondary training of robot and autonomous driving models with extreme precision, and Cosmos3Nano, which can perform high-quality video parsing and action reasoning in seconds. Both are now officially available. The Cosmos3Edge version, designed for real-time inference on edge devices, is also on the release roadmap.
Alongside the large model launch, NVIDIA co-founded the "NVIDIA Cosmos Coalition" with leading world model research teams and AI developers worldwide, including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI. NVIDIA founder and CEO Jensen Huang stated that with continuous breakthroughs in multimodal reasoning and world models, the transformative era of physical AI has arrived. The release of these cutting-edge open-source models will help developers around the world achieve technological leaps and create the next generation of intelligent systems capable of perceiving, reasoning, and acting in the real world.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











