option
Home
Flash News
Content
MatthewLewis
MatthewLewis
July 21, 2026

Google DeepMind launched GenCeption, a model built on Alibaba's Wan2.1 framework that reverses video generation into a visual analysis engine. It performs five core tasks simultaneously, including depth estimation, segmentation, and 3D pose estimation in a single forward pass. Trained on only 7500 single-person synthetic videos, it generalizes to multi-person, animal, and robot scenes. Small model processes 81 frames in 6 seconds; large 14B model takes 10 seconds, preserving fine details like cat whiskers and hair.

Google DeepMind launched GenCeption, a model built on Alibaba's Wan2.1 framework that reverses video generation into a visual analysis engine. It performs five core tasks simultaneously, including depth estimation, segmentation, and 3D pose estimation in a single forward pass. Trained on only 7500 single-person synthetic videos, it generalizes to multi-person, animal, and robot scenes. Small model processes 81 frames in 6 seconds; large 14B model takes 10 seconds, preserving fine details like cat whiskers and hair.
Comments (0)
0/300
OR