Tencent reshuffles large model team as Yao Shunyu takes charge of basic model

Tencent’s Hunyuan multimodal team has recently completed a significant restructuring and strategic pivot. Lu Xudong, formerly leading xAI’s multimodal understanding efforts, has joined Hunyuan to head multimodal content generation algorithms. This shift coincides with key personnel changes, including the departure of former multimodal understanding head Hu Han to pursue entrepreneurship and the arrival of Tian Yonglong, an ex-OpenAI researcher. These moves signal a major realignment in Tencent’s approach to developing multimodal large models.
Tencent Hunyuan’s multimodal evolution has centered on several core products. In late 2024, the team launched HunyuanVideo, the largest open-source text-to-video model, followed by 2025 expansions into image-to-video, custom generation, and digital human animation. Image generation capabilities have advanced from HunyuanImage to version 2.1, which supports native 2K resolution, and version 3.0, featuring 80 billion parameters, marking a shift toward industrial-grade performance. Meanwhile, Hy3D, a unified 3D model, is widely used in e-commerce and product design. In spatial intelligence, Tencent introduced Hy World 1.0, the world’s first open-source, simulation-ready immersive 3D generation model, with subsequent versions continuing to advance spatial intelligence research.
As Tencent consolidated its large language model and multimodal model divisions into the “Basic Model Department” under Yao Shunyu’s leadership, its R&D strategy underwent a fundamental shift. The current technical roadmap reflects industry debate over the role of multimodal systems in AI’s main trajectory. Fei-Fei Li’s World Labs argues that world models, as components of spatial intelligence, are central to achieving AGI. In contrast, DeepSeek advocates using visual modalities as tools to support language models, de-emphasizing standalone 3D or world model development.
Yao Shunyu’s recent strategic initiatives, including the flagship product Hy3, suggest a preference for the latter approach. He has consistently stated that AI’s next competitive edge lies in context handling and reasoning, not merely parameter count. By integrating multimodal understanding into the core model architecture and improving tool-calling stability and long-context support, Tencent aims to boost model efficiency in real-world office environments and GUI Agent applications. Lu Xudong’s addition strengthens multimodal understanding, addressing common issues like generation inconsistency and spatial instability. This restructuring highlights Tencent’s internal strategic evolution and its commitment to focusing on productivity and core capabilities in the next phase of AGI development.
Related article
Trump’s AI Strategy: What’s Next for US Policy
On August 3, five Democratic senators urged the Trump administration to clarify its oversight of frontier AI models, highlighting concerns regarding the competitive threat posed by Chinese AI systems. Credit: GettyFollowing recent cyber incidents, th
Tongyi Qianwen Open Platform Launches, Bringing Conversational-as-a-Service to Daily Life
The Tongyi Qianwen Open Platform has officially launched, providing developers with access to services across mobile devices, PCs, and AI glasses. Users no longer need to switch between apps; they can complete various daily service operations directl
Swiftlet Packs 80B Qwen Into Mac With 4.3GB Memory; iPhone 17 to Run 35B Natively
A Swift + Metal runtime named Swiftlet is redefining "where large models can run." It is specifically tailored for the Qwen3-Next and Qwen3.5/3.6 family of MoE (Mixture of Experts) models. The core idea is quite clever: only the small dense core of t
Related Special Topic Recommendations
Comments (0)
0/500

Tencent’s Hunyuan multimodal team has recently completed a significant restructuring and strategic pivot. Lu Xudong, formerly leading xAI’s multimodal understanding efforts, has joined Hunyuan to head multimodal content generation algorithms. This shift coincides with key personnel changes, including the departure of former multimodal understanding head Hu Han to pursue entrepreneurship and the arrival of Tian Yonglong, an ex-OpenAI researcher. These moves signal a major realignment in Tencent’s approach to developing multimodal large models.
Tencent Hunyuan’s multimodal evolution has centered on several core products. In late 2024, the team launched HunyuanVideo, the largest open-source text-to-video model, followed by 2025 expansions into image-to-video, custom generation, and digital human animation. Image generation capabilities have advanced from HunyuanImage to version 2.1, which supports native 2K resolution, and version 3.0, featuring 80 billion parameters, marking a shift toward industrial-grade performance. Meanwhile, Hy3D, a unified 3D model, is widely used in e-commerce and product design. In spatial intelligence, Tencent introduced Hy World 1.0, the world’s first open-source, simulation-ready immersive 3D generation model, with subsequent versions continuing to advance spatial intelligence research.
As Tencent consolidated its large language model and multimodal model divisions into the “Basic Model Department” under Yao Shunyu’s leadership, its R&D strategy underwent a fundamental shift. The current technical roadmap reflects industry debate over the role of multimodal systems in AI’s main trajectory. Fei-Fei Li’s World Labs argues that world models, as components of spatial intelligence, are central to achieving AGI. In contrast, DeepSeek advocates using visual modalities as tools to support language models, de-emphasizing standalone 3D or world model development.
Yao Shunyu’s recent strategic initiatives, including the flagship product Hy3, suggest a preference for the latter approach. He has consistently stated that AI’s next competitive edge lies in context handling and reasoning, not merely parameter count. By integrating multimodal understanding into the core model architecture and improving tool-calling stability and long-context support, Tencent aims to boost model efficiency in real-world office environments and GUI Agent applications. Lu Xudong’s addition strengthens multimodal understanding, addressing common issues like generation inconsistency and spatial instability. This restructuring highlights Tencent’s internal strategic evolution and its commitment to focusing on productivity and core capabilities in the next phase of AGI development.
Trump’s AI Strategy: What’s Next for US Policy
On August 3, five Democratic senators urged the Trump administration to clarify its oversight of frontier AI models, highlighting concerns regarding the competitive threat posed by Chinese AI systems. Credit: GettyFollowing recent cyber incidents, th
Tongyi Qianwen Open Platform Launches, Bringing Conversational-as-a-Service to Daily Life
The Tongyi Qianwen Open Platform has officially launched, providing developers with access to services across mobile devices, PCs, and AI glasses. Users no longer need to switch between apps; they can complete various daily service operations directl
Swiftlet Packs 80B Qwen Into Mac With 4.3GB Memory; iPhone 17 to Run 35B Natively
A Swift + Metal runtime named Swiftlet is redefining "where large models can run." It is specifically tailored for the Qwen3-Next and Qwen3.5/3.6 family of MoE (Mixture of Experts) models. The core idea is quite clever: only the small dense core of t





Home






