Home
Tencent Hunyuan integrates critical point size theory into PPO reinforcement learning, boosting throughput 2.29x and cutting GRPO training time by 29 percent

As reinforcement learning for large models scales up to larger GPU clusters and denser training datasets, training efficiency has emerged as a critical priority. Recent research from the Tencent Huan Yuan team addresses a frequently overlooked challenge: when models generate their own training data, how should batch sizes be redefined given the differing paces of rollout generation and training?
Building on the classic theory of critical batch size, the study re-derives this concept for online large language model reinforcement learning. The findings offer a practical solution: by adjusting the learning rate within a specific range as batch size increases, both GRPO and PPO algorithms can ensure that each response learns effectively without being diluted. This indicates that batch size is not simply "the larger, the better" nor a fixed value; rather, it exists within an optimal range that can be fine-tuned for maximum efficiency.
The impact on real-world hardware costs is substantial. With fixed hardware configurations, increasing the batch size can boost PPO generation throughput by up to 2.29 times. Meanwhile, the optimal GRPO configuration achieves the same validation targets in 29% less time. For teams heavily invested in GPU resources, this translates to either completing tasks faster with the same cluster or processing more data within the same timeframe, effectively reducing costs in the most expensive phase of reinforcement learning.
This research provides a practical lever for scaling online RL, where the model serves as both student and question setter. The efficiency gap arises from the mismatch between generation and training scales, which this study bridges by combining critical batch size theory with learning rate adjustments. While the industry competes to build the largest models, Tencent Huan Yuan is focusing on optimizing existing hardware for greater cost-effectiveness.
Related article
AI startups accelerate revenue growth
As established firms and emerging ventures scramble to leverage artificial intelligence, numerous AI startups report that their revenue is not merely expanding, but accelerating rapidly, achieving subsequent milestones in increasingly shorter periods
Replit’s Amjad Masad on the Cursor deal, fighting Apple, and why he’d rather not sell
Amjad Masad has spent the last decade building Replit, yet the past 18 months have been a transformative period. The AI coding platform has surged from $2.8 million in annual revenue in 2024 to a trajectory toward what Masad describes as a billion-do
AI Agents Boost Shrimp Farming as MIIT Issues Warning and Rockchip Responds
The conversation in the AI community has shifted from "training large models" to "raising lobsters." If you see someone on social media discussing how to feed data and tune parameters for a "lobster," don't be confused—they're not farming aquatic cre
Related Special Topic Recommendations
Comments (0)
0/500

As reinforcement learning for large models scales up to larger GPU clusters and denser training datasets, training efficiency has emerged as a critical priority. Recent research from the Tencent Huan Yuan team addresses a frequently overlooked challenge: when models generate their own training data, how should batch sizes be redefined given the differing paces of rollout generation and training?
Building on the classic theory of critical batch size, the study re-derives this concept for online large language model reinforcement learning. The findings offer a practical solution: by adjusting the learning rate within a specific range as batch size increases, both GRPO and PPO algorithms can ensure that each response learns effectively without being diluted. This indicates that batch size is not simply "the larger, the better" nor a fixed value; rather, it exists within an optimal range that can be fine-tuned for maximum efficiency.
The impact on real-world hardware costs is substantial. With fixed hardware configurations, increasing the batch size can boost PPO generation throughput by up to 2.29 times. Meanwhile, the optimal GRPO configuration achieves the same validation targets in 29% less time. For teams heavily invested in GPU resources, this translates to either completing tasks faster with the same cluster or processing more data within the same timeframe, effectively reducing costs in the most expensive phase of reinforcement learning.
This research provides a practical lever for scaling online RL, where the model serves as both student and question setter. The efficiency gap arises from the mismatch between generation and training scales, which this study bridges by combining critical batch size theory with learning rate adjustments. While the industry competes to build the largest models, Tencent Huan Yuan is focusing on optimizing existing hardware for greater cost-effectiveness.
AI startups accelerate revenue growth
As established firms and emerging ventures scramble to leverage artificial intelligence, numerous AI startups report that their revenue is not merely expanding, but accelerating rapidly, achieving subsequent milestones in increasingly shorter periods
Replit’s Amjad Masad on the Cursor deal, fighting Apple, and why he’d rather not sell
Amjad Masad has spent the last decade building Replit, yet the past 18 months have been a transformative period. The AI coding platform has surged from $2.8 million in annual revenue in 2024 to a trajectory toward what Masad describes as a billion-do
AI Agents Boost Shrimp Farming as MIIT Issues Warning and Rockchip Responds
The conversation in the AI community has shifted from "training large models" to "raising lobsters." If you see someone on social media discussing how to feed data and tune parameters for a "lobster," don't be confused—they're not farming aquatic cre











