option
Home
Flash News
Content
RoyLopez
RoyLopez
September 17, 2026

Xiaomi MiMo team publicly reveals real-time reinforcement learning progress for its new large model MiMo-V2.6 via a live training page. Led by former DeepSeek researcher Luo Fuli, the project features large-scale multi-task agent experiments with fully asynchronous parallelism. The training cost for two versions exceeded $1.28 million, taking roughly 30 hours combined. This transparent approach marks a shift from black-box trials to open-source engineering, highlighting Xiaomi's advancements in RL scaling laws and AGI exploration.

Comments (0)
0/300
OR