Home
Ali’s Black Tech: 0.6B Model Upgraded to 17B MoE with 5% Active Params, Running on CPU at 30 Tokens per Second
The Ali International Digital Commerce team has unveiled Marco-Mini-Instruct, the latest addition to the Marco-MoE series, showcasing how "smaller scales can deliver big results." With 17.3 billion total parameters, only 0.86 billion are active (roughly 5%), enabling exceptional inference speed that even runs smoothly on standard CPUs.

Ultra-Lightweight: Optimized for CPU Performance
According to official data, using 8-bit quantization with four DDR4 2400 memory modules, the model achieves an inference speed of approximately 30 tokens per second. This efficiency brings Mixture-of-Experts (MoE) architecture closer to mainstream accessibility, significantly reducing the barriers to local deployment.
Key Innovation: Upcycling Technology Transforms Potential
The standout feature of Marco-Mini-Instruct is not its size or speed, but its unique creation method. Instead of being trained from scratch, this model was evolved from the Qwen3-0.6B-Base model using upcycling technology.

This process involves splitting or copying components of a dense small model into multiple experts, introducing a routing mechanism. It also integrates fine-grained sub-matrix partitioning and Drop-Upcycling strategies—randomly discarding certain experts or routing paths during training to add regularization and improve robustness. This approach successfully upgrades a pure Dense model into an MoE architecture, offering the industry a new, cost-effective, and efficient path for MoE training.
Context Length and Training Setup
The model configuration expands max_position_embeddings to 32K, though the Supervised Fine-Tuning (SFT) phase utilizes an 8192-token context. Consequently, the default context length is well-suited for most practical application scenarios.
Post-Training Breakthrough: Cascaded On-Policy Distillation
The post-training phase is equally impressive: it begins with SFT preheating, followed by a cascaded On-Policy Distillation strategy. Initially, distillation uses Qwen3-30B-A3B-Instruct as the teacher model, then switches to the more powerful Qwen3-Next-80B-A3B-Instruct. The distillation data spans instruction following, complex reasoning, alignment security, and mathematical capabilities, ensuring the model maintains efficiency while significantly boosting overall intelligence.
Benchmark Results: 0.86B Active Parameters Beat 4B Dense Models
The final Marco-Mini-Instruct outperforms many dense models, such as Qwen3-4B, across most mainstream benchmarks. With only 0.86 billion active parameters, it fully validates the vast potential of MoE architecture in delivering "small yet powerful" performance.
Industry Impact: A New Open-Source MoE Training Paradigm
AIbase believes the greatest value of this achievement lies in opening new opportunities for developers. There is no longer a need to train large-scale MoE models from scratch; instead, teams can select a suitable small Dense model and strictly reproduce the upcycling and Drop-Upcycling processes outlined in the paper. The entire training cost remains manageable: the SFT phase requires 64 GPUs for 24 hours, and the distillation phase needs 64 GPUs for 110 hours, greatly lowering the threshold for small and medium-sized teams to experiment with MoE.
Alibaba's latest "modification" once again proves that breakthroughs in model efficiency do not necessarily rely on stacking parameters; innovative training paradigms can also drive qualitative leaps. The release of Marco-Mini-Instruct will undoubtedly accelerate the adoption of MoE technology in edge devices and personal developer scenarios, making it a key development for the entire industry to watch.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
The Ali International Digital Commerce team has unveiled Marco-Mini-Instruct, the latest addition to the Marco-MoE series, showcasing how "smaller scales can deliver big results." With 17.3 billion total parameters, only 0.86 billion are active (roughly 5%), enabling exceptional inference speed that even runs smoothly on standard CPUs.

Ultra-Lightweight: Optimized for CPU Performance
According to official data, using 8-bit quantization with four DDR4 2400 memory modules, the model achieves an inference speed of approximately 30 tokens per second. This efficiency brings Mixture-of-Experts (MoE) architecture closer to mainstream accessibility, significantly reducing the barriers to local deployment.
Key Innovation: Upcycling Technology Transforms Potential
The standout feature of Marco-Mini-Instruct is not its size or speed, but its unique creation method. Instead of being trained from scratch, this model was evolved from the Qwen3-0.6B-Base model using upcycling technology.

This process involves splitting or copying components of a dense small model into multiple experts, introducing a routing mechanism. It also integrates fine-grained sub-matrix partitioning and Drop-Upcycling strategies—randomly discarding certain experts or routing paths during training to add regularization and improve robustness. This approach successfully upgrades a pure Dense model into an MoE architecture, offering the industry a new, cost-effective, and efficient path for MoE training.
Context Length and Training Setup
The model configuration expands max_position_embeddings to 32K, though the Supervised Fine-Tuning (SFT) phase utilizes an 8192-token context. Consequently, the default context length is well-suited for most practical application scenarios.
Post-Training Breakthrough: Cascaded On-Policy Distillation
The post-training phase is equally impressive: it begins with SFT preheating, followed by a cascaded On-Policy Distillation strategy. Initially, distillation uses Qwen3-30B-A3B-Instruct as the teacher model, then switches to the more powerful Qwen3-Next-80B-A3B-Instruct. The distillation data spans instruction following, complex reasoning, alignment security, and mathematical capabilities, ensuring the model maintains efficiency while significantly boosting overall intelligence.
Benchmark Results: 0.86B Active Parameters Beat 4B Dense Models
The final Marco-Mini-Instruct outperforms many dense models, such as Qwen3-4B, across most mainstream benchmarks. With only 0.86 billion active parameters, it fully validates the vast potential of MoE architecture in delivering "small yet powerful" performance.
Industry Impact: A New Open-Source MoE Training Paradigm
AIbase believes the greatest value of this achievement lies in opening new opportunities for developers. There is no longer a need to train large-scale MoE models from scratch; instead, teams can select a suitable small Dense model and strictly reproduce the upcycling and Drop-Upcycling processes outlined in the paper. The entire training cost remains manageable: the SFT phase requires 64 GPUs for 24 hours, and the distillation phase needs 64 GPUs for 110 hours, greatly lowering the threshold for small and medium-sized teams to experiment with MoE.
Alibaba's latest "modification" once again proves that breakthroughs in model efficiency do not necessarily rely on stacking parameters; innovative training paradigms can also drive qualitative leaps. The release of Marco-Mini-Instruct will undoubtedly accelerate the adoption of MoE technology in edge devices and personal developer scenarios, making it a key development for the entire industry to watch.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











