Home
MiniCPM-V 4.6's 1.3B Model Dimension Reduction Attack Redefines Edge Multimodal Technology
On May 11, Mingshi Intelligence teamed up with Tsinghua University and the OpenBMB open-source community to unveil the new edge-side multimodal large model MiniCPM-V4.6. Despite its compact size—just 1.3 billion parameters—this lightweight model pushes the boundaries of larger models through exceptional intelligent density and cross-platform adaptability, speeding up the real-world deployment of edge-side AI.

1. Peak Performance: 1.3B Parameters Deliver Exceptional Results
MiniCPM-V4.6 comes in two versions—Instruct and Thinking—both demonstrating impressive reasoning and understanding across various benchmarks when compared to similarly sized models.
Global Leadership: On the Artificial Analysis (AA) leaderboard, MiniCPM-V4.6 scored an impressive 13 points, far surpassing similarly sized rivals (e.g., Alibaba's Qwen3.5-0.8B and Google's Gemma4-E2B-it) while nearly matching the performance of larger models like Qwen3.5-2B. This sets a new benchmark for 1B-parameter models.
Advanced Capabilities: From general image-text comprehension and complex STEM math reasoning to demanding document OCR and video temporal understanding, the model exhibits strong intelligence. The Thinking version, in particular, excels in multi-image reasoning and hallucination suppression.
2. Efficiency Breakthrough: Extreme Intelligent Density at the Edge
To tackle memory constraints in edge deployment, MiniCPM-V4.6 has been deeply optimized for inference speed and resource efficiency.
Low Memory Footprint: Memory requirements drop to just 6GB, enabling smooth operation on mainstream smartphones, PCs, and smart home devices.
Inference Efficiency: Powered by vLLM, inference throughput is 1.5× that of competitors. When processing a 3136² ultra-high-definition image at the edge, the first response latency is just 75.7ms—2.2× faster than competing models.
Throughput: A single GPU delivers text generation at 7,013 tokens per second and can process 1,344² images at 54.79 images per second, demonstrating remarkable efficiency.
3. Core Technology: LLaVA-UHD v4 Minimizes Overhead
The model's lightweight design is made possible by LLaVA-UHD v4, jointly developed by Mingshi Intelligence and Tsinghua University.
Encoding Redesign: Through a revamped ViT image encoding and shallow compression module, image encoding overhead is cut by 50% and high-resolution floating-point operations by 55.8%.
Hybrid Compression: It introduces a 4×/16× hybrid token compression that enables flexible switching between performance-first and speed-first modes. This technology has been validated in Kuaishou's recommendation model OneRec, handling massive traffic.
4. Ecosystem Deployment: From Lab to Industry
The open-source release of MiniCPM-V4.6 marks both a technical and ecosystem milestone.
Easy Development: It integrates seamlessly with fine-tuning frameworks like ms-swift and LLaMA-Factory, enabling full-scale fine-tuning on a single RTX 4090 GPU.
Cross-Platform Support: Compatible with vLLM, Ollama, and other major frameworks, it offers test versions for iOS, Android, and HarmonyOS, bringing AI to a wider range of hardware.
Real-World Impact: The model series is already deployed in automotive, PC, smart home, and industrial inspection sectors, with partners including Lenovo, Geely, SAIC Volkswagen, Xiaomi, and OPPO.
Related article
Li Feifei’s Team Unveils Method to Generate Infinite Robot Training Grounds From Single Video
In embodied intelligence, bridging the gap from simulation to reality—known as "Sim2Real"—has long been a persistent hurdle. Simulation setups are expensive, and models frequently lose performance when deployed in physical settings. A recent collabor
Adobe Creator Report: 80% Say AI Drives Fan and Business Growth
As artificial intelligence accelerates its global expansion, the content creation sector is experiencing a fundamental shift. Creative AI tools are no longer just workflow enhancers; they are driving business growth and helping creators build larger
Google declares itself a major player in AI design
At its annual I/O conference on Tuesday, Google unveiled Pics, a new AI-driven design and image creation tool for Google Workspace. The company states the app is built for accessibility, serving everyone from educators to small business owners.Pics a
Related Special Topic Recommendations
Comments (0)
0/500
On May 11, Mingshi Intelligence teamed up with Tsinghua University and the OpenBMB open-source community to unveil the new edge-side multimodal large model MiniCPM-V4.6. Despite its compact size—just 1.3 billion parameters—this lightweight model pushes the boundaries of larger models through exceptional intelligent density and cross-platform adaptability, speeding up the real-world deployment of edge-side AI.

1. Peak Performance: 1.3B Parameters Deliver Exceptional Results
MiniCPM-V4.6 comes in two versions—Instruct and Thinking—both demonstrating impressive reasoning and understanding across various benchmarks when compared to similarly sized models.
Global Leadership: On the Artificial Analysis (AA) leaderboard, MiniCPM-V4.6 scored an impressive 13 points, far surpassing similarly sized rivals (e.g., Alibaba's Qwen3.5-0.8B and Google's Gemma4-E2B-it) while nearly matching the performance of larger models like Qwen3.5-2B. This sets a new benchmark for 1B-parameter models.
Advanced Capabilities: From general image-text comprehension and complex STEM math reasoning to demanding document OCR and video temporal understanding, the model exhibits strong intelligence. The Thinking version, in particular, excels in multi-image reasoning and hallucination suppression.
2. Efficiency Breakthrough: Extreme Intelligent Density at the Edge
To tackle memory constraints in edge deployment, MiniCPM-V4.6 has been deeply optimized for inference speed and resource efficiency.
Low Memory Footprint: Memory requirements drop to just 6GB, enabling smooth operation on mainstream smartphones, PCs, and smart home devices.
Inference Efficiency: Powered by vLLM, inference throughput is 1.5× that of competitors. When processing a 3136² ultra-high-definition image at the edge, the first response latency is just 75.7ms—2.2× faster than competing models.
Throughput: A single GPU delivers text generation at 7,013 tokens per second and can process 1,344² images at 54.79 images per second, demonstrating remarkable efficiency.
3. Core Technology: LLaVA-UHD v4 Minimizes Overhead
The model's lightweight design is made possible by LLaVA-UHD v4, jointly developed by Mingshi Intelligence and Tsinghua University.
Encoding Redesign: Through a revamped ViT image encoding and shallow compression module, image encoding overhead is cut by 50% and high-resolution floating-point operations by 55.8%.
Hybrid Compression: It introduces a 4×/16× hybrid token compression that enables flexible switching between performance-first and speed-first modes. This technology has been validated in Kuaishou's recommendation model OneRec, handling massive traffic.
4. Ecosystem Deployment: From Lab to Industry
The open-source release of MiniCPM-V4.6 marks both a technical and ecosystem milestone.
Easy Development: It integrates seamlessly with fine-tuning frameworks like ms-swift and LLaMA-Factory, enabling full-scale fine-tuning on a single RTX 4090 GPU.
Cross-Platform Support: Compatible with vLLM, Ollama, and other major frameworks, it offers test versions for iOS, Android, and HarmonyOS, bringing AI to a wider range of hardware.
Real-World Impact: The model series is already deployed in automotive, PC, smart home, and industrial inspection sectors, with partners including Lenovo, Geely, SAIC Volkswagen, Xiaomi, and OPPO.
Li Feifei’s Team Unveils Method to Generate Infinite Robot Training Grounds From Single Video
In embodied intelligence, bridging the gap from simulation to reality—known as "Sim2Real"—has long been a persistent hurdle. Simulation setups are expensive, and models frequently lose performance when deployed in physical settings. A recent collabor
Adobe Creator Report: 80% Say AI Drives Fan and Business Growth
As artificial intelligence accelerates its global expansion, the content creation sector is experiencing a fundamental shift. Creative AI tools are no longer just workflow enhancers; they are driving business growth and helping creators build larger
Google declares itself a major player in AI design
At its annual I/O conference on Tuesday, Google unveiled Pics, a new AI-driven design and image creation tool for Google Workspace. The company states the app is built for accessibility, serving everyone from educators to small business owners.Pics a











