Home
DeepSeek V4.1 Flash Launches With Native Multimodal Vision, Outperforming Pro Version as Prices Fall
DeepSeek has officially launched the smallest model in its latest architecture series today: DeepSeek V4.1Flash. Designed as a lightweight model with native multi-modal visual understanding, it aims to deliver higher performance, faster inference, greater throughput, and scalability to larger parameter sizes.
At its core, DeepSeek V4.1Flash utilizes a 552B parameter Mixture of Experts (MoE) architecture, featuring an innovative Causal-Encoder-Decoder structure. Its asymmetric design results in only 8B input activations and 16B output activations, significantly lowering costs compared to similarly sized models. Enhanced by a new pre-training approach and large-scale reinforcement learning, this model outperforms flagship models like DeepSeek V4Pro across various benchmarks.

Regarding efficiency and cost control, the new model drastically reduces KV Cache size. Compared to its predecessor, V4.1Flash requires only 1/4 of the high-bandwidth memory (HBM) and 1/8 of the solid-state drive (SSD) capacity. This compression is particularly beneficial for Agent scenarios, where context storage and cache hit costs are high, significantly lowering operational expenses. Statistics show the KV Cache has shrunk to 1/437 of the initial model's size.
With the new model's release, DeepSeek has updated its product lineup and billing strategy. DeepSeek V4.1Flash is now available via the DeepSeek API under the model name deepseek-flash. Previous versions, V4Flash and V4Flash Vision Exp, have been discontinued, with old names (deepseek-v4-flash and deepseek-v4-flash-vision-exp) temporarily routing to V4.1Flash. Since V4.1Flash surpasses V4Pro in performance, cost, speed, and usage time, the company plans to phase out V4Pro: after 12:00 on September 14, 2026, requests to deepseek-v4-pro will be routed to V4.1Flash and billed at its rate. Official partners like WorkBuddy, CodeBuddy, and Tencent's OpenCode have already integrated this model.
Related article
Aliyun Qwen Open Platform Launches, Enabling AI on Phones, PCs, and Glasses Across Renting, Shipping, and More
As AI applications expand rapidly, Qwen has significantly upgraded its ecosystem. Launched on August 10, the Qwen Open Platform now provides broad access for mobile phones, PCs, and AI glasses to ecosystem partners and developers.In its initial phase
Dou Bao Set to Launch Paid Tiers in Mid-June, E-Commerce Push Planned for Q3
Douyin, ByteDance’s flagship large model application, is set to roll out paid content features by late June, with updates scheduled for the Force conference. This move marks a major milestone in monetizing China’s largest large model.Reports indicate
CCTV Exposes AI Dating Scam Targeting Middle-Aged and Older Women, 28 Arrested
Recent reports from CCTV News highlight a new wave of AI-driven dating scams targeting middle-aged and elderly women. Criminals are leveraging AI-generated personas to lure victims into financial fraud. Authorities in Chengdu, Sichuan, recently buste
Related Special Topic Recommendations
Comments (0)
0/500
DeepSeek has officially launched the smallest model in its latest architecture series today: DeepSeek V4.1Flash. Designed as a lightweight model with native multi-modal visual understanding, it aims to deliver higher performance, faster inference, greater throughput, and scalability to larger parameter sizes.
At its core, DeepSeek V4.1Flash utilizes a 552B parameter Mixture of Experts (MoE) architecture, featuring an innovative Causal-Encoder-Decoder structure. Its asymmetric design results in only 8B input activations and 16B output activations, significantly lowering costs compared to similarly sized models. Enhanced by a new pre-training approach and large-scale reinforcement learning, this model outperforms flagship models like DeepSeek V4Pro across various benchmarks.

Regarding efficiency and cost control, the new model drastically reduces KV Cache size. Compared to its predecessor, V4.1Flash requires only 1/4 of the high-bandwidth memory (HBM) and 1/8 of the solid-state drive (SSD) capacity. This compression is particularly beneficial for Agent scenarios, where context storage and cache hit costs are high, significantly lowering operational expenses. Statistics show the KV Cache has shrunk to 1/437 of the initial model's size.
With the new model's release, DeepSeek has updated its product lineup and billing strategy. DeepSeek V4.1Flash is now available via the DeepSeek API under the model name deepseek-flash. Previous versions, V4Flash and V4Flash Vision Exp, have been discontinued, with old names (deepseek-v4-flash and deepseek-v4-flash-vision-exp) temporarily routing to V4.1Flash. Since V4.1Flash surpasses V4Pro in performance, cost, speed, and usage time, the company plans to phase out V4Pro: after 12:00 on September 14, 2026, requests to deepseek-v4-pro will be routed to V4.1Flash and billed at its rate. Official partners like WorkBuddy, CodeBuddy, and Tencent's OpenCode have already integrated this model.
Aliyun Qwen Open Platform Launches, Enabling AI on Phones, PCs, and Glasses Across Renting, Shipping, and More
As AI applications expand rapidly, Qwen has significantly upgraded its ecosystem. Launched on August 10, the Qwen Open Platform now provides broad access for mobile phones, PCs, and AI glasses to ecosystem partners and developers.In its initial phase
Dou Bao Set to Launch Paid Tiers in Mid-June, E-Commerce Push Planned for Q3
Douyin, ByteDance’s flagship large model application, is set to roll out paid content features by late June, with updates scheduled for the Force conference. This move marks a major milestone in monetizing China’s largest large model.Reports indicate
CCTV Exposes AI Dating Scam Targeting Middle-Aged and Older Women, 28 Arrested
Recent reports from CCTV News highlight a new wave of AI-driven dating scams targeting middle-aged and elderly women. Criminals are leveraging AI-generated personas to lure victims into financial fraud. Authorities in Chengdu, Sichuan, recently buste











