DeepSeek-V4.1-Flash Debuts on Qwen AI Platform With API and Token Plan Live
DeepSeek-V4.1-Flash is now live on the Qwen AI platform, offering both API access and a flexible Token Plan. Developers can seamlessly integrate the model into their workflows via standard APIs or leverage the Token Plan directly within tools like Qoder, the Qwen APP, and Codex for coding, documentation, visual analysis, and agent-driven tasks.

Described as DeepSeek’s new lightweight flagship, this model utilizes a MoE architecture with 552B total parameters. It employs a Causal-Encoder-Decoder asymmetric structure, activating only ~8B parameters for input and ~16B for output. Natively supporting text and image understanding, it handles context windows up to 1 million tokens and generates outputs of approximately 393K. Official benchmarks indicate significant improvements in Agent and code performance, delivering high throughput and low latency despite reduced activation costs.
Cost efficiency is a key highlight. Advanced cache compression reduces KV Cache HBM requirements to 25% and SSD storage needs to 12.5% of previous generations, optimizing resources for long contexts and multi-turn interactions. Pricing on the Qwen platform is tiered: input costs 1 CNY per million tokens during off-peak hours (2 CNY peak), while output is 4 CNY per million tokens (8 CNY peak). Rate limits stand at 15K RPM and 1M TPM, supporting prefix completion, function calling, caching, structured output, batch processing, and online search.
Alibaba Cloud BaiLian has partnered with the vLLM open-source community to ensure broad deployment across China and international regions, including Beijing, Singapore, the US, Germany, Japan, and Hong Kong. The model supports multi-turn conversations, function calls, online search, context caching, and structured output, with a max_tokens limit of ~393,216. This makes it ideal for enterprises migrating long document parsing, customer service knowledge bases, code repository Q&A, and multimodal order reviews to a unified interface.
Industry analysts suggest that integrating "large parameters, small activation, and robust caching" into the Qwen ecosystem lowers the trial-and-error costs for long-context agents. For small teams, the combination of low off-peak pricing and flexible Token Plan subscriptions enables more agile batch inference and rapid prototype validation.
Related article
GitHub Copilot Integrates GPT-5.4 in Hours
GitHub Copilot has once again proven its rapid response capabilities. Mere hours after OpenAI launched its newest flagship model, GPT-5.4, GitHub announced full integration, providing developers worldwide with intelligent coding assistance powered by
Anthropic Leads the Way in IPO, AI Industry Enters a New Era of Trillion-Dollar Revenue
The global AI sector has hit another major milestone. Recent market data reveals that AI startup Anthropic filed a confidential IPO application on June 1st. If approved, Anthropic is poised to become the largest publicly traded company in the AI lab
How to optimize SEO for Google Japan and Yahoo Japan?
The connected city is already here, but it looks ordinaryThe most useful IoT infrastructure rarely looks futuristic. It looks like a waste bin that knows when it is full, a van that reports engine faults before it breaks down, a parking bay that tell
Related Special Topic Recommendations
Comments (0)
0/500
DeepSeek-V4.1-Flash is now live on the Qwen AI platform, offering both API access and a flexible Token Plan. Developers can seamlessly integrate the model into their workflows via standard APIs or leverage the Token Plan directly within tools like Qoder, the Qwen APP, and Codex for coding, documentation, visual analysis, and agent-driven tasks.

Described as DeepSeek’s new lightweight flagship, this model utilizes a MoE architecture with 552B total parameters. It employs a Causal-Encoder-Decoder asymmetric structure, activating only ~8B parameters for input and ~16B for output. Natively supporting text and image understanding, it handles context windows up to 1 million tokens and generates outputs of approximately 393K. Official benchmarks indicate significant improvements in Agent and code performance, delivering high throughput and low latency despite reduced activation costs.
Cost efficiency is a key highlight. Advanced cache compression reduces KV Cache HBM requirements to 25% and SSD storage needs to 12.5% of previous generations, optimizing resources for long contexts and multi-turn interactions. Pricing on the Qwen platform is tiered: input costs 1 CNY per million tokens during off-peak hours (2 CNY peak), while output is 4 CNY per million tokens (8 CNY peak). Rate limits stand at 15K RPM and 1M TPM, supporting prefix completion, function calling, caching, structured output, batch processing, and online search.
Alibaba Cloud BaiLian has partnered with the vLLM open-source community to ensure broad deployment across China and international regions, including Beijing, Singapore, the US, Germany, Japan, and Hong Kong. The model supports multi-turn conversations, function calls, online search, context caching, and structured output, with a max_tokens limit of ~393,216. This makes it ideal for enterprises migrating long document parsing, customer service knowledge bases, code repository Q&A, and multimodal order reviews to a unified interface.
Industry analysts suggest that integrating "large parameters, small activation, and robust caching" into the Qwen ecosystem lowers the trial-and-error costs for long-context agents. For small teams, the combination of low off-peak pricing and flexible Token Plan subscriptions enables more agile batch inference and rapid prototype validation.
GitHub Copilot Integrates GPT-5.4 in Hours
GitHub Copilot has once again proven its rapid response capabilities. Mere hours after OpenAI launched its newest flagship model, GPT-5.4, GitHub announced full integration, providing developers worldwide with intelligent coding assistance powered by
Anthropic Leads the Way in IPO, AI Industry Enters a New Era of Trillion-Dollar Revenue
The global AI sector has hit another major milestone. Recent market data reveals that AI startup Anthropic filed a confidential IPO application on June 1st. If approved, Anthropic is poised to become the largest publicly traded company in the AI lab





Home






