DeepSeek API Input Cache Price Slashed to One-Tenth of Original
The leading domestic large language model, DeepSeek, recently announced a significant price cut, reducing the input cache hit price across all API series to one-tenth of the original rate. This move marks a new phase in cost management for domestic AI, aiming to attract more developers and businesses by offering exceptional value for money.
Core Price Cuts Address Industry Pain Points
This price adjustment covers the entire V4-Pro and V4-Flash series. The input cache price for V4-Pro has been reduced to 0.1 RMB per million tokens, and with a limited-time promotion, the actual payment is only 0.025 RMB. Compared to overseas competitors, the input cache price is just 1/700 of GPT-5.5 Pro, demonstrating strong market competitiveness.
In addition to cache hit scenarios, prices for cache miss and output scenarios have also been reduced to one-quarter of the original price. This pricing strategy precisely targets high-frequency use cases such as RAG knowledge bases, intelligent customer service, and document analysis, potentially cutting enterprise operational costs by over 90%.

DeepSeek's ability to significantly reduce prices stems from its self-developed sparse attention architecture. This technology supports ultra-long context processing of up to 160K, improving efficiency in handling long texts while effectively lowering underlying computing power consumption and storage costs.
Related article
Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation
Reports indicate that the launch of Google’s highly anticipated next-generation Gemini model has been pushed back. Internal disagreements over development priorities and resource allocation, combined with limited computing capacity and complex approv
OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating
In response to external scrutiny regarding decelerating sales growth and missed internal benchmarks, AI leader OpenAI issued a confident statement on Tuesday, April 28. The company clarified that its consumer products and enterprise services are adva
Alibaba Super Cup: Qwen3.8-Max Debuts with Boosted Coding and Office Tools
Alibaba has officially unveiled Qwen3.8-Max, a next-generation foundation large model boasting 2.4 trillion parameters. This significant AI advancement delivers substantial performance gains in core areas like coding and professional office tasks, sh
Related Special Topic Recommendations
Comments (0)
0/500
The leading domestic large language model, DeepSeek, recently announced a significant price cut, reducing the input cache hit price across all API series to one-tenth of the original rate. This move marks a new phase in cost management for domestic AI, aiming to attract more developers and businesses by offering exceptional value for money.
Core Price Cuts Address Industry Pain Points
This price adjustment covers the entire V4-Pro and V4-Flash series. The input cache price for V4-Pro has been reduced to 0.1 RMB per million tokens, and with a limited-time promotion, the actual payment is only 0.025 RMB. Compared to overseas competitors, the input cache price is just 1/700 of GPT-5.5 Pro, demonstrating strong market competitiveness.
In addition to cache hit scenarios, prices for cache miss and output scenarios have also been reduced to one-quarter of the original price. This pricing strategy precisely targets high-frequency use cases such as RAG knowledge bases, intelligent customer service, and document analysis, potentially cutting enterprise operational costs by over 90%.

DeepSeek's ability to significantly reduce prices stems from its self-developed sparse attention architecture. This technology supports ultra-long context processing of up to 160K, improving efficiency in handling long texts while effectively lowering underlying computing power consumption and storage costs.
Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation
Reports indicate that the launch of Google’s highly anticipated next-generation Gemini model has been pushed back. Internal disagreements over development priorities and resource allocation, combined with limited computing capacity and complex approv
OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating
In response to external scrutiny regarding decelerating sales growth and missed internal benchmarks, AI leader OpenAI issued a confident statement on Tuesday, April 28. The company clarified that its consumer products and enterprise services are adva
Alibaba Super Cup: Qwen3.8-Max Debuts with Boosted Coding and Office Tools
Alibaba has officially unveiled Qwen3.8-Max, a next-generation foundation large model boasting 2.4 trillion parameters. This significant AI advancement delivers substantial performance gains in core areas like coding and professional office tasks, sh





Home






