option
Home
List of Al models
DeepSeek-V2-Chat-0628

DeepSeek-V2-Chat-0628

Add comparison
Add comparison
Model parameter quantity
236B
Model parameter quantity
Affiliated organization
DeepSeek
Affiliated organization
Open Source
License Type
Release time
May 6, 2024
Release time

Model Introduction
DeepSeek-V2 is a strong Mixture-of-Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token. Compared with DeepSeek 67B, DeepSeek-V2 achieves stronger performance, and meanwhile saves 42.5% of training costs, reduces the KV cache by 93.3%, and boosts the maximum generation throughput to 5.76 times.
Swipe left and right to view more
Language comprehension ability Language comprehension ability
Language comprehension ability
Often makes semantic misjudgments, leading to obvious logical disconnects in responses.
4.6
Knowledge coverage scope Knowledge coverage scope
Knowledge coverage scope
Possesses core knowledge of mainstream disciplines, but has limited coverage of cutting-edge interdisciplinary fields.
7.8
Reasoning ability Reasoning ability
Reasoning ability
Unable to maintain coherent reasoning chains, often causing inverted causality or miscalculations.
4.7
Related model
DeepSeek-V4-Flash DeepSeek-V4-Flash is DeepSeek’s efficiency-oriented lightweight model for low-latency interaction, high-concurrency usage, and everyday intelligent applications.
DeepSeek-V4-Pro DeepSeek-V4-Pro is DeepSeek’s flagship high-performance model, further enhancing reasoning, coding, long-text processing, and overall task capability.
DeepSeek-V3.2 The latest version of Deepseek V3 series models.
DeepSeek-V3.2-Exp The latest experimental version of Deepseek V3 series models.
DeepSeek-R1-0528 The latest version of Deepseek R1.
Relevant documents
Anthropic Expands Claude AI Coding Tools to Japan in Push for Overseas Growth Anthropic, a prominent U.S. artificial intelligence firm, is intensifying its global outreach. On Wednesday, the company hosted a major developer gathering, "Code with Claude," in Tokyo, drawing close to 500 software engineers. This initiative seeks
India’s Software Giant Cuts Hiring, Promises No Layoffs as AI Agents Scale Up As artificial intelligence reshapes traditional labor-intensive business models, Tata Consultancy Services (TCS), a premier Indian software outsourcing firm, has unveiled its strategic response. During Tuesday’s annual general meeting, the chairman o
China Locks AI Models During Gaokam to Block Instant Homework Help With the 2026 Gaokao fast approaching, rumors regarding the suspension of AI tools during the exam period have ignited intense online debate. In response to public concern, major AI platforms and educational apps have clarified their stance: rather t
Survey reveals FDA-approved breast cancer AI diagnostic tool underperforms radiologists' expectations A recent study published in "Clinical Imaging" by researchers at the UC San Diego Health Center surveyed 215 members of the American Society of Breast Imaging. The findings reveal that AI tools for breast cancer detection fall short of radiologists'
Ant Brain Full-Stack 2.0 Debuts at WAIC, Powering Smart Pharmacy With Unified Intelligence The 2026 World Artificial Intelligence Conference (WAIC) kicks off on July 17, with the prestigious "Treasures of the Exhibition" awards taking center stage. This year’s top ten selections highlight groundbreaking innovations, including Ant Group’s "
Model comparison
Start the comparison
OR