DeepSeek Unveils First Open-Source Multimodal Model Engineered for AI Agents
DeepSeek has released DeepSeek-V4-Flash-Vision-Exp on Hugging Face, marking the debut of the first experimental multimodal model in the V4 series under the MIT License. With 305B parameters and built on the V4-Flash-0731 foundation, this model integrates a visual encoder and Aligner into its core language architecture, enabling robust image comprehension through continuous training.

Targeting Multimodal Agents, Not Just Image Description
Unlike traditional multimodal models focused on visual question answering, DeepSeek has redefined the Vision version’s purpose: prioritizing multimodal agent capabilities. The official documentation highlights that agents can directly interpret visual inputs—such as web screenshots, software interfaces, and charts—and execute tasks by invoking tools. Essentially, this “eye” is designed for machine agents rather than human users.
The open-source package is comprehensive, featuring model weights, Tokenizer, Prompt Encoding reference implementations, and a minimal PyTorch inference setup. It covers essential modules including the visual encoder, Aligner, DFlash Attention, MoE, and Hyper-Connections. The community responded swiftly, with seven quantized versions already available on Hugging Face for llama.cpp, LM Studio, and Ollama.
Notable timeline details reveal a strategic rollout: the model first appeared on the DeepSeek API on August 21, accessible only via API without weight distribution. Developers could input text and images simultaneously, with image processing charged by token. Ten days later, weights were officially released, enabling local deployment and secondary development. This “API-first, then open-source” strategy underscores DeepSeek’s primary positioning as an API service provider.

Performance Near Claude Opus 4.8, with Measured Official Claims
Benchmark results indicate that adding visual capabilities does not significantly compromise pure text agent performance: Terminal Bench 2.1 scores rose from 82.7 to 83.9, and DeepSWE improved from 54.4 to 59.3, surpassing Opus 4.8’s 58.0. Multimodal agent improvements are more pronounced: ApexBench Pass@1 reached 36.5, Agents' Last Exam scored 27.3 (exceeding Opus 4.8’s 25.7), and ZeroBench Pass@5 hit 35.0, outperforming the competitor’s 34.0.
However, DeepSeek avoids claiming “comprehensive superiority”: on the NL2Repo project, it scored 57.7, trailing Opus 4.8’s 69.7. The official statement remains measured, noting only that “multimodal agent capabilities are close to Claude Opus 4.8.” Recent updates show DeepSeek assembling a complete agent technology stack: the model handles reasoning and tool calls, Harness manages continuous task execution, and Vision enables agents to directly interpret computer visual information. This open-source release is a critical step in completing this ecosystem.
Related article
US health agencies evaluate OpenAI and Anthropic AI models
Public health agencies nationwide are set to evaluate generative AI through a new initiative led by the Coalition for Health AI, in partnership with OpenAI, Anthropic, and Accenture.The Public Health Use Case and Learning Scaling Engine (PULSE) will
Alipay’s Touch and Pay to Transform 30 Million Offline Touchpoints Into an AI-Driven Business Network
Following the launch of a full-stack AI payment system and the AI-powered Alipay assistant "Abao," Alipay announced on July 8 that its "Touch and Pay" solution has undergone a comprehensive AI upgrade. The millions of "Touch and Pay" devices deployed
Tealium: Why RAG Quality Depends on Real-Time Data
Freddy Berlanti, Principal Product Manager of Applied AI at Tealium, explores retrieval-augmented generation (RAG). Image: Getty ImagesFreddy Berlanti, Principal Product Manager of Applied AI at Tealium, breaks down the critical divide between RAG sy
Related Special Topic Recommendations
Comments (0)
0/500
DeepSeek has released DeepSeek-V4-Flash-Vision-Exp on Hugging Face, marking the debut of the first experimental multimodal model in the V4 series under the MIT License. With 305B parameters and built on the V4-Flash-0731 foundation, this model integrates a visual encoder and Aligner into its core language architecture, enabling robust image comprehension through continuous training.

Targeting Multimodal Agents, Not Just Image Description
Unlike traditional multimodal models focused on visual question answering, DeepSeek has redefined the Vision version’s purpose: prioritizing multimodal agent capabilities. The official documentation highlights that agents can directly interpret visual inputs—such as web screenshots, software interfaces, and charts—and execute tasks by invoking tools. Essentially, this “eye” is designed for machine agents rather than human users.
The open-source package is comprehensive, featuring model weights, Tokenizer, Prompt Encoding reference implementations, and a minimal PyTorch inference setup. It covers essential modules including the visual encoder, Aligner, DFlash Attention, MoE, and Hyper-Connections. The community responded swiftly, with seven quantized versions already available on Hugging Face for llama.cpp, LM Studio, and Ollama.
Notable timeline details reveal a strategic rollout: the model first appeared on the DeepSeek API on August 21, accessible only via API without weight distribution. Developers could input text and images simultaneously, with image processing charged by token. Ten days later, weights were officially released, enabling local deployment and secondary development. This “API-first, then open-source” strategy underscores DeepSeek’s primary positioning as an API service provider.

Performance Near Claude Opus 4.8, with Measured Official Claims
Benchmark results indicate that adding visual capabilities does not significantly compromise pure text agent performance: Terminal Bench 2.1 scores rose from 82.7 to 83.9, and DeepSWE improved from 54.4 to 59.3, surpassing Opus 4.8’s 58.0. Multimodal agent improvements are more pronounced: ApexBench Pass@1 reached 36.5, Agents' Last Exam scored 27.3 (exceeding Opus 4.8’s 25.7), and ZeroBench Pass@5 hit 35.0, outperforming the competitor’s 34.0.
However, DeepSeek avoids claiming “comprehensive superiority”: on the NL2Repo project, it scored 57.7, trailing Opus 4.8’s 69.7. The official statement remains measured, noting only that “multimodal agent capabilities are close to Claude Opus 4.8.” Recent updates show DeepSeek assembling a complete agent technology stack: the model handles reasoning and tool calls, Harness manages continuous task execution, and Vision enables agents to directly interpret computer visual information. This open-source release is a critical step in completing this ecosystem.
US health agencies evaluate OpenAI and Anthropic AI models
Public health agencies nationwide are set to evaluate generative AI through a new initiative led by the Coalition for Health AI, in partnership with OpenAI, Anthropic, and Accenture.The Public Health Use Case and Learning Scaling Engine (PULSE) will
Alipay’s Touch and Pay to Transform 30 Million Offline Touchpoints Into an AI-Driven Business Network
Following the launch of a full-stack AI payment system and the AI-powered Alipay assistant "Abao," Alipay announced on July 8 that its "Touch and Pay" solution has undergone a comprehensive AI upgrade. The millions of "Touch and Pay" devices deployed
Tealium: Why RAG Quality Depends on Real-Time Data
Freddy Berlanti, Principal Product Manager of Applied AI at Tealium, explores retrieval-augmented generation (RAG). Image: Getty ImagesFreddy Berlanti, Principal Product Manager of Applied AI at Tealium, breaks down the critical divide between RAG sy





Home






