Home
Baidu Wenyin Unveils PaddleOCR-VL-1.6, Achieving 96.33% Accuracy and New SOTA in Document Parsing

Baidu has officially launched PaddleOCR-VL-1.6, a specialized variant of the ERNIE Large Model. In the authoritative OmnicDocBench v1.6 benchmark, it secured a 96.33% accuracy rate, outperforming leading models like Gemini-3-Pro, GPT-5.2, and GLM-OCR. This achievement establishes a new industry SOTA and ranks first globally in comprehensive performance, marking a major leap forward in how multi-modal large models comprehend complex documents and analyze real-world contexts.
As a key element of the ERNIE Large Model’s multi-modal ecosystem, PaddleOCR is built upon the ERNIE foundation and currently supports over 100 languages, serving users in more than 170 countries and regions. The upgraded PaddleOCR-VL-1.6 retains a lightweight 0.9B architecture while significantly enhancing core recognition capabilities in challenging scenarios—including tables, ancient texts, rare characters, seals, and charts—through a model-driven data construction mechanism and progressive training optimization.
In the Real5-OmniDocBench evaluation, which tests performance in real-world complex environments, the model maintained its lead with a total score of 93.19%. It successfully overcame widely recognized parsing challenges such as scanned documents, bent pages, screen captures, lighting variations, and tilted text.
Leveraging the established architecture, enterprises and developers can migrate smoothly without requiring additional adaptation. PaddleOCR has already surpassed 79.2K stars on GitHub, overtaking Google’s Tesseract OCR to become the world’s most popular open-source OCR project. The new model is now available on the official website, with both code and weights open-sourced. As large models increasingly evolve toward multi-modal depth, PaddleOCR-VL-1.6 offers a more efficient industrial-grade solution for document digitization and will further accelerate AI deployment in complex multi-modal applications.
Related article
Anthropic Launches Claude for Small Business, an AI Skill Pack for SMEs
Anthropic officially launched Claude for Small Business on Wednesday, introducing a suite of automation tools tailored for small and medium-sized enterprises (SMBs). This initiative seeks to reduce technical barriers, enabling local businesses withou
German Research Consortium Unveils Open-Source AI Model Soofi
The German Association for Artificial Intelligence has officially launched Soofi S30B-A3B, a new open-source large language model. This milestone advances Europe’s sovereign AI infrastructure and revitalizes the open-source sector with a focus on hig
Toyota’s $6.4bn robotics estimate puts physical AI in focus
Toyota Motor projects that rolling out automation across its manufacturing facilities, subsidiaries, and key suppliers may require approximately 400,000 robots and an annual expenditure of roughly 1 trillion yen ($6.4 billion) starting in 2028.Accord
Related Special Topic Recommendations
Comments (0)
0/500

Baidu has officially launched PaddleOCR-VL-1.6, a specialized variant of the ERNIE Large Model. In the authoritative OmnicDocBench v1.6 benchmark, it secured a 96.33% accuracy rate, outperforming leading models like Gemini-3-Pro, GPT-5.2, and GLM-OCR. This achievement establishes a new industry SOTA and ranks first globally in comprehensive performance, marking a major leap forward in how multi-modal large models comprehend complex documents and analyze real-world contexts.
As a key element of the ERNIE Large Model’s multi-modal ecosystem, PaddleOCR is built upon the ERNIE foundation and currently supports over 100 languages, serving users in more than 170 countries and regions. The upgraded PaddleOCR-VL-1.6 retains a lightweight 0.9B architecture while significantly enhancing core recognition capabilities in challenging scenarios—including tables, ancient texts, rare characters, seals, and charts—through a model-driven data construction mechanism and progressive training optimization.
In the Real5-OmniDocBench evaluation, which tests performance in real-world complex environments, the model maintained its lead with a total score of 93.19%. It successfully overcame widely recognized parsing challenges such as scanned documents, bent pages, screen captures, lighting variations, and tilted text.
Leveraging the established architecture, enterprises and developers can migrate smoothly without requiring additional adaptation. PaddleOCR has already surpassed 79.2K stars on GitHub, overtaking Google’s Tesseract OCR to become the world’s most popular open-source OCR project. The new model is now available on the official website, with both code and weights open-sourced. As large models increasingly evolve toward multi-modal depth, PaddleOCR-VL-1.6 offers a more efficient industrial-grade solution for document digitization and will further accelerate AI deployment in complex multi-modal applications.
Anthropic Launches Claude for Small Business, an AI Skill Pack for SMEs
Anthropic officially launched Claude for Small Business on Wednesday, introducing a suite of automation tools tailored for small and medium-sized enterprises (SMBs). This initiative seeks to reduce technical barriers, enabling local businesses withou
German Research Consortium Unveils Open-Source AI Model Soofi
The German Association for Artificial Intelligence has officially launched Soofi S30B-A3B, a new open-source large language model. This milestone advances Europe’s sovereign AI infrastructure and revitalizes the open-source sector with a focus on hig
Toyota’s $6.4bn robotics estimate puts physical AI in focus
Toyota Motor projects that rolling out automation across its manufacturing facilities, subsidiaries, and key suppliers may require approximately 400,000 robots and an annual expenditure of roughly 1 trillion yen ($6.4 billion) starting in 2028.Accord











