Home
27B Math SOTA and 3-Second Emotional Cloning: Youdao Open-Sources Zi Yue 4 Multimodal and TTS Engine
NetEase Youdao recently announced a comprehensive upgrade of its "Confucius4" large model to version 4.0. Now fully multimodal, Confucius4 supports integrated text, image, and audio interactions. Youdao also open-sourced its core multimodal and text-to-speech (TTS) models, while the translation model underwent a deep technical overhaul, delivering improvements in both quality and efficiency.
The multimodal model achieves SOTA in vision and mathematics, with superior performance on pure-text math problems
According to the announcement, the open-source Confucius4 multimodal model, with 27 billion parameters, has brought visual-input-based math capabilities to an industry-leading level (SOTA) in educational scenarios. Among models of similar scale, Confucius4 excels at handling advanced visual math and physics problems that incorporate charts. It also shows significant improvement on Chinese pure-text math problems, achieving an accuracy of 81.4%, which is industry-leading.

▲ Confucius4 achieves best-in-class results on multiple visual math benchmarks among models of the same scale
Image source: https://huggingface.co/netease-youdao/Confucius4
A more critical breakthrough lies in practical cost-effectiveness. According to relevant officials, the new model employs a refined reasoning chain reconstruction scheme. By aggregating a large volume of high-quality, concise reasoning samples for deep optimization, it compresses the output length of the reasoning chain by 43.2%.
This means it can deliver answers faster with fewer tokens and shorter reasoning paths, significantly reducing inference costs in real-world business scenarios for enterprises and developers.

▲ Confucius4 significantly reduces output tokens on multiple visual math benchmarks
Image source: https://huggingface.co/netease-youdao/Confucius4
Additionally, the Confucius4 research team deeply optimized the model for real-world homework, exams, and problem-solving scenarios faced by Chinese students. This enables it to address authentic learning challenges, making it a more empathetic digital assistant.
Open-source TTS: supports 14 languages, clones original voice in 3 seconds with no accent issues across languages
Alongside the multimodal model, the speech synthesis (TTS) engine was also open-sourced. Built on a cutting-edge "speech encoder + LLM" architecture, it offers developers and content creators zero-shot, low-barrier voice cloning and emotional synthesis capabilities.
Currently, it fully supports 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese. The system can naturally transfer a speaker's voice across different languages without additional training, maintaining voice consistency while ensuring the synthesized results sound native and fluent, with no accent leakage during cross-language cloning.
For voice cloning, Confucius4 achieves full "upload and clone" support. Users simply provide any audio material, and the system replicates the original voice within three seconds. According to the announcement, the engine's accuracy on cloning tasks exceeds 97%, and the similarity between the cloned voice and the original voice is over 85%. It preserves the speaker's unique vocal characteristics while accurately reproducing their emotional tone, with comprehensive capabilities that rank among the top in the field.
Furthermore, this open-source model demonstrates strong robustness in real multilingual scenarios. It can handle various synthesis needs, including daily conversations, news broadcasts, corporate promotions, and complex emotional expressions.
Translation model quality upgraded comprehensively, with 80% faster inference speed
As one of Youdao's most deeply rooted technological assets, the translation model also received significant technical upgrades in this update, further enhancing its performance on translation tasks.
Regarding data, the Confucius team collected and cleaned billions of multilingual items and hired professionals with TEM-8 certification for multidimensional manual evaluation, ensuring high-quality corpora from the start.
On the algorithmic side, the model uses an innovative "multi-expert OPD" mode, adopting a smarter "soft approach" to leverage the strengths of various experts. It also introduces format rewards and language detection mechanisms through reinforcement learning, effectively resolving common issues like out-of-context translations and mixed-language outputs in machine translation.
To meet the demands of high-frequency, high-concurrency industrial applications, the updated translation model is equipped with an efficient acceleration mechanism that directly boosts overall inference speed by 80%. Combined with a customized solution of automated large model evaluation and random manual sampling, the new generation translation model demonstrates extremely high standards of speed and quality across multiple scenarios, including text, image, and document translation.
Reflecting on Youdao's journey in AI, from the initial launch of Confucius as the first education-focused large model, introducing the "virtual speaking coach Hi Echo" that overturned traditional oral practice methods, to the comprehensive integration of Confucius 2.0 and 3.0 into software and hardware ecosystems, Youdao has consistently led AI empowerment in real-world scenarios. In 2026, Youdao accelerated AI application with a series of AI Agent products such as LobsterAI, Youdao Treasure, Youdao Conference Agent, and Thinkflow, realizing a forward-looking full-scenario AI Agent matrix.
The upgrade of Confucius4 and the full open-sourcing of core models not only significantly lower barriers for developers in multimodal and speech synthesis fields but also demonstrate an ecological closed loop where underlying core technologies nurture upper-layer Agent matrices. Youdao hopes that, with the joint contributions of global developers and the open-source community, this full-modal large model ecosystem will unleash true productivity transformation across a wide range of industries.
Appendix: Open-source addresses:
"Confucius4" multimodal model: https://huggingface.co/netease-youdao/Confucius4
"Confucius4" TTS model: https://github.com/netease-youdao/Confucius4-TTS
Related article
Cybersecurity Experts Criticize Guardrails on Anthropic’s Fable
Anthropic launched its newest model, Fable, on Tuesday, positioning it as a public, restricted iteration of its highly anticipated cybersecurity-focused model, Mythos.However, the limitations have sparked dissatisfaction among several cybersecurity r
Vbot VITA Power Raises Nearly 500 Million Yuan in Pre-A Funding, Begins Mass Delivery of First Embodied Intelligence Product
Vbot Weita Power, a pioneer in embodied intelligence, has secured nearly 500 million yuan in pre-A financing. Dongfang Jiafu, Huatai Zijin, and Fosun Ruizheng led the round, with additional backing from Shangqi Capital (SAIC Group) and existing share
Baidu Wenyin Unveils PaddleOCR-VL-1.6, Achieving 96.33% Accuracy and New SOTA in Document Parsing
Baidu has officially launched PaddleOCR-VL-1.6, a specialized variant of the ERNIE Large Model. In the authoritative OmnicDocBench v1.6 benchmark, it secured a 96.33% accuracy rate, outperforming leading models like Gemini-3-Pro, GPT-5.2, and GLM-OCR
Related Special Topic Recommendations
Comments (1)
0/500
NetEase Youdao recently announced a comprehensive upgrade of its "Confucius4" large model to version 4.0. Now fully multimodal, Confucius4 supports integrated text, image, and audio interactions. Youdao also open-sourced its core multimodal and text-to-speech (TTS) models, while the translation model underwent a deep technical overhaul, delivering improvements in both quality and efficiency.
The multimodal model achieves SOTA in vision and mathematics, with superior performance on pure-text math problems
According to the announcement, the open-source Confucius4 multimodal model, with 27 billion parameters, has brought visual-input-based math capabilities to an industry-leading level (SOTA) in educational scenarios. Among models of similar scale, Confucius4 excels at handling advanced visual math and physics problems that incorporate charts. It also shows significant improvement on Chinese pure-text math problems, achieving an accuracy of 81.4%, which is industry-leading.

▲ Confucius4 achieves best-in-class results on multiple visual math benchmarks among models of the same scale
Image source: https://huggingface.co/netease-youdao/Confucius4
A more critical breakthrough lies in practical cost-effectiveness. According to relevant officials, the new model employs a refined reasoning chain reconstruction scheme. By aggregating a large volume of high-quality, concise reasoning samples for deep optimization, it compresses the output length of the reasoning chain by 43.2%.
This means it can deliver answers faster with fewer tokens and shorter reasoning paths, significantly reducing inference costs in real-world business scenarios for enterprises and developers.

▲ Confucius4 significantly reduces output tokens on multiple visual math benchmarks
Image source: https://huggingface.co/netease-youdao/Confucius4
Additionally, the Confucius4 research team deeply optimized the model for real-world homework, exams, and problem-solving scenarios faced by Chinese students. This enables it to address authentic learning challenges, making it a more empathetic digital assistant.
Open-source TTS: supports 14 languages, clones original voice in 3 seconds with no accent issues across languages
Alongside the multimodal model, the speech synthesis (TTS) engine was also open-sourced. Built on a cutting-edge "speech encoder + LLM" architecture, it offers developers and content creators zero-shot, low-barrier voice cloning and emotional synthesis capabilities.
Currently, it fully supports 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese. The system can naturally transfer a speaker's voice across different languages without additional training, maintaining voice consistency while ensuring the synthesized results sound native and fluent, with no accent leakage during cross-language cloning.
For voice cloning, Confucius4 achieves full "upload and clone" support. Users simply provide any audio material, and the system replicates the original voice within three seconds. According to the announcement, the engine's accuracy on cloning tasks exceeds 97%, and the similarity between the cloned voice and the original voice is over 85%. It preserves the speaker's unique vocal characteristics while accurately reproducing their emotional tone, with comprehensive capabilities that rank among the top in the field.
Furthermore, this open-source model demonstrates strong robustness in real multilingual scenarios. It can handle various synthesis needs, including daily conversations, news broadcasts, corporate promotions, and complex emotional expressions.
Translation model quality upgraded comprehensively, with 80% faster inference speed
As one of Youdao's most deeply rooted technological assets, the translation model also received significant technical upgrades in this update, further enhancing its performance on translation tasks.
Regarding data, the Confucius team collected and cleaned billions of multilingual items and hired professionals with TEM-8 certification for multidimensional manual evaluation, ensuring high-quality corpora from the start.
On the algorithmic side, the model uses an innovative "multi-expert OPD" mode, adopting a smarter "soft approach" to leverage the strengths of various experts. It also introduces format rewards and language detection mechanisms through reinforcement learning, effectively resolving common issues like out-of-context translations and mixed-language outputs in machine translation.
To meet the demands of high-frequency, high-concurrency industrial applications, the updated translation model is equipped with an efficient acceleration mechanism that directly boosts overall inference speed by 80%. Combined with a customized solution of automated large model evaluation and random manual sampling, the new generation translation model demonstrates extremely high standards of speed and quality across multiple scenarios, including text, image, and document translation.
Reflecting on Youdao's journey in AI, from the initial launch of Confucius as the first education-focused large model, introducing the "virtual speaking coach Hi Echo" that overturned traditional oral practice methods, to the comprehensive integration of Confucius 2.0 and 3.0 into software and hardware ecosystems, Youdao has consistently led AI empowerment in real-world scenarios. In 2026, Youdao accelerated AI application with a series of AI Agent products such as LobsterAI, Youdao Treasure, Youdao Conference Agent, and Thinkflow, realizing a forward-looking full-scenario AI Agent matrix.
The upgrade of Confucius4 and the full open-sourcing of core models not only significantly lower barriers for developers in multimodal and speech synthesis fields but also demonstrate an ecological closed loop where underlying core technologies nurture upper-layer Agent matrices. Youdao hopes that, with the joint contributions of global developers and the open-source community, this full-modal large model ecosystem will unleash true productivity transformation across a wide range of industries.
Appendix: Open-source addresses:
"Confucius4" multimodal model: https://huggingface.co/netease-youdao/Confucius4
"Confucius4" TTS model: https://github.com/netease-youdao/Confucius4-TTS
Cybersecurity Experts Criticize Guardrails on Anthropic’s Fable
Anthropic launched its newest model, Fable, on Tuesday, positioning it as a public, restricted iteration of its highly anticipated cybersecurity-focused model, Mythos.However, the limitations have sparked dissatisfaction among several cybersecurity r
Vbot VITA Power Raises Nearly 500 Million Yuan in Pre-A Funding, Begins Mass Delivery of First Embodied Intelligence Product
Vbot Weita Power, a pioneer in embodied intelligence, has secured nearly 500 million yuan in pre-A financing. Dongfang Jiafu, Huatai Zijin, and Fosun Ruizheng led the round, with additional backing from Shangqi Capital (SAIC Group) and existing share
Baidu Wenyin Unveils PaddleOCR-VL-1.6, Achieving 96.33% Accuracy and New SOTA in Document Parsing
Baidu has officially launched PaddleOCR-VL-1.6, a specialized variant of the ERNIE Large Model. In the authoritative OmnicDocBench v1.6 benchmark, it secured a 96.33% accuracy rate, outperforming leading models like Gemini-3-Pro, GPT-5.2, and GLM-OCR











