Qwen launched Qwen-Audio-3.0-TTS, a voice synthesis model with free-style natural language instruction control. Users can describe desired tone or style in everyday language instead of adjusting engineering parameters. Two versions offered: Flash for real-time interaction with 300ms first-packet delay, and Plus for high-quality audiobooks and dubbing. This lowers the barrier for voice synthesis, turning it into a tool for everyone.





Home
