Home
Tencent, Tsinghua Launch SongGeneration 2 With 8.55% Error Rate, Intensifying Pressure on Suno
The AI music landscape experienced another significant shift in early 2026. On March 9, the music foundation model SongGeneration2, developed jointly by Tencent and Tsinghua University's Human-Computer Speech Interaction Lab , was officially released. This model not only marked a qualitative leap in technical architecture but also outperformed current mainstream open-source models across multiple core dimensions, even rivaling top commercial models in overall quality.

Three Breakthroughs: No More 'Plastic' AI Music
SongGeneration2 ’s core advantages come from a comprehensive upgrade of its underlying architecture, addressing three key pain points that plagued earlier AI music:
High musicality: Unlike basic melody stacking, this model handles complex multi-track arrangements with impressive spatial depth.
High lyric accuracy: Unclear pronunciation and hallucinated pitch shifts are now history. Its phoneme error rate (PER) is as low as 8.55%, notably outperforming the leading commercial model Suno v5 (12.4%), and trailing only MiniMax2.5 .
Strong controllability: Whether through text descriptions or audio prompts, it precisely follows instructions, enabling deep customization of style and emotion.

'Dual-Core' Drive: A Dream Collaboration Between LLM and Diffusion Models
Architecturally, SongGeneration2 employs an innovative hybrid LLM-diffusion architecture:
Composing Brain (LeLM): Handles global structure planning and vocal details, solving the challenge of "how to sing".
High-Fidelity Renderer (Diffusion): Synthesizes extremely complex acoustic details under the guidance of the language model.
Hierarchical Representation: The first to adopt parallel modeling of mixed and multi-track representations, balancing melodic stability with sound quality refinement.
True Open Source, Low Barrier: Ordinary Computers Can Also 'Write Songs'
What excited developers most was Tencent's remarkable openness. The SongGeneration-v2-large model with 4 billion parameters is now fully open-source, supporting multilingual generation including Chinese and English. Remarkably, it runs smoothly on consumer-grade hardware with just 22GB VRAM, enabling local and private music creation.
To let users try it right away, the team also released the SongGeneration-v2-Fast version on HuggingFace, trading a minimal drop in audio quality for ultra-fast generation—producing a complete song in under a minute.
Based on the performance of SongGeneration2 , AI music has officially evolved from a 'geek toy' into the realm of 'commercial-grade applications.' With the upcoming open-sourcing of a Medium model supporting 12GB VRAM and an automated evaluation framework, the era of everyone becoming a 'composer' may finally be upon us.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
The AI music landscape experienced another significant shift in early 2026. On March 9, the music foundation model SongGeneration2, developed jointly by

Three Breakthroughs: No More 'Plastic' AI Music
High musicality: Unlike basic melody stacking, this model handles complex multi-track arrangements with impressive spatial depth.
High lyric accuracy: Unclear pronunciation and hallucinated pitch shifts are now history. Its phoneme error rate (PER) is as low as 8.55%, notably outperforming the leading commercial model
Strong controllability: Whether through text descriptions or audio prompts, it precisely follows instructions, enabling deep customization of style and emotion.

'Dual-Core' Drive: A Dream Collaboration Between LLM and Diffusion Models
Architecturally,
Composing Brain (LeLM): Handles global structure planning and vocal details, solving the challenge of "how to sing".
High-Fidelity Renderer (Diffusion): Synthesizes extremely complex acoustic details under the guidance of the language model.
Hierarchical Representation: The first to adopt parallel modeling of mixed and multi-track representations, balancing melodic stability with sound quality refinement.
True Open Source, Low Barrier: Ordinary Computers Can Also 'Write Songs'
What excited developers most was Tencent's remarkable openness. The SongGeneration-v2-large model with 4 billion parameters is now fully open-source, supporting multilingual generation including Chinese and English. Remarkably, it runs smoothly on consumer-grade hardware with just 22GB VRAM, enabling local and private music creation.
To let users try it right away, the team also released the SongGeneration-v2-Fast version on HuggingFace, trading a minimal drop in audio quality for ultra-fast generation—producing a complete song in under a minute.
Based on the performance of
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











