Fish Audio Launches S2: Open-Source Model Enables Word-Level Emotion Control

Fish Audio has officially launched its new text-to-speech model, S2, representing a significant leap forward in expressiveness and controllability for open-source TTS technology.
Named Fish Audio S2, this model prioritizes powerful emotional control. Users can make fine-grained adjustments to prosody and emotion using natural language instructions. By inserting tags like [laugh], [whisper], or [super happy], or even using free-form descriptions such as [professional broadcast tone] or [pitch up], it enables precise, word-level control to generate highly expressive and naturally vivid speech.
Key features include:
Completely open source: The model weights, fine-tuning code, and streaming inference engine based on SGLang are all publicly available on GitHub and Hugging Face. S2-Pro is the flagship version with approximately 4.4 billion parameters.Ultra-low latency: Inference latency is under 150 milliseconds, making it ideal for real-time applications like chatbots and virtual streamers.Native multi-speaker support: It can process multiple speakers in a single inference, handling conversational turns, interruptions, and natural emotional delivery while maintaining consistent voice quality without extra processing.Fish Audio reports that S2 was trained on roughly 10 million hours of audio data spanning nearly 50 languages. Utilizing reinforcement learning alignment and a dual autoregressive architecture, it demonstrates leading naturalness and expressiveness in multiple benchmarks. It is considered one of the most emotionally intelligent TTS systems available, open-source or proprietary. "True linguistic freedom starts now," Fish Audio announced, marking the arrival of AI speech with genuine emotion and personality.
GitHub:https://github.com/fishaudio/fish-speech/
HuggingFace:https://huggingface.co/fishaudio/s2-pro/
Related article
Apple Smart Glasses Could Debut at WWDC27, Highlighting Privacy Protection
Bloomberg’s Mark Gurman reports that Apple’s smart glasses, codenamed N50, are slated for a WWDC27 debut in June 2027, with a retail launch expected in autumn 2027. Originally targeted for late this year and early 2027, the device’s release has been
Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation
Reports indicate that the launch of Google’s highly anticipated next-generation Gemini model has been pushed back. Internal disagreements over development priorities and resource allocation, combined with limited computing capacity and complex approv
OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating
In response to external scrutiny regarding decelerating sales growth and missed internal benchmarks, AI leader OpenAI issued a confident statement on Tuesday, April 28. The company clarified that its consumer products and enterprise services are adva
Related Special Topic Recommendations
Comments (0)
0/500

Fish Audio has officially launched its new text-to-speech model, S2, representing a significant leap forward in expressiveness and controllability for open-source TTS technology.
Named Fish Audio S2, this model prioritizes powerful emotional control. Users can make fine-grained adjustments to prosody and emotion using natural language instructions. By inserting tags like [laugh], [whisper], or [super happy], or even using free-form descriptions such as [professional broadcast tone] or [pitch up], it enables precise, word-level control to generate highly expressive and naturally vivid speech.
Key features include:
Completely open source: The model weights, fine-tuning code, and streaming inference engine based on SGLang are all publicly available on GitHub and Hugging Face. S2-Pro is the flagship version with approximately 4.4 billion parameters.Ultra-low latency: Inference latency is under 150 milliseconds, making it ideal for real-time applications like chatbots and virtual streamers.Native multi-speaker support: It can process multiple speakers in a single inference, handling conversational turns, interruptions, and natural emotional delivery while maintaining consistent voice quality without extra processing.Fish Audio reports that S2 was trained on roughly 10 million hours of audio data spanning nearly 50 languages. Utilizing reinforcement learning alignment and a dual autoregressive architecture, it demonstrates leading naturalness and expressiveness in multiple benchmarks. It is considered one of the most emotionally intelligent TTS systems available, open-source or proprietary. "True linguistic freedom starts now," Fish Audio announced, marking the arrival of AI speech with genuine emotion and personality.
GitHub:https://github.com/fishaudio/fish-speech/
HuggingFace:https://huggingface.co/fishaudio/s2-pro/
Apple Smart Glasses Could Debut at WWDC27, Highlighting Privacy Protection
Bloomberg’s Mark Gurman reports that Apple’s smart glasses, codenamed N50, are slated for a WWDC27 debut in June 2027, with a retail launch expected in autumn 2027. Originally targeted for late this year and early 2027, the device’s release has been
Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation
Reports indicate that the launch of Google’s highly anticipated next-generation Gemini model has been pushed back. Internal disagreements over development priorities and resource allocation, combined with limited computing capacity and complex approv
OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating
In response to external scrutiny regarding decelerating sales growth and missed internal benchmarks, AI leader OpenAI issued a confident statement on Tuesday, April 28. The company clarified that its consumer products and enterprise services are adva





Home






