option
Home
Flash News
Content
TerryYoung
TerryYoung
March 30, 2026

Microsoft open-sourced its VibeVoice AI models for speech-to-text and text-to-speech. The suite includes VibeVoice-ASR-7B for transcribing hour-long audio, VibeVoice-TTS-1.5B for generating multi-speaker dialogues up to 90 minutes, and a real-time TTS model with 300ms latency. Released under MIT license, it supports local deployment and aims to advance speech synthesis innovation.

Microsoft open-sourced its VibeVoice AI models for speech-to-text and text-to-speech. The suite includes VibeVoice-ASR-7B for transcribing hour-long audio, VibeVoice-TTS-1.5B for generating multi-speaker dialogues up to 90 minutes, and a real-time TTS model with 300ms latency. Released under MIT license, it supports local deployment and aims to advance speech synthesis innovation.
Comments (0)
0/300
OR