Home
Microsoft Unveils MAI-Transcribe-2, Its Most Accurate Transcription Model Yet, Supporting 60 Languages at Just 0.67 Yuan Per Hour
Microsoft has launched MAI-Transcribe-2, positioning it as the industry’s fastest, most precise, and cost-efficient AI speech-to-text solution. Priced at a promotional rate of $0.10 (approx. 0.67 RMB) per hour, this offer remains valid through the end of 2026.

Unmatched Accuracy and Speed: 10x Faster than GPT-Transcribe
Performance benchmarks on the FLEURS dataset show MAI-Transcribe-2 achieving a 5.2% average word error rate (WER) in both forced-language and auto-detection modes. By comparison, Gemini 3.1 Pro scored 5.3% and 5.8%, GPT-Transcribe reached 10.4% and 10.6%, while Whisper v3-Large recorded 22.8% and 23.5%.
Regarding processing speed, Microsoft cites Artificial Analysis data indicating MAI-Transcribe-2 operates 10 times faster than GPT-Transcribe, 7 times faster than Scribe v2, and 5 times faster than Gemini 3.5 Transcribe.
Advanced Features: Speaker Diarization, Timestamps, and 60 Languages
Key features include speaker separation, which identifies distinct voices within a recording and assigns text to the correct speaker. Word-level timestamps enable precise audio navigation, editing, and subtitle synchronization.
Users can customize transcription styles: verbatim mode retains filler words and errors for legal or analytical accuracy, while clean mode removes them for readable notes and subtitles. The model also supports keyword bias, automatic language detection, dynamic language switching, and robust performance in noisy environments. Covering 60 languages, it handles mixed-language conversations seamlessly without requiring pre-specified language settings.
Developers can access or test the model via Microsoft Foundry, MAI Playground, and OpenRouter.
Related article
Tencent Accelerates AI Push With Doubled-Size Hunyuan Hy4, Morgan Stanley Lifts Target to HKD 550
Tencent has officially unveiled the preview of HuanYuan Hy4, accelerating its timeline ahead of Morgan Stanley’s September forecast and fulfilling the company’s Q2 earnings commitment to launch the model within the year. Since appointing Yao Shunyu a
Pearson’s Dave Treat: AI Should Guide Learning, Not Cheat
Pearson’s Chief Technology Officer, Dave TreatOn academic integrity, AI adoption, and the critical skills that matter as agentic AI transforms educationDave Treat joined Pearson nearly two years ago as Chief Technology Officer, later assuming the rol
OpenAI Launches GPT-5.6-Cyber, Boosting Zero-Day Detection with AI Automation
OpenAI has expanded its Daybreak cybersecurity initiative, introducing a two-tier access system and launching GPT-5.6-Cyber, a model purpose-built for security operations. This move addresses the growing reality that threat actors are leveraging AI t
Related Special Topic Recommendations
Comments (0)
0/500
Microsoft has launched MAI-Transcribe-2, positioning it as the industry’s fastest, most precise, and cost-efficient AI speech-to-text solution. Priced at a promotional rate of $0.10 (approx. 0.67 RMB) per hour, this offer remains valid through the end of 2026.

Unmatched Accuracy and Speed: 10x Faster than GPT-Transcribe
Performance benchmarks on the FLEURS dataset show MAI-Transcribe-2 achieving a 5.2% average word error rate (WER) in both forced-language and auto-detection modes. By comparison, Gemini 3.1 Pro scored 5.3% and 5.8%, GPT-Transcribe reached 10.4% and 10.6%, while Whisper v3-Large recorded 22.8% and 23.5%.
Regarding processing speed, Microsoft cites Artificial Analysis data indicating MAI-Transcribe-2 operates 10 times faster than GPT-Transcribe, 7 times faster than Scribe v2, and 5 times faster than Gemini 3.5 Transcribe.
Advanced Features: Speaker Diarization, Timestamps, and 60 Languages
Key features include speaker separation, which identifies distinct voices within a recording and assigns text to the correct speaker. Word-level timestamps enable precise audio navigation, editing, and subtitle synchronization.
Users can customize transcription styles: verbatim mode retains filler words and errors for legal or analytical accuracy, while clean mode removes them for readable notes and subtitles. The model also supports keyword bias, automatic language detection, dynamic language switching, and robust performance in noisy environments. Covering 60 languages, it handles mixed-language conversations seamlessly without requiring pre-specified language settings.
Developers can access or test the model via Microsoft Foundry, MAI Playground, and OpenRouter.
Tencent Accelerates AI Push With Doubled-Size Hunyuan Hy4, Morgan Stanley Lifts Target to HKD 550
Tencent has officially unveiled the preview of HuanYuan Hy4, accelerating its timeline ahead of Morgan Stanley’s September forecast and fulfilling the company’s Q2 earnings commitment to launch the model within the year. Since appointing Yao Shunyu a
Pearson’s Dave Treat: AI Should Guide Learning, Not Cheat
Pearson’s Chief Technology Officer, Dave TreatOn academic integrity, AI adoption, and the critical skills that matter as agentic AI transforms educationDave Treat joined Pearson nearly two years ago as Chief Technology Officer, later assuming the rol
OpenAI Launches GPT-5.6-Cyber, Boosting Zero-Day Detection with AI Automation
OpenAI has expanded its Daybreak cybersecurity initiative, introducing a two-tier access system and launching GPT-5.6-Cyber, a model purpose-built for security operations. This move addresses the growing reality that threat actors are leveraging AI t











