chiprook
← AI
AIOctober 2, 2026, 01:30

Microsoft launches MAI-Transcribe-2 and MAI-Voice-2 speech models

Microsoft introduced three MAI speech models: MAI-Transcribe-2 for speech-to-text with diarization and word-level timestamps, MAI-Voice-2 for high-fidelity text-to-speech with prosody and emotion control, and MAI-Voice-2-Flash for low-latency voice agents. The models cover 60 languages, but native real-time streaming transcription is not documented.

Microsoft launches MAI-Transcribe-2 and MAI-Voice-2 speech models
#Microsoft#MAI-Transcribe-2#MAI-Voice-2
Read next
AI

ElevenLabs launches v4 and v4 Turbo speech models with 90 languages

AI

Sarvam AI launches Saaras V4 multilingual speech recognition model

AI

Kyutai releases Voice of Reason: speech-native models solve spoken math

AI

Oído: open-source speech recognition for ESP32-S3 without a command list