Microsoft launches MAI-Transcribe-2 and MAI-Voice-2 speech models
Microsoft introduced three MAI speech models: MAI-Transcribe-2 for speech-to-text with diarization and word-level timestamps, MAI-Voice-2 for high-fidelity text-to-speech with prosody and emotion control, and MAI-Voice-2-Flash for low-latency voice agents. The models cover 60 languages, but native real-time streaming transcription is not documented.
- MAI-Transcribe-2 supports 60 languages, diarization and word-level timestamps
- MAI-Voice-2 offers prosody and emotion control for synthesized speech
- MAI-Voice-2-Flash targets low-latency voice-agent interactions
- Native real-time streaming transcription is not listed in the docs
Read next
AI