chiprook
← AI
AIOctober 3, 2026, 03:31

Microsoft Ships First Streaming Transcription Model, Completing Voice Agent Pipeline on Foundry

On October 1, 2026, Microsoft launched MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash through Foundry and Azure AI Speech. The streaming model claims 0.13s end-of-speech latency versus 0.49s for Grok Voice Transcribe 2.0 and costs $0.54 per audio hour against $0.10 for the batch model.

Microsoft Ships First Streaming Transcription Model, Completing Voice Agent Pipeline on Foundry
#Microsoft#Azure#Foundry
Read next
AI

FlashAttention-4 ships for Blackwell B200 with FP4 and async pipelining

AI

Microsoft Foundry adds native real-time voice agents to rival Gemini Live

AI

Microsoft MAI voice models arrive in LiveKit for TTS agents

Software

GitLab Duo Self-Hosted adds support for models via Microsoft Foundry