chiprook
← AI
AISeptember 24, 2026, 22:43

Pipecat benchmarked 23 real-time STT models for voice agents: no single winner

Pipecat by Daily published an open benchmark of 23 real-time speech-to-text models using 1,000 real utterances, scoring latency (TTFS) and semantic accuracy. Results form a Pareto frontier: NVIDIA Nemotron 3.0 ASR is fastest (221 ms median), while Meta muse-voice-transcribe-1.0 is most accurate (0.83% semantic WER).

Pipecat benchmarked 23 real-time STT models for voice agents: no single winner
#Pipecat#Nvidia#Meta#Speechmatics
Read next
AI

ElevenLabs reportedly valued at $22B with $600M ARR

AI

Deutsche Telekom is replacing phone menus with AI agents

AI

OpenAI tests Sponsored Agents, opening brand chats from ChatGPT ads

AI

Alpha School students were monitored by offshore call center staff, not AI