Pipecat benchmarked 23 real-time STT models for voice agents: no single winner
Pipecat by Daily published an open benchmark of 23 real-time speech-to-text models using 1,000 real utterances, scoring latency (TTFS) and semantic accuracy. Results form a Pareto frontier: NVIDIA Nemotron 3.0 ASR is fastest (221 ms median), while Meta muse-voice-transcribe-1.0 is most accurate (0.83% semantic WER).
- Benchmark: 1,000 real utterances, streaming, TTFS and semantic WER, open code and data
- Nemotron 3.0 ASR: 221 ms median TTFS at 1.95% semantic WER
- Meta muse-voice-transcribe-1.0: 0.83% semantic WER at 392 ms TTFS
- 16 of 23 models are beaten on both speed and accuracy by frontier models
Read next
AI