Sarvam AI launches Saaras V4 multilingual speech recognition model
Indian startup Sarvam AI has released Saaras V4, an automatic speech recognition model built on a 3-billion-parameter hybrid state-space language model. It supports 22 Indian languages, code-mixed speech and noisy audio, offers five output formats and streams with under 150 ms time to first token.
- Language identification error rate is 5.22% across 22 Indian languages and 2.9% across the top 10
- Five output modes: verbatim, normalised, codemix, translit and translate
- Streaming recognition with time to first token under 150 ms
- Available via API with Python and Node.js SDKs, plus Vercel, LiveKit and Pipecat integrations
Read next
AI