Sarvam AI Saaras V4: How the Multilingual Speech Model Is Built for India’s Diverse Languages
Sarvam AI, an Indian startup, released Saaras V4, a multilingual speech model designed for India's diverse languages. The model uses a 3-billion-parameter decoder trained from scratch to handle code-mixed speech and regional accents. Sarvam reports state-of-the-art performance across all 22 supported Indian languages. It also posts the lowest average error rate in English across seven public benchmarks.
Saaras V4's value is in its direct training on Indian speech patterns. Most ASR tools adapt models built for English first. Sarvam's choice to train its 3-billion-parameter decoder from scratch gives it full control. This approach delivers better performance across India's 22 official languages.
The key term prompting feature is a direct play for enterprise customers. Developers can supply up to 50 specific terms. This boosts recognition accuracy for brand names and technical jargon. This directly addresses a pain point for call centers and support platforms in India.
The real test for Sarvam AI is market adoption beyond early adopters. Its 150-millisecond response time for the first output token supports real-time streaming. This speed is critical for voice assistants and live captioning tools. The question is whether enterprises will switch from established ASR providers.
Related reading
6 stories
Lightspeed targets $250M for new India fund, focusing on early-stage AI
Lightspeed's new India fund focuses on early-stage AI, highlighting the investment landscape for companies like Sarvam AI.
Make In India @12: The Numbers Behind The Manufacturing Push

Mantle Hits Back-to-Back All-Time Highs with 1,473 Tokenized Assets and $476M in Distributed Asset Value

Decentralisation must not lead to fragmented responsibility: NA President

Pine Labs Taps Google Cloud to Build Agentic AI Infrastructure

