Illustration by tuput
English
Sarvam's Saaras V3 scored about 19% word error on the IndicVoices test in February 2026, and Meta's open model lists a 3.6 character error rate for Hindi. The only benchmark built on Indian phone calls keeps its test set private, and Google's and OpenAI's own pages for their newest speech products give language counts but no Indian-language error rates.
Sarvam AI’s Saaras V3 speech model scored about 19% word error on the IndicVoices benchmark, according to Sarvam’s own post of 11 February 2026. Meta’s open Omnilingual ASR model, released on 10 November 2025, lists a character error rate of 3.6 for Hindi in its results table. The two numbers measure different things on different audio, and most of the other published scores are just as hard to line up.
Word error rate (WER) is the share of words a system gets wrong against a human transcript. Character error rate (CER) counts wrong characters instead. Indian languages add a further problem. Researchers at AI4Bharat, the language-technology lab at IIT Madras, wrote in a 1 March 2026 preprint that flexible spelling and optional splitting or joining of words make plain WER rate systems worse than users find them.
Sarvam’s 22-language models
Saaras V3 handles 22 official Indian languages plus English in one model. Sarvam’s post gives about 19% WER on IndicVoices for all languages and 19.31% for the ten most widely spoken. It says the earlier Saaras V2.5, which covered 11 languages, scored about 22% on the same set.
IndicVoices is an AI4Bharat dataset. Its 2024 paper describes 7,348 hours of speech from 16,237 speakers in 145 districts across 22 languages, of which 1,639 hours were transcribed at the time. Sarvam’s post names GPT-4o Transcribe, Gemini 3 Pro, Deepgram Nova3 and ElevenLabs Scribe v2 as rivals it beats, but it prints no scores for them. The post does not name an outside evaluator.
Saaras V4 followed on 25 September 2026. Sarvam reports 16.03% WER on IndicContextEval, an AI4Bharat test where the user supplies key terms, and a language-identification error of 5.22% across all 22 languages and 2.9% across the top ten. Most V4 accuracy results appear only as charts. The synchronous API takes audio under 30 seconds, and batch jobs take files up to two hours.
Sarvam Audio, announced on 3 February 2026, is a separate model, an audio extension of Sarvam’s 3-billion-parameter language model. The post compares it with GPT-4o Transcribe and Gemini 3 Flash on IndicVoices without giving numbers, and says it was not yet generally available.
The March 2026 AI4Bharat preprint, which has a Sarvam researcher among its authors and has not been peer reviewed, measured how much spelling variants distort the scores. Its orthographically informed WER (OIWER) accepts any listed spelling of a word. It came out 6.3 points lower on average for one model, and the average gap was 14.4 points for Tamil and 13.9 for Malayalam against about 5.3 for Marathi and Gujarati.
Meta’s 1,600 languages
Omnilingual ASR covers more than 1,600 languages, and Meta says over 500 of them had never been transcribed by AI. Models run from 300 million to 7 billion parameters. Code and models are under the Apache 2.0 licence, and the paper’s decoder borrows from large language models. Meta’s own headline is a CER below 10 for 78% of languages with its 7-billion-parameter model.
The project’s results table on GitHub gives the Indian picture. Hindi scores 3.6 CER after 236.7 hours of training audio, Tamil 5.2 after 379.3 hours, Bengali 2.8 and Odia 5.8 after 51.8 hours. Bhojpuri has 5.5 training hours and a CER of 16.5. Santali is at 4.9 on 0.4 hours. The table has no column naming the test audio.
Two limits are on record. Meta says teaching the model a new language from a few examples is less accurate than full training. And the GitHub page says most checkpoints reject clips of 40 seconds or longer, though a December 2025 update added “Unlimited” variants for longer audio.
For translation, Meta’s Omnilingual MT paper of 17 March 2026 describes what the authors call the first system supporting more than 1,600 languages. Its 1 billion to 8 billion parameter models match or beat a 70-billion-parameter baseline, the authors say. They also report that baseline models often understand a low-resource language but generate it poorly. tuput found no per-language Indian scores in the abstract or the pages it opened.
The one test built on Indian phone calls
Voice of India, run by Josh Talks with AI4Bharat and accepted at Interspeech 2026, uses 536 hours of unscripted telephone speech from 36,691 speakers in 15 languages. It scores with OI-WER, a lattice of valid spellings. The core test set is private, so outsiders cannot reproduce the results, and only a small public sample is planned.
Josh Talks’s April 2026 write-up scores 15 systems in its table. These are OI-WER percentages, lower is better.
| Model | Hindi | Bengali | Tamil |
|---|---|---|---|
| Sarvam Audio | 5.0 | 6.0 | 14.2 |
| Gemini 3 Pro | 6.0 | 8.5 | 15.7 |
| ElevenLabs Scribe v2 | 7.7 | 10.0 | 20.4 |
| IndicConformer (AI4Bharat) | 8.2 | 10.7 | 19.9 |
| OmniASR LLM 7B (Meta) | 13.7 | 22.8 | 43.1 |
| GPT-4o Transcribe | 34.0 | 44.8 | 64.2 |
Sarvam Audio, a Sarvam product, ranks first on all three. GPT-4o Transcribe is an OpenAI model from March 2025. TechCrunch reported at its launch, citing OpenAI’s internal benchmark, a word error rate approaching 30% for Tamil, Telugu, Malayalam and Kannada. That is a different test from Voice of India’s.
The write-up lists its own limits. Hypotheses behind the scoring lattices come from an internal ensemble of recognition models. Named entities, domain terms and numbers are not yet tested. Male speakers show a relative error penalty of about 19 to 21% against female speakers.
Bodhan’s September release and numbers that do not match
In September 2026 Bodhan AI, an IIT Madras-incubated group working with AI4Bharat and NVIDIA, released Indic-Transcribe, open models for speech recognition. The Core model’s card reports an average OIWER of 8.7 on Voice of India, against 10.7 for Saaras V3, 17.8 for IndicConformer and 21.1 for Gemini 3 Pro. It covers 25 languages: English, the 22 scheduled languages, Bhojpuri and Bhili. The card does not say who ran the evaluation.
The card lists its limits. It was trained on clips of up to 30 seconds, outputs native script only, and assumes one speaker. Automatic language identification is uneven, and Bhojpuri, Maithili and Urdu are often absorbed into Hindi.
The Bodhan card and the Josh Talks page also disagree on the same models. The card gives Gemini 3 Pro 9.3 for Hindi and OmniASR LLM 7B 9.6, where the Josh Talks table has 6.0 and 13.7. Neither page explains the difference, so scores from the two cannot be pooled.
Google and OpenAI publish language counts
Google’s Chirp 3 documentation, updated on 7 October 2026, lists Hindi and Indian English as generally available. Ten more Indian languages are in preview, namely Assamese, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil and Telugu. Urdu does not appear. The page gives no error rates.
Google’s Gemini 3.5 Live Translate, announced on 9 June 2026, translates speech from 70 or more languages and aims to keep the speaker’s pacing and pitch. Google Meet’s speech translation grows from five languages to 70 or more, and from English-only pairs to over 2,000 language combinations in one meeting. The post says the translation stays a few seconds behind the speaker and prints no benchmark figures or Indian language names.
OpenAI’s gpt-realtime-translate, reported in May 2026, accepts speech in more than 70 languages, including Hindi, Bengali, Gujarati, Malayalam, Punjabi and Telugu in its cookbook list. Speech output comes in 13 languages, and Hindi is the only Indian one. OpenAI’s cookbook warns that the model can substitute wrong names or entities, and it gives no latency figures.
Bhashini, the government platform covered in India’s push for AI in its own languages, also offers speech services. tuput found no published error rates for its models, so it is not scored here. The Supreme Court’s own use of AI for transcription and translation is set out in AI in Indian courts.
The Josh Talks write-up lists named-entity, domain-specific and numeric evaluation as future work, so no score above covers names, jargon or digits.
Sources & further reading
- Meta AI: Omnilingual ASR, advancing automatic speech recognition (10 November 2025)
- arXiv: Omnilingual ASR, Open-Source Multilingual Speech Recognition for 1600+ Languages (12 November 2025)
- GitHub: facebookresearch/omnilingual-asr, code, model list and 7B per-language results table
- arXiv: Omnilingual MT, Machine Translation for 1,600 Languages (17 March 2026)
- Sarvam AI: Saaras V3 speech recognition (11 February 2026)
- Sarvam AI: Introducing Saaras V4 (25 September 2026)
- Sarvam AI: Sarvam Audio (3 February 2026)
- arXiv: IndicVoices, Towards building an Inclusive Multilingual Speech Dataset for Indian Languages (4 March 2024)
- arXiv: Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages (1 March 2026)
- arXiv: Voice of India, A Large-Scale Benchmark for Real-World Speech Recognition in India (21 April 2026)
- Josh Talks AI: Voice of India, Building India's National ASR Benchmark (April 2026)
- Hugging Face: Bodhan AI Indic-Transcribe-core model card (September 2026)
- DT Next: IIT-M-incubated Bodhan AI unveils open models for Indian languages (12 September 2026)
- Google Cloud: Chirp 3 transcription model documentation (last updated 7 October 2026)
- Google: Fluid, natural voice translation with Gemini 3.5 Live Translate (9 June 2026)
- OpenAI Cookbook: Build Live Translation Apps with gpt-realtime-translate
- Quantum Zeitgeist: OpenAI's New Model Translates 70+ Languages in Realtime (11 May 2026)
- TechCrunch: OpenAI upgrades its transcription and voice-generating AI models (20 March 2025)
Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.
Enjoyed this? Get the next one.
One good read at a time, straight to your inbox. No spam, unsubscribe anytime.