AI SPEECH / ANALYSIS

SraVaani 1.0 needs a real accent test, not a headline score

The open Indian-language speech model covers 65 languages and dialects. Before adopting it, build a small test from the voices, noise and names your product actually handles.

WASSAI editorial desk ·

SraVaani research diagram showing speech and image encoders learning from paired audio and images before speech recognition fine-tuning.
Official SraVaani research diagram. The team used paired speech and images during pretraining, then fine-tuned the speech encoder for recognition. · Image source

ARTPARK at the Indian Institute of Science published SraVaani 1.0 in August, an open automatic-speech-recognition model for 65 Indian languages and dialects. Its authors trained the speech encoder from scratch on 29,912 hours of unlabelled audio across 105 languages, then fine-tuned it on labelled speech. The release is interesting not because one model solves Indian speech, but because it gives builders another testable baseline for voices often poorly represented in mainstream benchmarks. [1] [2]

Vision helped the model learn speech

Project Vaani collected spontaneous speech by asking participants to describe images. SraVaani reused that pairing: an audio encoder and an image encoder learned to bring matching speech and pictures closer in a shared representation, without requiring new transcriptions. The research team reports 11.8 million audio-image pairs, covering 16,580 hours, in this alignment stage. [1] [2]

On a held-out Vaani test covering 48 languages, the authors report that this visual alignment reduced average word error rate from 28.09 to 27.40. Results improved for 37 languages, with statistically significant gains for 20 and no significant regressions. Those are research results on a defined test set, not a promise about every microphone, accent or use case. [1]

Build a test from your real workflow

Before switching a subtitle or voice product, assemble 40 to 60 consented clips from the conditions it actually encounters. Include clean studio speech, phone microphones, fan or traffic noise, code-switching, regional accents, proper names and spoken numbers. Make human-checked transcripts, then run exactly the same files through SraVaani and your current baseline.

Track word or character error rate, but keep separate counts for names, numbers and language switches. Record latency, memory use and failed outputs as well. For captioning, inspect punctuation, casing and timestamps explicitly: the published training pipeline normalised text to lowercase and removed punctuation and digits, so product-ready formatting needs its own evaluation. [1]

Treat openness as an invitation to measure

The authors made the model available on Hugging Face and published a repository with fine-tuning and baseline-evaluation scripts. Coverage is still uneven: 65 languages and dialects is substantial, but the documented training hours vary from more than 1,000 for some languages to less than one for others. A good decision therefore comes from a per-language scorecard, not one blended average. [1]

Publish the test set design, record the model version and review mistakes with fluent speakers. If SraVaani wins only on clean clips but loses names in noisy, code-switched speech, that is still a useful result. Open speech models become valuable when teams can reproduce where they work, where they fail and which voices need better data next.

Sources

  1. ARTPARK-IISc: SraVaani — How Vision Helps Hearing · Published 2026-08-11; checked 5 October 2026.
  2. VAANI: Capturing the language landscape for an inclusive digital India · Published 2026-03-30; checked 5 October 2026.