- IISc SPIRE Lab has released the SraVaani speech AI model, which covers 65 Indian languages.
- Project Vaani recorded over 31,000 hours of speech from 156,000 people across 165 districts in 28 states and three Union Territories.
- SraVaani was trained on the complete Vaani speech corpus, followed by audio-image alignment using 11.8 million Vaani audio-image pairs.
- SraVaani was further trained on transcribed speech from Vaani and other public Indian speech datasets including IndicVoices, RESPIN, SPRING-INX, SPICOR, and SYSPIN.
- The model produces text across 10 different scripts and can automatically identify the language being spoken, eliminating the need for a language tag to be specified in advance.
IISc's SPIRE Lab has unveiled SraVaani, a groundbreaking speech AI model that recognizes 65 Indian languages, trained on over 31,000 hours of data from Project Vaani. This initiative, supported by ARTPARK and Google, aims to enhance accessibility for approximately 25 crore people whose languages are often overlooked by existing systems.1234
SraVaani encompasses 20 scheduled languages and 45 regional dialects, ensuring comprehensive coverage across India. It includes languages from various regions, such as Garo, Angika, Chakma, Kokborok, Tulu, Bundeli, and Bajjika. The model is publicly available on Hugging Face under an MIT license, allowing developers to experiment and adapt it.
In terms of performance, SraVaani has been evaluated against eight public benchmark datasets, achieving an impressive 9.5% word error rate on Garo, significantly better than the next-best system at 69.4%. The model's training involved a diverse dataset, with speakers describing images in their own words, capturing natural speech and regional variations. This innovative approach positions SraVaani as a leader in the field of speech recognition in India.
The model's ability to produce text across 10 different scripts and automatically identify spoken languages without prior tagging further enhances its usability and accessibility.
“SraVaani was trained on 11.8 million audio-image pairs and can output text in 10 scripts while auto-identifying the spoken language. The underlying Project Vaani corpus spans 165 districts across 28 states and three Union Territories, capturing natural speech from 156,000 speakers.”










