遇见数据集

Rasa

收藏
魔搭社区2026-07-20 更新2026-08-23 收录
官方服务:

资源简介:

# Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages **Funded by**: Bhashini, Ministry of Electronics and Information Technology, Government of India **Supported by**: EkStep Foundation and Nilekani Philanthropies ## Overview We introduce **Rasa**, the first high-quality **multilingual expressive Text-to-Speech (TTS) dataset** for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering a female and male voice for each of the 22 officially recognized languages of India. In our initial version, we explore a practical recipe for collecting high-quality data for resource-constrained languages, prioritizing easily obtainable neutral data alongside smaller amounts of expressive data. This approach enables us to extend our dataset to encompass a diverse array of speaking styles and contexts. These include neutral readings from Wikipedia and IndicTTS texts, expressive speech capturing the six Ekman emotions (happy, sad, angry, fear, disgust, and surprise), as well as command-based interactions from platforms like Alexa, BigBasket, UMANG, and DigiPay. Additionally, Rasa includes natural conversations on various topics, news-reading, and narration from book readings. Currently, we release the data for 44 speaker-language pairs across all 22 Indian languages. Through this release, we aim to provide a valuable resource for developing expressive TTS models in multilingual settings for the officially recognized languages of India. ### Key Features - **Multilingual Coverage**: Covers diverse Indian languages - **Expressive Speech**: Includes **Ekman emotions** (happy, sad, angry, fear, disgust, and surprise) - **Multiple Speaking Styles**: - Neutral speech from Wikipedia texts - Command-based interactions from Alexa, BigBasket, UMANG, and DigiPay - Natural conversations on various topics - News reading and narration from book readings - **High-Quality Data**: 48 KHz, Mono - **Current Release**: 44 speaker-language pairs available now in all 22 Indian languages. Through this release, we aim to provide a **valuable resource for multilingual expressive TTS models**, helping advance text-to-speech synthesis for Indian languages. --- ## Dataset Statistics | Language | Speaker | Hours | Utterances | |-----------|---------|---------|------------| | Assamese | Female | 29.09 | 15,084 | | Assamese | Male | 29.23 | 16,614 | | Bengali | Female | 29.69 | 15,570 | | Bengali | Male | 27.06 | 15,639 | | Bodo | Female | 27.32 | 16,329 | | Bodo | Male | 24.99 | 13,162 | | Dogri | Female | 25.69 | 13,177 | | Dogri | Male | 22.35 | 10,332 | | Gujarati | Female | 25.39 | 12,151 | | Gujarati | Male | 26.01 | 15,799 | | Hindi | Female | 27.05 | 15,107 | | Hindi | Male | 23.78 | 13,464 | | Kannada | Female | 27.02 | 14,913 | | Kannada | Male | 27.59 | 15,999 | | Kashmiri | Female | 10.95 | 6,653 | | Kashmiri | Male | 25.34 | 16,059 | | Konkani | Female | 26.33 | 17,585 | | Konkani | Male | 25.43 | 19,324 | | Maithili | Female | 21.79 | 12,446 | | Maithili | Male | 29.34 | 12,917 | | Malayalam | Female | 26.42 | 16,974 | | Malayalam | Male | 25.22 | 16,877 | | Manipuri | Female | 27.49 | 13,516 | | Manipuri | Male | 23.55 | 12,579 | | Marathi | Female | 28.80 | 15,472 | | Marathi | Male | 26.55 | 14,483 | | Nepali | Female | 28.74 | 16,016 | | Nepali | Male | 26.36 | 15,239 | | Odia | Female | 24.07 | 11,757 | | Odia | Male | 23.85 | 13,611 | | Punjabi | Female | 26.72 | 13,420 | | Punjabi | Male | 29.12 | 15,620 | | Sanskrit | Female | 25.27 | 12,757 | | Sanskrit | Male | 26.20 | 11,002 | | Santali | Female | 27.77 | 14,469 | | Santali | Male | 25.04 | 13,513 | | Sindhi | Female | 24.61 | 13,588 | | Sindhi | Male | 27.58 | 15,659 | | Tamil | Female | 29.94 | 19,871 | | Tamil | Male | 26.18 | 16,790 | | Telugu | Female | 27.28 | 15,404 | | Telugu | Male | 24.98 | 15,006 | | Urdu | Female | 26.23 | 13,980 | | Urdu | Male | 26.03 | 15,023 | |-----------|---------|---------|------------| | **Total** | | 1145.44 | 640,950 | | 1035.51 | 582,199 | --- ## License CC-BY-4.0 ## Citation If you use this dataset, please cite: ```bibtex @inproceedings{ai4bharat2024rasa, author={Praveen Srinivasa Varadhan and Ashwin Sankar and Giri Raju and Mitesh M. Khapra}, title={{Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings}}, year=2024, booktitle={Proc. INTERSPEECH 2024}, } ```

提供机构:
maas
创建时间:
2025-12-24
二维码
社区交流群
二维码
科研交流群
商业服务