SRUTI
收藏资源简介:
SRUTI是一个为印度农村博杰普里语妇女设计的语音识别基准数据集,包含约64.8小时的语音数据,其中17小时已转录。数据集涵盖健康、农业、治理和金融等关键领域,旨在促进农村妇女的数字包容性。数据收集面临诸多挑战,包括信任障碍、社会规范、人口统计学考虑、数据收集提示设计、社区参与和意识、现场数据收集以及转录等。最终,SRUTI数据集的创建为低资源语言的语音识别系统训练提供了宝贵的资源。
SRUTI is a speech recognition benchmark dataset designed for rural Bhojpuri-speaking women in India. It contains approximately 64.8 hours of speech data, among which 17 hours have been transcribed. The dataset covers key domains including health, agriculture, governance and finance, aiming to promote digital inclusion for rural women. Numerous challenges were encountered during data collection, including trust barriers, social norms, demographic considerations, data collection prompt design, community engagement and awareness, on-site data collection and transcription. Ultimately, the creation of the SRUTI dataset provides a valuable resource for training speech recognition systems for low-resource languages.




