RegSpeech12
收藏资源简介:
RegSpeech12是一个包含孟加拉语方言的自发语音数据集,旨在记录和分析这些方言的语音和形态学特性,同时探索构建适合地区变体的计算模型(特别是自动语音识别系统)的可行性。该数据集包含来自孟加拉国不同地区的215位说话者的录音,涵盖了64个不同的主题,包括教育、家庭生活、经济、体育和政治等。数据集的收集和验证过程严格遵循特定协议,以确保数据的多样性和自然性。
RegSpeech12 is a spontaneous speech dataset focused on Bengali dialects. It aims to document and analyze the phonetic and morphological properties of these dialects, while exploring the feasibility of developing computational models—particularly automatic speech recognition systems—tailored to regional language variants. The dataset includes recordings from 215 speakers across various regions of Bangladesh, covering 64 distinct topics such as education, family life, economy, sports, politics, and more. The collection and validation workflows of the dataset strictly follow specific protocols to ensure the diversity and naturalness of the collected data.

- 1RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across DialectsBRAC University · 2025年



