Persian Dialect IDentification (PDID)
收藏资源简介:
PDID数据集是一个多口音语料库,涵盖了10个地区波斯口音,为波斯语音识别中的口音变化提供了第一个系统性的基准,填补了多语言语音研究中的一个关键空白,并为未来低资源、语言多样性语言的研究提供了基础。数据集包含来自200多个小时的原始语音中约23小时的干净口音标注数据,样本被标准化为16kHz、单声道、16位WAV格式,并以3-30秒的片段进行分割。数据集的创建过程包括语音活动检测、说话人分割、基于静默的分割和语音-音乐分离等预处理步骤。该数据集旨在解决语音识别系统对音调和方言变化的敏感性,并通过抑制口音相关特征并鼓励语音识别模型学习口音中性表示来提高模型的鲁棒性。
The PDID dataset is a multi-accent corpus covering 10 regional Persian accents. It serves as the first systematic benchmark for accent variations in Persian speech recognition, filling a critical gap in multilingual speech research and laying a foundation for future studies on low-resource, linguistically diverse languages. The dataset contains approximately 23 hours of clean, accent-annotated speech data extracted from over 200 hours of raw speech. The samples are standardized to 16kHz, mono-channel, 16-bit WAV format, and segmented into 3-30 second clips. The dataset's creation includes preprocessing steps such as voice activity detection, speaker diarization, silence-based segmentation, and speech-music separation. This dataset aims to address the sensitivity of speech recognition systems to tonal and dialectal variations, and improve model robustness by suppressing accent-related features and encouraging speech recognition models to learn accent-neutral representations.




