ArabicSpeech/ADI20
收藏资源简介:
ADI-20是一个阿拉伯语方言识别数据集和模型集合,它是先前发布的ADI-17数据集的扩展。该数据集覆盖了所有阿拉伯语国家的方言,包含3,556小时的音频数据,来自19种阿拉伯方言以及现代标准阿拉伯语(MSA),用于训练和评估最先进的阿拉伯语方言识别系统。扩展内容包括:加入了MSA以区分地区方言和标准化阿拉伯语;新增了突尼斯和巴林方言,其中突尼斯方言基于整个TunSwitch数据集(包括代码切换子集),总时长达142.7小时,并补充了YouTube内容以达到每个分割2小时;巴林方言则完全来自YouTube,训练数据约271小时,验证和测试各2小时。数据集需在下载ADI17后使用。
ADI-20 is a collection of Arabic dialect identification datasets and models, which is an extension of the previously released ADI-17 dataset. This dataset covers dialects from all Arabic-speaking countries, containing 3,556 hours of audio data sourced from 19 Arabic dialects and Modern Standard Arabic (MSA), and is used for training and evaluating state-of-the-art Arabic dialect identification systems. The extensions include: adding MSA to distinguish regional dialects from standardized Arabic; newly adding Tunisian and Bahraini dialects. Specifically, the Tunisian dialect is based on the entire TunSwitch dataset (including its code-switching subset), with a total duration of 142.7 hours, and YouTube content is supplemented to ensure 2 hours per data split; the Bahraini dialect is entirely sourced from YouTube, with approximately 271 hours of training data, and 2 hours each for validation and testing. This dataset can only be used after downloading ADI-17.




