UBC-NLP/NADI_2026_ADI20_micro
收藏资源简介:
这是一个小型版本的阿拉伯语方言识别数据集,基于ADI-20/ADI-17数据集,专为NADI 2026共享任务设计。数据集包含20种阿拉伯语方言的音频数据,每种方言提供10小时的录音,这些录音源自原始ADI-17数据集以及ADI-20中新增的方言。它旨在用于训练阿拉伯语方言识别模型,覆盖的方言包括:MSA(现代标准阿拉伯语)、BAH(巴林)、TUN(突尼斯)、ALG(阿尔及利亚)、EGY(埃及)、IRA(伊拉克)、JOR(约旦)、KSA(沙特阿拉伯)、KUW(科威特)、LEB(黎巴嫩)、LIB(利比亚)、MAU(毛里塔尼亚)、MOR(摩洛哥)、OMA(阿曼)、PAL(巴勒斯坦)、QAT(卡塔尔)、SUD(苏丹)、SYR(叙利亚)、UAE(阿联酋)和YEM(也门)。
This is a smaller version of the ADI-20/ADI-17 Arabic dialect identification dataset, designed for the NADI 2026 shared task. The dataset consists of audio data from 20 Arabic dialects, with 10 hours of recordings per dialect, sourced from the original ADI-17 dataset and the new dialects introduced in ADI-20. It is intended for training Arabic dialect identification models and covers the following dialects: MSA, BAH, TUN, ALG, EGY, IRA, JOR, KSA, KUW, LEB, LIB, MAU, MOR, OMA, PAL, QAT, SUD, SYR, UAE, YEM.




