NADI2026_subtask1.1_Robust_ASR_test
收藏资源简介:
该数据集是NADI2026竞赛子任务1.1(鲁棒自动语音识别测试)的专用测试集。数据集包含八个独立配置,分别对应八个不同的阿拉伯语国家或地区:阿尔及利亚、埃及、约旦、毛里塔尼亚、摩洛哥、巴勒斯坦、阿联酋和也门。每个配置专门用于评估自动语音识别模型在特定地域口音或方言上的鲁棒性。数据集仅包含测试集,每个配置包含500个音频样本。每个样本由两个字段构成:一个唯一的字符串标识符(id)和一个音频文件(audio),所有音频的采样率均为16000Hz。该数据集适用于多方言/多口音的阿拉伯语自动语音识别模型的性能评测与基准测试。
This dataset is a dedicated test set for the NADI2026 competition subtask 1.1 (Robust Automatic Speech Recognition Testing). It contains eight independent configurations, corresponding to eight different Arabic-speaking countries or regions: Algeria, Egypt, Jordan, Mauritania, Morocco, Palestine, UAE, and Yemen. Each configuration is specifically designed to evaluate the robustness of automatic speech recognition models on specific regional accents or dialects. The dataset only includes a test set, with each configuration containing 500 audio samples. Each sample consists of two fields: a unique string identifier (id) and an audio file (audio), with all audio sampled at 16000Hz. This dataset is suitable for performance evaluation and benchmarking of multi-dialect/multi-accent Arabic automatic speech recognition models.
数据集总览
- 数据集名称:NADI2026_subtask1.1_Robust_ASR_test
- 数据集链接:https://huggingface.co/datasets/UBC-NLP/NADI2026_subtask1.1_Robust_ASR_test
- 用途:该数据集是为NADI2026子任务1.1(鲁棒自动语音识别测试)设计的测试集。
数据子集与配置
数据集包含8个地区配置,每个配置对应一个阿拉伯语国家/地区的测试数据:
| 配置名称 | 测试样本数 | 测试集大小(字节) |
|---|---|---|
| Algeria | 500 | 72,457,316 |
| Egypt | 500 | 73,284,460 |
| Jordan | 500 | 68,236,940 |
| Mauritania | 500 | 62,880,236 |
| Morocco | 500 | 55,947,532 |
| Palestine | 500 | 89,750,652 |
| UAE | 500 | 69,841,764 |
| Yemen | 500 | 82,745,724 |
- 总计:每个地区500个测试样本,共4,000个测试样本。
数据特征
每个样本包含两个字段:
- id(字符串类型):样本的唯一标识符。
- audio(音频类型):音频数据,采样率为16,000 Hz。
数据划分
数据集仅包含一个划分:test(测试集)。每个子配置的测试数据文件路径格式为:{配置名}/test-*。
补充说明
当前数据集卡片内容尚不完整,更多信息可能需要参考贡献指南或数据集发布者的后续补充。




