NADI2026_subtask1_ASR
收藏资源简介:
NADI2026_subtask1_ASR数据集是一个用于阿拉伯语自动语音识别(ASR)任务的多方言语音数据集,专为NADI2026竞赛的子任务1设计。数据集包含来自八个不同阿拉伯语国家或地区的语音样本,分别为:阿尔及利亚、埃及、约旦、毛里塔尼亚、摩洛哥、巴勒斯坦、阿联酋和也门。每个地区的数据作为一个独立的配置(config)提供。数据集由音频文件及其对应的文本转录组成,每个样本包含三个字段:唯一标识符(id)、音频数据(audio,采样率为16kHz)和文本转录(transcription)。所有配置均划分为训练集和验证集,训练集固定包含1600个样本,验证集样本数量因地区而异,范围从727到1600个。该数据集适用于训练和评估针对特定阿拉伯语方言的语音识别模型,支持多方言ASR研究。
The NADI2026_subtask1_ASR dataset is a multi-dialectal speech dataset for Arabic automatic speech recognition (ASR) tasks, specifically designed for Subtask 1 of the NADI2026 competition. The dataset includes speech samples from eight distinct Arabic-speaking countries/regions: Algeria, Egypt, Jordan, Mauritania, Morocco, Palestine, the United Arab Emirates, and Yemen. Data from each region is provided as an independent configuration (config). The dataset comprises audio files paired with their corresponding text transcriptions, with each sample containing three fields: a unique identifier (id), audio data (audio, with a sampling rate of 16 kHz), and the text transcription (transcription). All configurations are split into training and validation sets. The training set consistently contains 1600 samples, while the number of validation samples varies by region, ranging from 727 to 1600. This dataset is suitable for training and evaluating speech recognition models tailored to specific Arabic dialects, and supports multi-dialect ASR research.
数据集概述:NADI2026_subtask1_ASR
该数据集用于NADI2026竞赛的子任务1,聚焦于阿拉伯语方言的自动语音识别(ASR)。
数据集构成
数据集包含8个子集(config),每个子集对应一个阿拉伯语国家或地区的方言,分别为:阿尔及利亚(Algeria)、埃及(Egypt)、约旦(Jordan)、毛里塔尼亚(Mauritania)、摩洛哥(Morocco)、巴勒斯坦(Palestine)、阿联酋(UAE)和也门(Yemen)。
数据特征
每个数据样本包含三个字段:
- id(string):样本的唯一标识符。
- audio(audio):音频数据,采样率为16000 Hz。
- transcription(string):音频对应的转录文本。
数据拆分
每个子集均划分为训练集(train)和验证集(validation),具体样本数量如下:
| 子集 | 训练集样本数 | 验证集样本数 |
|---|---|---|
| Algeria | 1600 | 727 |
| Egypt | 1600 | 1600 |
| Jordan | 1600 | 1600 |
| Mauritania | 1600 | 1600 |
| Morocco | 1600 | 1600 |
| Palestine | 1600 | 900 |
| UAE | 1600 | 1600 |
| Yemen | 1600 | 1183 |
数据规模
数据集的总下载大小约为 3.12 GB,各子集的详细大小如下:
| 子集 | 下载大小 | 数据集大小 |
|---|---|---|
| Algeria | 297.06 MB | 297.11 MB |
| Egypt | 422.63 MB | 422.70 MB |
| Jordan | 426.50 MB | 426.56 MB |
| Mauritania | 370.78 MB | 370.85 MB |
| Morocco | 365.50 MB | 365.59 MB |
| Palestine | 415.10 MB | 415.17 MB |
| UAE | 434.09 MB | 434.16 MB |
| Yemen | 387.87 MB | 387.93 MB |
数据文件路径
每个子集的数据文件存储在对应的文件夹中,训练集文件匹配 train-* 模式,验证集文件匹配 validation-* 模式。例如,阿尔及利亚子集的训练文件位于 Algeria/train-*。
用途
该数据集专为多方言阿拉伯语语音识别任务设计,可用于训练和评估针对不同阿拉伯语方言的ASR模型。




