NADI2026_subtask2_MixedASR
收藏资源简介:
该数据集名为“NADI2026_subtask2_MixedASR”,推测为NADI2026竞赛的第二个子任务(混合自动语音识别)相关。数据集包含一个开发集(dev),共3152个样本。每个样本包含三个字段:唯一标识符(id)、音频数据(audio,采样率为16000Hz)以及对应的文本转录(transcription)。数据形式为音频-文本配对,适用于自动语音识别(ASR)任务的模型训练或评估。数据集总大小约为838.7MB。
The dataset is named "NADI2026_subtask2_MixedASR", which is presumed to be related to the second subtask (Mixed Automatic Speech Recognition) of the NADI2026 competition. The dataset includes a development set (dev) with a total of 3,152 samples. Each sample contains three fields: a unique identifier (id), audio data (audio, with a sampling rate of 16000 Hz), and the corresponding text transcription (transcription). The data is in the form of audio-text pairs, suitable for training or evaluating models for automatic speech recognition (ASR) tasks. The total size of the dataset is approximately 838.7 MB.
- 数据集名称:NADI2026_subtask2_MixedASR
- 数据集链接:https://huggingface.co/datasets/UBC-NLP/NADI2026_subtask2_MixedASR
- 数据集描述:该数据集用于NADI2026子任务2,涉及混合自动语音识别(ASR)任务。
- 配置:
default - 数据文件:开发集(dev)数据位于
data/dev-*路径下。 - 数据集特征:
id:字符串类型,样本标识符。audio:音频数据,采样率为16000 Hz。transcription:字符串类型,转录文本。
- 数据集划分:
- 开发集(dev):共3152个样本,数据集大小为838,679,480字节(约800 MB),下载大小为819,946,622字节。




