NADI2026_subtask1.2_MixedASR_test
收藏资源简介:
该数据集名为 NADI2026_subtask2_MixedASR_test,是一个用于自动语音识别(ASR)任务的测试集,推测其与 NADI2026 评测任务(子任务2)相关,可能涉及混合语言或方言的语音识别场景。数据集包含 427 个音频样本,仅提供测试分割。每个样本包含两个字段:id(样本的唯一标识符,字符串类型)和 audio(音频数据,采样率为 16000 Hz)。数据集的下载大小约为 100 MB,在磁盘上的大小约为 101.5 MB。该数据集适用于对混合语言或方言环境下的语音识别模型进行测试和评估。
The dataset is named NADI2026_subtask2_MixedASR_test, which is a test set for automatic speech recognition (ASR) tasks. It is presumed to be related to the NADI2026 evaluation task (Subtask 2) and may involve mixed-language or dialect speech recognition scenarios. The dataset contains 427 audio samples, with only the test split provided. Each sample includes two fields: id (the unique identifier of the sample, string type) and audio (audio data with a sampling rate of 16000 Hz). The approximate download size of the dataset is 100 MB, and its on-disk size is about 101.5 MB. This dataset is suitable for testing and evaluating speech recognition models in mixed-language or dialect environments.
数据集概述:NADI2026_subtask1.2_MixedASR_test
基本信息
- 数据集名称:NADI2026_subtask1.2_MixedASR_test
- 数据集地址:https://huggingface.co/datasets/UBC-NLP/NADI2026_subtask1.2_MixedASR_test
数据集配置
- 默认配置名:
default
数据文件
- 测试集(test)数据文件路径:
data/test-*
数据集特征
- id:字符串类型(dtype: string)
- audio:音频数据类型,采样率为16000 Hz
数据划分
- 测试集(test)
- 样本数量:427 条
- 数据大小:101,520,786 字节(约96.8 MB)
- 数据文件下载大小:100,102,446 字节(约95.5 MB)
备注
- 该数据集卡片当前需要补充更多信息(More Information needed),详细内容可参考:https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards




