NADI2026_subtask1.3_codeswitched_asr
收藏资源简介:
该数据集是一个面向音频转录任务的数据集,包含音频文件及其对应的文本转录。数据集中定义了两个核心特征字段:audio字段存储音频数据,transcription字段存储对应的文本转录字符串。数据集仅提供一个测试划分(test split),共包含2,123个样本,总数据大小约为234 MB,下载大小约为210 MB。数据文件存储路径遵循data/test-*的模式。该数据集适用于语音识别、音频内容转录等相关任务的研究与评估。
This dataset is tailored for audio transcription tasks, comprising audio files and their corresponding text transcriptions. Two core feature fields are defined in this dataset: the 'audio' field stores the audio data, and the 'transcription' field stores the corresponding text transcription string. The dataset only provides one test split, containing a total of 2,123 samples. The overall data size is approximately 234 MB, and the download size is around 210 MB. The data files follow the storage path pattern 'data/test-*'. This dataset is suitable for research and evaluation of related tasks such as speech recognition and audio content transcription.
数据集概述
该数据集名为 NADI2026_subtask1.3_codeswitched_asr,由 UBC-NLP 提供,主要用于代码混合语音的自动语音识别(ASR)任务。
数据特征
- audio:音频数据,类型为
audio。 - transcription:对应的文本转录,类型为
string。
数据划分
- 测试集(test):包含 2,123 个样本,总大小为 233,594,856 字节。
文件信息
- 配置名称:
default - 数据文件路径:
data/test-*,下载大小为 209,691,691 字节。




