nadi26_subtask_5_SLU_test
收藏资源简介:
NADI 2026 SLU测试集是NADI 2026竞赛中Subtask 5: Spoken Language Understanding(口语理解)的官方测试数据集。该数据集包含一个名为clean的分割,共计989个样本,每个样本由两个字段构成:audio字段存储为直接嵌入在Parquet文件中的音频数据;id字段为样本的唯一标识字符串。数据集专为口语理解相关任务(如语音指令的语义解析)的评估而设计,适用于研究人员和开发者用于模型测试与基准评估。
The NADI 2026 SLU test set is the official test dataset for Subtask 5: Spoken Language Understanding in the NADI 2026 competition. It contains a split named clean with a total of 989 samples. Each sample consists of two fields: the audio field stores audio data directly embedded in Parquet files, and the id field is a unique identifier string for the sample. This dataset is designed for evaluating spoken language understanding tasks, such as semantic parsing of voice commands, and is suitable for researchers and developers for model testing and benchmark evaluation.




