URO-Bench
收藏资源简介:
URO-Bench是由上海交通大学MoE Key Lab of Artificial Intelligence和X-LANCE Lab提出的一种全面评估端到端语音对话模型的数据集。该数据集包含基础轨道和高级轨道两个难度级别,共有36个测试集,覆盖了语音对话场景中的多语言、多轮对话和副语言等方面,旨在评估模型在理解、推理和口语对话三个维度的能力。
URO-Bench is a comprehensive benchmark dataset for evaluating end-to-end spoken dialogue models, proposed by the MoE Key Lab of Artificial Intelligence and X-LANCE Lab at Shanghai Jiao Tong University. This dataset includes two difficulty levels, namely the Basic Track and the Advanced Track, with a total of 36 test sets. It covers multiple key aspects of spoken dialogue scenarios, including multilingualism, multi-turn dialogue, and paralinguistic cues. The dataset aims to evaluate a model's capabilities across three dimensions: comprehension, reasoning, and spoken dialogue interaction.




