voicebench-ja
收藏资源简介:
该数据集旨在定量评估语音语言模型在接收音频输入与文本输入时表现出的智能和推理能力差异。数据集由四个子集构成,这些子集基于三个文本基准(Elyza-tasks-100、M-IFEval和JamC-QA)的样本,并应用了语音合成技术。语音合成使用了SB Intuitions公司内部的TTS模型,并以JVS语料库的语音作为提示。数据集包含以下子集: 1. Elyza:从elyza-tasks-100中选取36个样本,经过文本修正和语音合成。 2. Spoken-Elyza:调整Elyza子集的参考文本,去除不适合语音传递的标记和符号,并进行人工听力验证,最终保留34个样本。 3. M-IFEval:将M-IFEval的输入提示转换为可朗读形式并合成语音,保留原始评估约束。 4. JamC-QA:对JamC-QA的多选题样本添加标签并合成语音,从2309个样本中筛选出1452个适合评估的样本。 数据集文本部分采用CC BY-SA 4.0许可,语音数据禁止商用和再分发。
This dataset aims to quantitatively evaluate the differences in intelligence and reasoning capabilities of speech language models when receiving audio inputs versus text inputs. The dataset comprises four subsets derived from samples of three text benchmarks (Elyza-tasks-100, M-IFEval, and JamC-QA) via speech synthesis technology. Speech synthesis was performed using an internal TTS model from SB Intuitions, with voices from the JVS corpus as prompts. The dataset includes the following subsets: 1. Elyza: 36 samples selected from Elyza-tasks-100, which underwent text correction and speech synthesis. 2. Spoken-Elyza: The reference texts of the Elyza subset were adjusted to remove markers and symbols unsuitable for oral transmission, followed by manual auditory verification, with 34 samples finally retained. 3. M-IFEval: The input prompts of M-IFEval were converted into orally readable forms and synthesized into speech, while the original evaluation constraints were preserved. 4. JamC-QA: Labels were added to the multiple-choice samples of JamC-QA, followed by speech synthesis; 1,452 eligible samples for evaluation were screened from the original 2,309 samples. The text portion of the dataset is licensed under CC BY-SA 4.0, while commercial use and redistribution of the audio data are prohibited.




