yanyan666/SpeechEval
收藏资源简介:
SpeechEval是一个大规模多语言数据集,用于通用、可解释的语音质量评估。该数据集包含32,207个独特的语音片段和128,754个人工验证的标注,涵盖英语、中文、日语和法语四种语言。它结合了结构化标签和丰富的自然语言解释,适用于经典监督学习和语音大语言模型的指令调优。数据集覆盖四种核心评估任务:语音质量评估(提供单话语的自由形式、多方面的描述)、语音质量比较(对两个话语进行成对比较,包括决策和理由)、语音质量改进建议(为次优话语提供可操作的改进建议)以及深度伪造语音检测(在质量相关上下文中分类语音为人类或合成/操纵)。数据集结构包括音频文件夹(按语言组织)和元数据文件(以JSONL格式存储),总分割大小包括训练集73,123、验证集20,501和测试集35,130。许可证为CC BY-NC-SA 4.0。
SpeechEval is a large-scale multilingual dataset for general-purpose, interpretable speech quality evaluation. It contains 32,207 unique speech clips and 128,754 human-verified annotations across English, Chinese, Japanese, and French languages. Each example combines structured labels and rich natural-language explanations, making it suitable for both classic supervised learning and instruction-tuning of SpeechLLMs. The dataset covers four core evaluation tasks: Speech Quality Assessment (free-form, multi-aspect descriptions for a single utterance), Speech Quality Comparison (pairwise comparison of two utterances with decision and justification), Speech Quality Improvement Suggestion (actionable suggestions to improve a suboptimal utterance), and Deepfake Speech Detection (classify speech as human vs synthetic/manipulated, with quality-related context). The directory structure includes audio folders organized by language and metadata files in JSONL format. Total split sizes are: Train: 73,123, Validation: 20,501, Test: 35,130. License: CC BY-NC-SA 4.0.



