EchoMind
收藏资源简介:
EchoMind是一个多层次的对话评估基准,旨在全面评估语音语言模型(SLMs)的共情能力。该数据集模拟了人类对话的认知过程,通过三个相互关联的阶段进行评估:理解(内容理解和语音感知)、推理(综合推理)和对话(开放式回复生成)。EchoMind的数据集包含1,137个对话脚本,每个脚本都有三种语音风格变体,用于测试语音表达对模型的影响。该数据集的构建旨在解决SLMs在理解和生成共情式对话方面的挑战,以促进更高级的语音语言模型的发展。
EchoMind is a multi-level dialogue evaluation benchmark designed to comprehensively evaluate the empathic capabilities of Spoken Language Models (SLMs). This dataset simulates the cognitive process of human dialogue, and conducts evaluations across three interrelated stages: comprehension (content understanding and speech perception), reasoning (comprehensive reasoning), and dialogue (open-ended response generation). The EchoMind dataset contains 1,137 dialogue scripts, each with three speech style variants to test the impact of vocal expression on models. This benchmark is constructed to address the challenges faced by SLMs in understanding and generating empathic dialogues, so as to promote the development of more advanced spoken language models.
EchoMind数据集概述
数据集基本信息
- 数据集名称: EchoMind
- 维护机构: 香港中文大学人类语言技术中心
- 许可证: CC-BY-NC-SA-4.0
- 联系方式: lizhou21@cuhk.edu.cn
数据集定位
EchoMind是首个相互关联的多阶段基准测试,专门用于评估语音语言模型的共情对话能力。该基准通过顺序的、上下文关联的任务模拟共情对话的认知过程。
核心特征
- 多级评估框架: 包含三个主要层级
- 层级1: 通过内容理解和语音感知进行评估
- 层级2: 整合内容和语音进行推理
- 层级3: 生成上下文和情感对齐的开放域响应
- 控制变量设计: 所有任务共享相同的语义中性脚本,同时通过受控的语音风格变化测试表达效果
- 共情导向框架: 涵盖3个粗粒度和12个细粒度维度,包含39个语音属性
评估维度
- 评估方法: 结合客观指标和主观指标
- 核心能力: 口语内容理解、非词汇语音线索感知、整合推理、响应生成
- 测试范围: 已测试12个先进的语音语言模型
主要发现
- 现有最先进模型在处理高表达性语音线索方面存在困难
- 在指令遵循、自然语音变化的适应性以及有效利用语音线索实现共情方面存在持续弱点
相关资源
- 论文: https://arxiv.org/abs/2510.22758
- 官方网站: https://hlt-cuhksz.github.io/EchoMind/
- 排行榜: https://hlt-cuhksz.github.io/EchoMind/
- HuggingFace数据集页面: https://huggingface.co/datasets/hlt-cuhksz/EchoMind

- 1EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models香港中文大学(深圳)大数据研究院 · 2025年



