ASLP-lab/HumDial-EIBench
收藏资源简介:
HumDial-EIBench是一个用于评估音频语言模型情感理解能力的人类录制多轮对话情感智能基准。该基准旨在诊断模型是否真正理解语音中的情感,而非依赖文本转录的捷径。数据集基于ICASSP 2026 HumDial挑战赛的真实人类录音对话构建,包含中文和英文子集,总计1,077个样本。核心目标是通过四个任务评估模型在记忆、推理、生成和跨模态鲁棒性方面的情感智能:任务1为情感轨迹检测(多轮跟踪情感变化),任务2为隐式因果推理(从上下文线索推断情感触发因素),任务3为共情响应生成(评估文本共情、声学共情和音频质量),任务4为声学-语义冲突(测试文本情感与声学情感矛盾时的鲁棒性)。数据设计强调真实人类多轮音频、客观对抗性多选题任务,以及声学-语义冲突任务,以解决现有基准中合成语音、单轮设置和主观评分等问题。
HumDial-EIBench is a human-recorded multi-turn emotional intelligence benchmark for audio language models, designed to evaluate whether ALMs truly understand emotion in speech rather than relying on text transcription shortcuts. Built from authentic human-recorded dialogues from the ICASSP 2026 HumDial Challenge, it includes both Chinese and English subsets with a total of 1,077 samples. The core goal is to diagnose emotional intelligence in ALMs across memory, reasoning, generation, and cross-modal robustness through four tasks: Task 1 (Emotional Trajectory Detection) tracks emotion changes across dialogue turns, Task 2 (Implicit Causal Reasoning) infers latent emotional triggers from scattered context clues, Task 3 (Empathetic Response Generation) evaluates responses in textual empathy, vocal empathy, and audio quality, and Task 4 (Acoustic-Semantic Conflict) tests robustness when text sentiment contradicts vocal affect. The benchmark addresses gaps in existing ALM evaluations by combining real human multi-turn audio, objective adversarial MCQ tasks, and a dedicated acoustic-semantic conflict task.




