MME-Emotion
收藏资源简介:
MME-Emotion是一个用于评估多模态大语言模型情感智能的系统性基准测试,包含6,500个精选视频片段和任务特定的问答对,涵盖广泛场景以构建8个情感任务。该基准测试具有可扩展能力、多样化设置和统一协议,是MLLMs领域最大的情感智能基准测试
MME-Emotion is a systematic benchmark for evaluating the emotional intelligence of multimodal large language models. It contains 6,500 curated video clips and task-specific question-answer pairs, covering a wide range of scenarios to establish 8 emotional tasks. This benchmark features scalability, diverse experimental settings and a unified evaluation protocol, making it the largest emotional intelligence benchmark in the MLLMs field.
MME-Emotion 数据集概述
数据集基本信息
- 数据集名称:MME-Emotion
- 核心目标:评估多模态大语言模型(MLLMs)在情感智能方面的理解和推理能力
- 数据规模:6,500个精选视频片段,附带任务特定的问答对
- 任务范围:涵盖8种情感任务,涉及广泛场景
主要任务类型
- 视频问答(Video QA)
- 情感推理(Emotion Reasoning)
- 情感识别(Emotion Recognition)
评估体系
- 评估指标:
- 识别分数(Recognition Score)
- 推理分数(Reasoning Score)
- 思维链分数(CoT Score)
- 评估框架:采用多智能体系统框架进行分析
- 验证方式:由五位人类专家全面验证评估策略的有效性
数据集特点
- 可扩展能力:支持大规模评估
- 多样化设置:覆盖广泛场景
- 统一协议:采用标准化的评估流程
相关资源
- 项目页面:https://mme-emotion.github.io/
- 论文地址:https://www.arxiv.org/pdf/2508.09210
- 数据集地址:https://github.com/FunAudioLLM/MME-Emotion/blob/main/MME-Emotion-Dataset/dataset-readme.md
- 排行榜:https://mme-emotion.github.io/#leaderboard
评估模型
已系统评估20个开源和闭源的前沿多模态大语言模型,包括:
- GPT-4
- Gemini
- Qwen-VL
引用信息
latex @article{zhang2025mme, title={MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models}, author={Zhang, Fan and Cheng, Zebang and Deng, Chong and Li, Haoxuan and Lian, Zheng and Chen, Qian and Liu, Huadai and Wang, Wen and Zhang, Yi-Fan and Zhang, Renrui and others}, journal={arXiv preprint arXiv:2508.09210}, year={2025} }




