MDSEval
收藏资源简介:
MDSEval是一个针对多模态对话摘要任务的元评估基准数据集,由AWS AI Labs和剑桥大学语言技术实验室的研究人员创建。该数据集包含了198个高质量的图像分享对话,每个对话配对5个由SOTA MLLMs生成的摘要,并由人类专家从八个维度进行评估。MDSEval旨在推动鲁棒、人类对齐的多模态评估方法的发展,并促进更复杂的多模态对话代理的研究。
MDSEval is a meta-evaluation benchmark dataset for multimodal dialogue summarization tasks. It was developed by researchers from AWS AI Labs and the Language Technology Laboratory at the University of Cambridge. The dataset consists of 198 high-quality image-sharing dialogues, each paired with 5 summaries generated by state-of-the-art (SOTA) multimodal large language models (MLLMs), and evaluated by human experts across eight dimensions. MDSEval aims to advance the development of robust, human-aligned multimodal evaluation methodologies and foster research on more sophisticated multimodal dialogue agents.

- 1MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue SummarizationAWS AI Labs, Language Technology Lab, University of Cambridge · 2025年



