EMMA (Enhanced MultiModal reAsoning)
收藏资源简介:
EMMA(增强多模态推理)数据集由电子科技大学、中山大学、华盛顿大学、微软和香港中文大学的研究团队共同创建,旨在评估多模态大语言模型在数学、物理、化学和编程领域的多模态推理能力。该数据集包含2788个问题,其中1796个是新构建的,问题类型涵盖选择题和开放式问题,涉及图像和文本的多模态推理任务。数据集的构建过程包括从现有基准中筛选问题,并通过与领域专家合作手动创建新问题。EMMA的应用领域主要集中在多模态推理能力的评估,旨在解决当前MLLMs在处理复杂多模态和多步推理任务时的局限性。
EMMA (Enhanced Multimodal Reasoning) dataset was co-created by research teams from the University of Electronic Science and Technology of China, Sun Yat-sen University, the University of Washington, Microsoft, and The Chinese University of Hong Kong. It is designed to evaluate the multimodal reasoning capabilities of multimodal large language models (MLLMs) across the domains of mathematics, physics, chemistry, and programming. This dataset contains 2,788 questions, among which 1,796 are newly constructed. The question types cover both multiple-choice questions and open-ended questions, involving multimodal reasoning tasks with both images and text. The construction process of EMMA includes screening questions from existing benchmarks and manually creating new questions in collaboration with domain experts. The primary application scenario of the EMMA dataset is the evaluation of multimodal reasoning capabilities, aiming to address the limitations of current MLLMs in handling complex multimodal and multi-step reasoning tasks.




