MFE-ETP
收藏资源简介:
MFE-ETP数据集由天津大学智能与计算学部创建,是一个针对具身任务规划的多模态基础模型综合评估基准。该数据集包含1184个高质量测试案例,覆盖100个具身任务,涉及对象理解、时空感知、任务理解和具身推理等多个能力维度。数据集的创建过程结合了从BEHAVIOR-100和VirtualHome平台收集的典型家庭任务数据,并通过人工标注和设计任务指令进行精细化处理。MFE-ETP数据集主要应用于提升多模态基础模型在具身人工智能领域的任务规划能力,旨在解决模型在复杂任务场景中的性能瓶颈问题。
The MFE-ETP dataset, developed by the College of Intelligence and Computing at Tianjin University, is a comprehensive evaluation benchmark for multimodal foundation models focused on embodied task planning. It comprises 1,184 high-quality test cases spanning 100 embodied tasks, covering multiple capability dimensions including object understanding, spatio-temporal perception, task comprehension, and embodied reasoning. The dataset was constructed by integrating typical household task data collected from the BEHAVIOR-100 and VirtualHome platforms, followed by fine-grained processing via manual annotation and task instruction design. The MFE-ETP dataset is primarily applied to improve the task planning capabilities of multimodal foundation models in the domain of embodied artificial intelligence, with the goal of resolving performance bottlenecks of models in complex task scenarios.
数据集概述
标题
MFE-ETP: An Embodied Task Planning Benchmark for Multi-modal Foundation Models
作者
- Min Zhang
- Jianye Hao
- Xian Fu
- Peilong Han
- Hao Zhang
- Lei Shi
- Hongyao Tang
机构
- Tianjin University
- Montreal Institute of Learning Algorithms (MILA)
摘要
近年来,多模态基础模型(MFMs)和具身人工智能(EAI)以空前的速度并行发展,两者的结合引起了AI研究界的广泛关注。本工作旨在深入全面地评估MFMs在具身任务规划方面的性能,以揭示其在该领域的功能和局限性。为此,基于具身任务规划的特点,我们首先开发了一个系统的评估框架,该框架涵盖了MFMs的四个关键能力:对象理解、时空感知、任务理解和具身推理。随后,我们提出了一个新的基准,名为MFE-ETP,其特点是任务场景复杂多变、任务类型典型多样、任务实例难度不一,以及从多模态问题回答到具身任务推理的丰富测试案例类型。最后,我们提供了一个简单易用的自动评估平台,使多个MFMs能够在提出的基准上进行自动化测试。通过使用该基准和评估平台,我们评估了几个最先进的MFMs,发现它们与人类水平的性能存在显著差距。MFE-ETP是一个高质量、大规模且具有挑战性的基准,与现实世界任务相关。
相关链接




