EMBODIEDBENCH
收藏资源简介:
EMBODIEDBENCH是由伊利诺伊大学厄巴纳-香槟分校等机构创建的综合性评测数据集,包含四种环境下的1128个测试任务,覆盖从高级语义任务到低级原子动作任务。数据集经过精心设计,不仅具有多样的任务层次,还引入了面向能力的细粒度评估框架。该数据集能够全面评估多模态大型语言模型在视觉感知、常识推理、复杂指令理解、空间感知和长期规划等方面的性能。
EMBODIEDBENCH is a comprehensive benchmark dataset developed by the University of Illinois Urbana-Champaign and other institutions. It comprises 1128 test tasks across four distinct environments, covering tasks ranging from high-level semantic tasks to low-level atomic action tasks. Meticulously designed, this dataset not only features diverse task hierarchies but also introduces a capability-oriented fine-grained evaluation framework. This benchmark enables comprehensive performance evaluation of multimodal large language models across multiple domains including visual perception, commonsense reasoning, complex instruction understanding, spatial perception, and long-term planning.

- 1EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents伊利诺伊大学厄巴纳-香槟分校, Northwestern University, 多伦多大学, 芝加哥丰田技术研究所 · 2025年



