“关于环境和行为的推理” (REA) 数据集
收藏资源简介:
REA数据集是一个用于评估多模态大型语言模型在空间和时间理解能力的大规模数据集。该数据集由伊利诺伊大学厄巴纳-香槟分校的研究团队开发,旨在帮助模型理解3D环境的全局结构和视频中最近发生的动作。数据集包括五个任务,分别测试模型在相对方向、相对距离、物品定位、家具使用预测和动作规划方面的能力。REA数据集通过结合3D场景表示和2D视频表示,为模型提供丰富的时空上下文信息。为了创建REA数据集,研究团队使用EPIC-KITCHENS数据集的动作标注、VISOR数据集的对象分割标注以及EPIC-FIELDS数据集的稀疏点云。数据收集流程包括两个主要组件:问答生成和点云重建。问答生成包括视频采样、3D位置估计、空间关系估计和导航动作估计。点云重建则通过将2D分割掩码投影到COLMAP点云上,计算出查询对象在3D空间中的平均位置。REA数据集旨在解决多模态大型语言模型在时空理解方面的挑战,推动模型在现实世界中的应用,如具身AI、情景感知和时空问答等。
The REA Dataset is a large-scale dataset designed for evaluating the spatial and temporal understanding capabilities of multimodal large language models. Developed by a research team from the University of Illinois Urbana-Champaign, it aims to help models comprehend the global structure of 3D environments and recently occurred actions in videos. The dataset encompasses five tasks that respectively test the model's abilities in relative orientation, relative distance, object localization, furniture usage prediction, and motion planning. The REA Dataset provides rich spatio-temporal contextual information for models by combining 3D scene representations and 2D video representations. To construct the REA Dataset, the research team utilized action annotations from the EPIC-KITCHENS Dataset, object segmentation annotations from the VISOR Dataset, and sparse point clouds from the EPIC-FIELDS Dataset. The data collection process consists of two main components: question-answer generation and point cloud reconstruction. Question-answer generation includes video sampling, 3D position estimation, spatial relationship estimation, and navigation action estimation. Point cloud reconstruction calculates the average 3D position of queried objects by projecting 2D segmentation masks onto COLMAP point clouds. The REA Dataset aims to address the challenges in spatio-temporal understanding of multimodal large language models, and promote their real-world applications such as embodied AI, context awareness, and spatio-temporal question answering.




