IRef-VLA
收藏资源简介:
IRef-VLA数据集是由卡内基梅隆大学机器人学院创建的,旨在为交互式参考视觉和语言引导的3D场景动作提供基准数据集。该数据集包含超过11.5K个来自现有数据集的扫描3D房间,7.6M个启发式生成的语义关系和4.7M个参考语句。数据集还包含语义对象和房间注释、场景图、可导航自由空间注释,并加入了语言不完美的参考语句。该数据集适用于三维场景理解,有助于开发健壮的交互式导航系统。
The IRef-VLA dataset was created by the Robotics Institute at Carnegie Mellon University, designed to serve as a benchmark dataset for interactive reference-based vision-and-language guided 3D scene action tasks. It contains over 11.5K scanned 3D rooms sourced from existing datasets, 7.6M heuristically generated semantic relationships, and 4.7M referring expressions. The dataset also includes semantic object and room annotations, scene graphs, navigable free space annotations, as well as referring expressions with imperfect language. This dataset supports 3D scene understanding research and facilitates the development of robust interactive navigation systems.

- 1IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes卡内基梅隆大学机器人学院 · 2025年



