Grounded Language Dataset (GoLD)
收藏资源简介:
Grounded Language Dataset (GoLD) 是一个多模态数据集,包含常见家用物品的图像和深度点云,以及人们使用口语或书面语言对其进行的描述。该数据集由马里兰大学巴尔的摩分校创建,旨在支持机器人、自然语言处理和人类计算机交互领域的研究。数据集包含47个对象类别,分布在五个高级别类别中,每个类别包含四到五个实例,总计207个对象实例。在创建过程中,对象在转盘上旋转,从不同角度捕捉图像和深度数据。GoLD数据集的应用领域包括开发能够理解和响应自然语言命令的机器人,以及研究多模态(图像、文本和语音)之间的交互。
Grounded Language Dataset (GoLD) is a multimodal dataset comprising images and depth point clouds of everyday household objects, paired with their descriptions in either spoken or written language. Developed by the University of Maryland, Baltimore County, this dataset aims to support research across robotics, natural language processing, and human-computer interaction domains. The dataset encompasses 47 object categories, which are organized under five high-level classes. Each high-level category contains 4 to 5 individual instances, resulting in a total of 207 object instances. During its construction, objects were rotated on a turntable, with images and depth data captured from multiple perspectives. Application use cases for the GoLD dataset include developing robots capable of understanding and responding to natural language commands, as well as investigating cross-modal interactions between images, text, and speech.

- 1Presentation and Analysis of a Multimodal Dataset for Grounded Language Learning马里兰大学巴尔的摩分校 · 2020年



