Look and Tell
收藏资源简介:
Look and Tell数据集是由KTH皇家理工学院创建的,用于研究在第一人称和第三人称视角下的多模态指代交流。数据集包含了25名参与者使用Meta Project Aria智能眼镜和GoPro相机记录的同步注视、语音和视频数据,共计3.67小时的记录,包括2707个丰富的注释指代表达式。该数据集旨在推动能够理解和参与情境对话的具身智能体的发展。数据集提供了对2D和3D场景表示、第一人称和第三人称视角之间进行多模态接地比较的独特测试平台,对于解决共享自主性和人机协作中的挑战至关重要。
The Look and Tell Dataset was developed by KTH Royal Institute of Technology to research multimodal referring communication under both first-person and third-person perspectives. The dataset contains synchronized gaze, speech and video data recorded by 25 participants using Meta Project Aria smart glasses and GoPro cameras, with a total of 3.67 hours of recordings including 2707 richly annotated referring expressions. This dataset aims to advance the development of embodied AI agents that can understand and engage in situated dialogues. It provides a unique testbed for multimodal grounded comparisons between 2D and 3D scene representations, as well as between first-person and third-person perspectives, which is critical for addressing challenges in shared autonomy and human-robot collaboration.
KTH-ARIA-referential 数据集概述
数据集基本信息
- 数据集名称: Gaze-Speech Analysis in Referential Communication with ARIA Headset
- 研究机构: KTH Royal Institute of Technology
- 主要语言: 英语
- 许可协议: CC BY-NC-ND 4.0
数据集规模与统计
- 总时长: 2.259小时
- 样本数量: 96个片段
- 平均片段时长: 84.7秒
数据内容与结构
- 参与者: 20人(2名男性,18名女性)
- 实验任务: 参与者记忆食谱配料和步骤,佩戴ARIA眼镜进行口头指导
- 数据类型:
- 眼动数据(注视跟踪)
- 语音数据(音频录音)
- 第三人称视角视频
- 文本记录(话语内容)
数据格式
- 音频文件(.wav)
- 话语文本(.txt)
- 第一人称视角视频(.mp4)
- 注视固定数据(通过Python脚本生成)
适用研究领域
- 指称沟通分析
- 注视与语音同步研究
- 人机交互与多模态对话系统
- 任务环境中的眼动追踪研究
数据采集方法
- 硬件设备: ARIA智能眼镜、GoPro相机
- 采集方式: 实时捕获参与者在描述食谱时的注视轨迹和言语表达
研究目的
探索指称沟通中注视与语音的同步模式,以及物体位置对这种同步的影响。
数据处理与分析
- 使用Python脚本进行时间相关性检测
- 提供辅助函数用于绘制注视固定点和跟踪物体

- 1Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric ViewsKTH皇家理工学院 · 2025年



