基于THU-READ的行为识别研究数据集
收藏资源简介:
THU-READ 数据集主要面向 RGB-D 第一人称视角下的动作识别研究以及相关应用开发需求建设。它由清华大学研究团队基于 Primesense Carmine 相机采集产生,该相机可同时记录 RGB 和深度视频序列。在采集时,研究人员将相机安装在头盔上,模拟真实场景,邀请 8 位受试者(6 男 2 女)在实验室、浴室、会议室、宿舍和餐厅这 5 种不同场景下,自然地重复执行各类动作,每个动作重复 3 次。数据集主要内容包含 40 种日常生活动作,涵盖单手、双手操作,以及与物体交互和非交互的动作,如 “bounce ball”“cut fruit”“thumb” 等。从体量上看,原始数据经处理后,共得到 1920 个视频片段,343,626 个有效帧,平均每个动作视频实例包含 179 帧,分辨率为 640×480。该数据集的意义重大,弥补了此前相关数据集在深度模态利用和第一人称视角动作识别方面的不足,为动作识别算法研究提供了丰富多样的样本,有助于推动虚拟现实、生活记录、远程康复等领域的发展。
The THU-READ dataset was developed to meet the research and application development needs of first-person action recognition based on RGB-D modality. It was collected by the research team from Tsinghua University using a Primesense Carmine camera, which can simultaneously capture RGB and depth video sequences. During data collection, researchers mounted the camera on a helmet to simulate real-world scenarios, and invited 8 participants (6 males and 2 females) to naturally perform various actions repeatedly in 5 different scenarios: laboratory, bathroom, conference room, dormitory, and dining room, with each action being repeated 3 times. The dataset mainly includes 40 daily life action categories, covering one-handed and two-handed operations, as well as object-interactive and non-interactive actions, such as "bounce ball", "cut fruit", "thumb", etc. In terms of scale, after processing the raw data, a total of 1920 video clips and 343,626 valid frames are obtained. Each action video instance contains an average of 179 frames, with a resolution of 640×480. This dataset holds great significance, as it fills the gaps in existing related datasets regarding depth modality utilization and first-person action recognition. It provides diverse samples for action recognition algorithm research, and promotes the development of fields such as virtual reality, life logging, and remote rehabilitation.




