yxma/tactile-video-pretrain
收藏资源简介:
Tactile Video Pretrain(触觉视频预训练)数据集包含约30小时的GelSight-style触觉视频,并与手腕安装的场景RGB视频配对。每一行数据代表一次完整的接触丰富机器人操作演示,时长约6至20秒,帧率为30 FPS。数据集提供两种配置:tactile_only配置每行包含一个触觉手指视频,适用于触觉自监督预训练(如视频MAE、V-JEPA、对比学习或掩码学习);tactile_rgb配置每行包含配对的触觉视频和手腕RGB视频,适用于跨模态对齐(如CLIP-style对比学习或掩码跨模态学习)。数据集包含11,967行tactile_only数据和11,965行tactile_rgb数据,总计触觉视频时长29.71小时,场景视频时长15.70小时。数据来源于46个任务,共6,338次演示,视频分辨率为640×480,格式为MP4,总帧数达3,208,954帧。数据集还包含每帧的机器人状态轨迹(如工具中心点位置、方向、夹爪距离等),适用于机器人学、视频分类和特征提取等任务。
The Tactile Video Pretrain Dataset contains approximately 30 hours of GelSight-style tactile videos paired with wrist-mounted scene RGB videos. Each data entry represents a complete contact-rich robotic manipulation demonstration, lasting 6 to 20 seconds with a frame rate of 30 FPS. The dataset offers two configurations: the tactile_only configuration, where each entry contains a single tactile finger video, designed for tactile self-supervised pretraining tasks such as Video MAE, V-JEPA, contrastive learning, or masked learning; and the tactile_rgb configuration, where each entry contains paired tactile videos and wrist RGB videos, suitable for cross-modal alignment tasks including CLIP-style contrastive learning or masked cross-modal learning. The dataset includes 11,967 tactile_only entries and 11,965 tactile_rgb entries, with a total tactile video duration of 29.71 hours and a total scene video duration of 15.70 hours. The data is collected from 46 distinct tasks, totaling 6,338 demonstrations, with videos recorded at 640×480 resolution in MP4 format, amounting to a total of 3,208,954 frames. The dataset also provides per-frame robotic state trajectories, such as tool center point (TCP) position, orientation, gripper distance, and other relevant states, which is applicable to research areas including robotics, video classification, and feature extraction.




