molmo-motion-1m
收藏资源简介:
MolmoMotion-1M 是一个大规模、多领域的3D点轨迹标注数据集,专为3D点跟踪、轨迹预测及相关机器人视觉研究设计。该数据集整合了来自七个不同视频语料库的标注,涵盖第一人称操作、真实世界机器人遥操作、动态真实世界场景以及模拟器渲染等多种场景。核心内容是为每个视频片段提供精确的、经过运动过滤的3D点轨迹,大多数子数据集还提供对应的2D像素级轨迹。每个数据样本包含:与源视频帧对齐的3D/2D轨迹点序列、一个简短的动作描述文本、逐帧的相机姿态与内参,以及预设的训练/测试数据划分。数据以结构化格式组织,根目录下包含多个子数据集文件夹,其中annotations/目录下的JSON索引文件定义了视频片段元数据和数据集划分,实际的轨迹数据和相机参数存储在以视频ID命名的.npz文件中。3D轨迹和2D轨迹的轴顺序不同,存储格式因上游来源而异。数据集适用于计算机视觉和机器人学领域的多项任务,如3D点跟踪、长时轨迹预测、以物体为中心的视频理解,以及涉及第一人称视角和机器人操作场景的研究。数据集的混合许可证结构要求用户在使用每个子数据集前查阅对应的上游许可证条款。
MolmoMotion-1M is a large-scale, multi-domain 3D point trajectory annotation dataset designed specifically for 3D point tracking, trajectory prediction, and related robotic vision research. This dataset integrates annotations from seven distinct video corpora, covering various scenarios including first-person manipulation, real-world robotic teleoperation, dynamic real-world scenes, and simulator-rendered environments. The core content is to provide precise, motion-filtered 3D point trajectories for each video segment, and most sub-datasets also provide corresponding 2D pixel-level trajectories. Each data sample contains: 3D/2D trajectory point sequences aligned with source video frames, a brief action description text, per-frame camera poses and intrinsic parameters, as well as pre-defined train/test data splits. The data is organized in a structured format: the root directory contains multiple sub-dataset folders, where the JSON index file under the annotations/ directory defines video segment metadata and dataset splits, while actual trajectory data and camera parameters are stored in .npz files named after video IDs. The axis order of 3D trajectories and 2D trajectories differs, and storage formats vary depending on upstream sources. The dataset is applicable to multiple tasks in the fields of computer vision and robotics, such as 3D point tracking, long-term trajectory prediction, object-centric video understanding, and research involving first-person perspective and robotic manipulation scenarios. The dataset has a mixed license structure, requiring users to consult the corresponding upstream license terms before using each sub-dataset.




