GMOS-2K
收藏资源简介:
GMOS-2K是由牛津大学视觉几何组与上海交通大学联合创建的视频移动对象分割数据集,旨在为三维空间和时间细粒度运动分析提供基准资源。该数据集包含2,210个真实世界视频序列,总计涵盖4,648个独立运动对象,数据来源于五个成熟的视频对象分割基准(DAVIS17、YTVOS19、OVIS、MoCA-Mask和HOI4D),并经过严格的筛选与标注流程。数据创建过程通过对原始视频进行双重过滤,并对运动对象添加时间细粒度标注,精确标记每个对象在时间轴上的运动区间。该数据集主要应用于自动驾驶、视频监控和三维场景重建等领域,旨在解决传统移动对象分割方法在三维几何感知不足和时间粒度粗糙方面的局限性,推动实时在线运动分割技术的发展。
GMOS-2K is a video moving object segmentation dataset jointly developed by the Visual Geometry Group of the University of Oxford and Shanghai Jiao Tong University, designed to serve as a benchmark resource for fine-grained 3D spatial and temporal motion analysis. This dataset comprises 2,210 real-world video sequences, encompassing a total of 4,648 independently moving objects. The source data is collected from five mature video object segmentation benchmarks, namely DAVIS17, YTVOS19, OVIS, MoCA-Mask, and HOI4D, and has undergone strict screening and annotation workflows. In the dataset creation process, double filtering is performed on the original videos, and fine-grained temporal annotations are added to moving objects to accurately mark their motion intervals along the temporal axis. Primarily applied in fields such as autonomous driving, video surveillance, and 3D scene reconstruction, this dataset aims to alleviate the limitations of conventional moving object segmentation methods, including insufficient 3D geometric perception and overly coarse temporal granularity, and advance the development of real-time online motion segmentation technologies.
数据集概述
- 数据集名称:GMOS-2K
- 用途:用于运动物体分割(Moving Object Segmentation),支持在 RGB 视频上输出 3D 感知、时间精细粒度的多运动物体分割,并提供前景-背景变体 GMOS-S 以加速部署。
- 数据来源:从五个已建立的 视频物体分割(VOS) 基准数据集中筛选和标注:
- DAVIS
- YTVOS
- OVIS
- MoCA-Mask
- HOI4D
- 规模:
- 视频总数:2,210 个真实世界视频
- 标注运动物体数量:4,648 个
- 划分:1,930 个训练视频 / 280 个测试视频
- 标注类型:
- 每个物体的逐帧分割掩码
- 时间区间标签:精确标注每个物体在视频中 运动的时间区间,实现时间精细粒度的运动标注
- 配套评价协议:MOS-I(Instantaneous),包含三项互补指标,用于评估时间精细粒度的运动分割性能。
性能与特点
- GMOS 在 MOS、MOS-I 和无监督 VOS 基准上达到最优结果。
- 运行速度显著快于以往的多物体 MOS 方法。
- 支持在线推理,可用于流式部署场景。




