VIDEOMOLMO
收藏资源简介:
VIDEOMOLMO数据集是一套包含72,000个视频-字幕对和100,000个物体点的综合数据集,旨在支持基于文本描述的精细时空指向。该数据集由多个来源的视频数据构建而成,如Refer-YTVOS、Refer-DAVIS、MeViS等,通过半自动化的标注流程确保了高质量和可扩展性。数据集用于训练VIDEOMOLMO模型,该模型能够根据自然语言查询生成整个视频序列中目标物体的点级预测,并保持时间一致性。VIDEOMOLMO数据集的发布填补了当前时空指向数据集的空白,为视觉定位和推理任务提供了宝贵资源。
The VIDEOMOLMO dataset is a comprehensive corpus comprising 72,000 video-caption pairs and 100,000 object points, designed to support fine-grained spatial-temporal grounding based on textual descriptions. This dataset is constructed from video data sourced from multiple resources including Refer-YTVOS, Refer-DAVIS, MeViS and others, and a semi-automated annotation pipeline is adopted to ensure high annotation quality and scalability. The dataset is used to train the VIDEOMOLMO model, which can generate point-level predictions of target objects across entire video sequences based on natural language queries while maintaining temporal consistency. The release of the VIDEOMOLMO dataset fills the gap in current spatial-temporal grounding datasets, providing a valuable resource for visual grounding and reasoning tasks.




