STGR-CoT-30k, STGR-RL-36k
收藏资源简介:
Open-o3Video是一个非代理框架,它将显式的时空证据整合到视频推理中。为了支持这一功能,我们首先精心策划和构建了两个高质量的数据集,STGR-CoT-30k用于SFT,STGR-RL-36k用于RL。这些数据集集成了现有仅有时空标注的资源,并包含了5900个新标注的高质量时空样本。每个实例都包含一个问题-答案对、时间戳关键帧、本地化边界框以及一个将视觉证据明确链接到推理步骤的思考链。这些设计为SFT提供了同步的时空监督,以获取有根据的推理格式,并为RL提供了可靠、可验证的信号,以在复杂的视频动态下优化对齐。STGR-CoT-30k包含13.7%的时空数据和50.0%的一般问答数据,而STGR-RL-36k包含30.3%的时空数据和41.7%的问答数据。这些数据集旨在帮助模型学习如何在动态场景中进行一致的定位,并为强化学习提供可验证的奖励。Open-o3Video在V-STAR基准测试和其他视频理解任务上取得了最先进的性能,证明了其在长视频推理、感知导向任务和细粒度时空定位方面的优势。
Open-o3Video is a non-agent framework that integrates explicit spatiotemporal evidence into video reasoning. To support this capability, we first carefully curated and constructed two high-quality datasets: STGR-CoT-30k for Supervised Fine-Tuning (SFT) and STGR-RL-36k for Reinforcement Learning (RL). These datasets integrate existing resources with only spatiotemporal annotations, and additionally include 5,900 newly annotated high-quality spatiotemporal samples. Each instance contains a question-answer pair, timestamped keyframes, localization bounding boxes, and a Chain of Thought (CoT) that explicitly links visual evidence to reasoning steps. These designs provide synchronized spatiotemporal supervision for SFT to acquire well-grounded reasoning formats, and provide reliable, verifiable signals to optimize alignment under complex video dynamics for RL. STGR-CoT-30k contains 13.7% spatiotemporal data and 50.0% general question-answering data, while STGR-RL-36k contains 30.3% spatiotemporal data and 41.7% question-answering data. These datasets are designed to help models learn consistent localization in dynamic scenarios, and provide verifiable rewards for reinforcement learning. Open-o3Video achieves state-of-the-art performance on the V-STAR benchmark and other video understanding tasks, demonstrating its advantages in long-form video reasoning, perception-oriented tasks, and fine-grained spatiotemporal localization.




