johncmm/Video-R1-data
收藏资源简介:
该数据集来自论文《Video-R1: Reinforcing Video Reasoning in MLLMs》,旨在强化多模态大语言模型(MLLMs)中的视频推理能力。数据集包含视频和图像数据,视频数据文件夹包括CLEVRER、LLaVA-Video-178K、NeXT-QA、PerceptionTest和STAR,图像数据文件夹包括Chart、General、Knowledge、Math、OCR和Spatial。数据集提供两个JSON文件:Video-R1-COT-165k.json用于监督微调(SFT)冷启动,Video-R1-260k.json用于强化学习(RL)训练。数据格式涉及多模态问题回答,例如视频相关的问题,包含问题ID、问题文本、数据类型(视频或图像)、问题类型(如多项选择)、选项列表、详细的推理过程(以“思考”形式呈现)和答案。示例显示问题涉及视频内容分析,如识别屏幕上的俄语文本,并提供了逐步推理和最终答案。数据集规模在10万到100万之间,语言为英语,任务类别为视频-文本到文本。
This dataset is from the paper Video-R1: Reinforcing Video Reasoning in MLLMs and aims to enhance video reasoning capabilities in multimodal large language models (MLLMs). It includes both video and image data, with video data folders covering CLEVRER, LLaVA-Video-178K, NeXT-QA, PerceptionTest, and STAR, and image data folders covering Chart, General, Knowledge, Math, OCR, and Spatial. The dataset provides two JSON files: Video-R1-COT-165k.json for supervised fine-tuning (SFT) cold start, and Video-R1-260k.json for reinforcement learning (RL) training. The data format involves multimodal question-answering, such as video-related questions, containing fields like problem ID, problem text, data type (video or image), problem type (e.g., multiple choice), options list, a detailed reasoning process (presented as a thought process), and the solution. An example shows problems analyzing video content, such as identifying Russian text on screen, with step-by-step reasoning and final answers. The dataset size is between 100K and 1M, the language is English, and the task category is video-text-to-text.



