遇见数据集

May-apple/VBVR-Reorganized

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

VBVR-Reorganized数据集是VBVR(基于视频的视觉推理)数据集的重新组织、提示清洗和配对变体增强版本,专为视频生成训练设计。该数据集将每个任务划分为Pure_Reasoning(纯推理)和Instruction_Following(指令跟随)两类,对Pure_Reasoning任务的提示进行重写以去除可能泄露答案的短语,并引入了4个配对变体生成器,这些生成器与正向任务共享第一帧但要求模型根据提示生成不同的真实视频,以增强模型对提示的理解而非记忆。数据集通过硬链接引用原始VBVR媒体文件,不额外占用磁盘空间。它包括训练集(包含51个Pure_Reasoning生成器和53个Instruction_Following生成器,共1,040,000个样本)和测试集(分为域内和域外,共540个样本)。数据以parquet格式存储,每个样本包含提示文本、原始提示(可选)、第一帧图像、最终帧图像、真实视频和元数据等字段。数据集还提供了两个小型检查分割(train_pr_rewritten和train_pr_unchanged)以便于浏览重写效果。总体而言,该数据集旨在支持视频生成模型的训练和评估,强调推理能力和提示遵循。

The VBVR-Reorganized Dataset is a reorganized, prompt-cleaned and paired-variant-enhanced version of the VBVR (Video-Based Visual Reasoning) dataset, specifically designed for video generation training. This dataset divides each task into two categories: Pure_Reasoning and Instruction_Following. For the prompts of Pure_Reasoning tasks, it rewrites them to remove phrases that may leak the correct answer, and introduces 4 paired-variant generators. These generators share the first frame with their corresponding positive tasks, but require the model to generate distinct realistic videos based on the prompts, so as to enhance the model's understanding of prompts rather than rote memorization. The dataset references original VBVR media files via hard links, thus not occupying additional disk space. It includes a training set (containing 51 Pure_Reasoning generators and 53 Instruction_Following generators, totaling 1,040,000 samples) and a test set (divided into in-domain and out-of-domain, with 540 samples in total). The data is stored in parquet format, and each sample contains fields such as prompt text, original prompt (optional), first-frame image, final-frame image, ground-truth video and metadata. The dataset also provides two small check splits (train_pr_rewritten and train_pr_unchanged) to facilitate browsing the effect of prompt rewriting. Overall, this dataset aims to support the training and evaluation of video generation models, with an emphasis on reasoning ability and prompt following.

提供机构:
May-apple
二维码
社区交流群
二维码
科研交流群
商业服务