VIVID-10M
收藏资源简介:
VIVID-10M是由西北工业大学、北京理工大学和快手科技联合创建的第一个大规模混合图像-视频本地编辑数据集。该数据集包含976万条样本,涵盖广泛的编辑任务,如添加、修改和删除实体。数据集的创建过程结合了多种视觉感知模型和多模态大语言模型,确保了高质量的数据生成。每个样本包括真实数据、掩码数据、掩码和本地描述,适用于视频编辑模型的训练。VIVID-10M旨在解决现有视频编辑数据集的缺乏和高成本问题,适用于视频编辑模型的训练和评估,特别是在提高编辑质量和用户交互性方面。
VIVID-10M is the first large-scale mixed image-video local editing dataset jointly created by Northwestern Polytechnical University, Beijing Institute of Technology, and Kuaishou Technology. This dataset comprises 9.76 million samples, covering a wide range of editing tasks including adding, modifying, and deleting entities. The construction of this dataset integrates multiple visual perception models and multimodal large language models, ensuring high-quality data generation. Each sample includes real-world data, mask data, masks, and local descriptions, making it suitable for the training of video editing models. VIVID-10M aims to address the scarcity and high cost issues of existing video editing datasets, and is applicable to the training and evaluation of video editing models, particularly for enhancing editing quality and user interactivity.




