InsViE-1M
收藏资源简介:
InsViE-1M是一个包含100万高质量训练三元组(源视频、编辑视频、指令)的指令式视频编辑数据集。该数据集由香港理工大学和OPPO研究院共同构建,旨在提高指令式视频编辑模型的性能。数据集通过精心设计的两阶段编辑-过滤管道构建而成,包括从现实世界视频、图像编辑对以及真实图像生成的静态视频三种来源生成的三元组。数据集的创建过程涉及使用大型视觉语言模型生成指令,以及采用基于视频生成模型的方法来编辑和过滤视频,以确保数据的高质量和适用性。该数据集的应用领域主要是为了解决指令式视频编辑的问题,提升编辑模型的性能和泛化能力。
InsViE-1M is an instructional video editing dataset containing 1 million high-quality training triples (source video, edited video, instruction). This dataset was jointly constructed by The Hong Kong Polytechnic University and OPPO Research Institute, aiming to improve the performance of instructional video editing models. It is built via a meticulously designed two-stage editing-filtering pipeline, which generates training triples from three sources: real-world videos, video-image edit pairs, and static videos generated from real images. The dataset creation process involves leveraging large vision-language models to generate editing instructions, as well as adopting video generation model-based methods to edit and filter videos, ensuring the high quality and applicability of the collected data. The primary applications of this dataset focus on resolving the challenges in instructional video editing, and enhancing the performance and generalization capability of video editing models.




