InstructMove
收藏资源简介:
InstructMove数据集是由东京大学和Adobe合作创建的,用于训练基于指令的图像编辑模型。该数据集从互联网视频中提取帧对,并使用多模态大语言模型生成编辑指令,包含600万对图像和相应的编辑指令。数据集的创建过程包括从视频中采样帧对、使用多模态LLMs生成指令,并进行运动过滤以确保帧对的质量。该数据集主要用于解决复杂图像编辑任务,如调整主体姿态、重新排列元素和改变视角,旨在提高图像编辑的精确性和自然性。
The InstructMove dataset was co-created by The University of Tokyo and Adobe for training instruction-based image editing models. It extracts frame pairs from internet videos and generates editing instructions via multimodal large language models, containing 6 million image-instruction pairs. The dataset creation pipeline includes sampling frame pairs from videos, generating instructions using multimodal LLMs, and performing motion filtering to ensure the quality of the frame pairs. This dataset is primarily designed to address complex image editing tasks such as adjusting subject poses, rearranging elements, and altering perspectives, with the goal of improving the accuracy and naturalness of image editing.

- 1Instruction-based Image Manipulation by Watching How Things Move东京大学, Adobe · 2024年



