LongWriter-V22k
收藏资源简介:
LongWriter-V22k是由清华大学知识工程实验室收集的一个数据集,包含22,158个示例,每个示例包含多个输入图像、一个指令和相应的输出,输出长度从0到10,000词不等。该数据集通过两个阶段收集:首先是从MMEvol数据集中筛选出长输出指令,并通过LongWrite Agent-V管道生成对应的长时间文本作为SFT数据;其次是收集人类对VLM输出的细粒度修正,用于DPO数据。数据集旨在提高视觉语言模型的长文本生成能力。
LongWriter-V22k is a dataset collected by the Knowledge Engineering Laboratory of Tsinghua University. It contains 22,158 instances, each consisting of multiple input images, an instruction, and a corresponding output, with the output length ranging from 0 to 10,000 words. The dataset is collected in two stages: first, long-output instructions are screened from the MMEvol dataset, and the LongWrite Agent-V pipeline is used to generate corresponding long texts as Supervised Fine-Tuning (SFT) data; second, fine-grained human corrections to Vision-Language Model (VLM) outputs are collected for Direct Preference Optimization (DPO) data. The dataset aims to enhance the long-text generation capabilities of vision-language models.




