Long Visual Question Answering (Long-VQA), Long Multimodal Retrieval (Long-MR)
收藏资源简介:
本研究构建了两个增强型长上下文多模态数据集:Long Visual Question Answering (Long-VQA) 和 Long Multimodal Retrieval (Long-MR)。Long-VQA 数据集扩展了17个广泛使用的数据集,将其内容从短序列扩展到包含多达32K个tokens的长序列,旨在评估视觉语言模型在长序列中的理解和推理能力。Long-MR 数据集则通过插入目标图像或文本段到交错图像和文本的序列中,评估模型从超长多模态序列中检索特定目标的能力。这些数据集的创建旨在增强视觉语言模型在长上下文场景中的训练和评估,解决现有数据集在长上下文理解方面的不足。
This study develops two enhanced long-context multimodal datasets: Long Visual Question Answering (Long-VQA) and Long Multimodal Retrieval (Long-MR). The Long-VQA dataset adapts 17 widely used datasets by extending their content from short sequences to long sequences containing up to 32K tokens, aiming to evaluate the comprehension and reasoning capabilities of vision-language models in long-sequence scenarios. The Long-MR dataset, by inserting target images or text segments into interleaved image-text sequences, assesses models' ability to retrieve specific targets from ultra-long multimodal sequences. These datasets are created to strengthen the training and evaluation of vision-language models in long-context settings, addressing the gaps in long-context understanding of existing datasets.

- 1V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding清华大学, 商汤科技研究, 香港大学, 上海人工智能实验室 · 2024年



