MultiModal Needle-in-a-haystack (MMNeedle)
收藏资源简介:
MMNeedle数据集由罗格斯大学创建,旨在评估多模态大型语言模型(MLLMs)在长上下文理解能力方面的表现。该数据集包含40,000张图像、560,000个标题和280,000对针与干草堆的配对,通过图像拼接技术扩展输入上下文长度,以测试模型在复杂视觉上下文中定位目标子图像的能力。数据集构建过程中,采用了自动化的正负样本生成方法,确保了数据的多样性和平衡性。MMNeedle数据集的应用领域广泛,主要用于解决MLLMs在处理长上下文混合模态输入时的性能评估问题,有助于推动相关技术的发展和应用。
The MMNeedle dataset, created by Rutgers University, is designed to evaluate the performance of multimodal large language models (MLLMs) in long-context understanding tasks. This dataset consists of 40,000 images, 560,000 captions, and 280,000 needle-in-haystack pairs. It extends the input context length via image stitching technology, aiming to test the model's ability to locate target sub-images within complex visual contexts. During the dataset construction process, automated positive and negative sample generation methods were adopted to ensure data diversity and balance. The MMNeedle dataset has a wide range of application scenarios, primarily used for performance assessment of MLLMs when processing long-context mixed-modal inputs, and it helps advance the development and application of related technologies.




