MMVM Benchmark
收藏资源简介:
MMVM Benchmark是由武汉大学、字节跳动Seed等机构共同构建的多模态视觉匹配基准数据集,旨在评估多模态大语言模型在视觉匹配任务中的表现。该数据集包含1510个手动标注的多图像问答对,数据来源于15个公开数据集和互联网视频平台,涵盖了室内、城市、荒野等多种场景。数据集通过自动标注管道生成,包含220K视觉匹配数据,并带有推理标注。MMVM Benchmark的应用领域主要集中在视觉匹配任务,旨在解决多模态大语言模型在视觉对应性理解上的不足,提升其在视觉推理和匹配任务中的表现。
MMVM Benchmark is a multimodal visual matching benchmark dataset jointly constructed by Wuhan University, ByteDance Seed and other institutions, aiming to evaluate the performance of multimodal large language models (LLMs) in visual matching tasks. This dataset contains 1510 manually annotated multi-image question-answer pairs, which are sourced from 15 public datasets and online video platforms, covering diverse scenarios such as indoor, urban, wilderness and others. Generated via an automatic annotation pipeline, the dataset includes 220K visual matching samples with reasoning annotations. The MMVM Benchmark mainly focuses on applications in visual matching tasks, aiming to address the gaps in multimodal LLMs' understanding of visual correspondence and improve their performance in visual reasoning and matching tasks.




