VMCBench
收藏资源简介:
VMCBench是一个由20个现有VQA数据集转换而成的多选题基准数据集,包含9018个问题,旨在为视觉语言模型(VLM)提供标准化、可扩展的评估工具。数据集通过AutoConverter框架自动生成,确保了问题的正确性和挑战性。VMCBench涵盖了多种问题类型,能够全面评估VLM在不同任务中的表现。该数据集的应用领域主要集中在视觉语言模型的评估和优化,旨在解决开放式问题评估中的语义相似性测量难题,并提供更客观、可重复的评估方法。
VMCBench is a multiple-choice benchmark dataset converted from 20 existing VQA datasets, consisting of 9018 questions, aiming to provide a standardized and scalable evaluation tool for Vision-Language Models (VLMs). The dataset is automatically generated via the AutoConverter framework, ensuring the correctness and challenging nature of the questions. VMCBench covers a wide range of question types, enabling comprehensive evaluation of VLM performance across diverse tasks. The dataset is primarily applied to the evaluation and optimization of vision-language models, aiming to address the challenge of semantic similarity measurement in open-ended question evaluation and provide more objective and reproducible evaluation methods.




