AUTOBENCH-V
收藏资源简介:
AUTOBENCH-V是由阿卜杜拉国王科技大学等机构创建的一个自动化视觉语言模型评估框架。该数据集旨在通过生成相关的图像样本和视觉问答任务,灵活高效地评估大型视觉语言模型(LVLMs)的特定能力。数据集包含多个难度级别的图像描述和问答任务,涵盖了基础理解、空间理解、语义理解、推理能力和大气理解等多个维度。通过该数据集,研究者可以全面评估LVLMs在不同视觉任务中的表现,揭示其在抽象理解和细节推理方面的优势与不足,为未来的研究提供重要参考。
AUTOBENCH-V is an automated visual-language model evaluation framework developed by King Abdullah University of Science and Technology (KAUST) and other institutions. This dataset aims to flexibly and efficiently evaluate the specific capabilities of large vision-language models (LVLMs) by generating relevant image samples and visual question answering (VQA) tasks. It includes image captioning and question answering tasks across multiple difficulty levels, covering multiple dimensions such as basic comprehension, spatial comprehension, semantic comprehension, reasoning ability, and atmospheric comprehension. Through this dataset, researchers can comprehensively assess the performance of LVLMs in various visual tasks, reveal their strengths and weaknesses in abstract comprehension and fine-grained reasoning, and provide valuable references for future research.

- 1AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?南本德大学、MBZUAI、KAUST · 2024年



