IllusionBench
收藏资源简介:
IllusionBench是由上海交通大学图像通信与网络工程研究所创建的一个大规模视觉幻觉理解基准数据集。该数据集包含1051张图像、5548个问答对和1051个黄金文本描述,涵盖了经典认知幻觉、真实场景幻觉、陷阱幻觉等多种类型。数据集通过手动标注的问答对和图像描述,详细记录了幻觉的存在、原因和内容。IllusionBench旨在评估视觉语言模型(VLMs)在真实场景中对视觉幻觉的理解能力,并揭示模型在幻觉识别中的局限性。数据集的应用领域主要集中在视觉语言模型的性能评估和幻觉理解能力的提升,旨在解决模型在复杂视觉场景中的幻觉识别和解释问题。
IllusionBench is a large-scale visual illusion understanding benchmark dataset created by the Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University. This dataset contains 1051 images, 5548 question-answer pairs, and 1051 gold reference text descriptions, covering multiple types including classic cognitive illusions, real-scene illusions, and trap illusions. Through manually annotated question-answer pairs and image descriptions, the dataset comprehensively documents the existence, underlying causes and specific content of visual illusions. The core objective of IllusionBench is to evaluate the visual illusion understanding capabilities of vision-language models (VLMs) in real-world scenarios, and to reveal the limitations of these models in illusion recognition. The main application areas of this dataset focus on the performance evaluation of vision-language models and the enhancement of their illusion understanding abilities, aiming to solve the problems of illusion recognition and explanation of models in complex visual scenes.




