ColorBlindnessEval
收藏资源简介:
ColorBlindnessEval数据集由Apply U机构创建,旨在评估视觉语言模型(VLMs)在视觉对抗场景中的鲁棒性,灵感来源于石原色盲测试。该数据集包含500张类似于石原测试的图像,每张图像中都嵌入了一个从0到99的数字,颜色组合各不相同,旨在挑战VLMs准确识别复杂视觉模式中嵌入的数字信息。数据集的创建过程分为三个阶段:首先生成包含数字的参考图像;然后使用蒙特卡洛方法生成无颜色的圆盘;最后根据参考图像中圆盘的位置分配颜色。该数据集为评估和提高VLMs在实际应用中的可靠性和安全性提供了有价值的工具。
The ColorBlindnessEval dataset was created by the Apply U institution, aiming to evaluate the robustness of Vision-Language Models (VLMs) in visual adversarial scenarios, with inspiration drawn from Ishihara color blindness tests. This dataset contains 500 Ishihara test-like images, each embedding a number ranging from 0 to 99 with varying color combinations, designed to challenge VLMs to accurately recognize the embedded numerical information in complex visual patterns. The dataset's creation process is divided into three stages: first, generate reference images containing embedded numbers; second, use the Monte Carlo method to generate colorless disks; finally, assign colors based on the positions of the disks in the reference images. This dataset offers a valuable tool for evaluating and improving the reliability and safety of VLMs in real-world applications.
ColorBlindnessEval 数据集概述
数据集简介
ColorBlindnessEval 是一个新颖的基准测试数据集,旨在评估视觉语言模型在受石原色盲测试启发的视觉对抗场景中的鲁棒性。
数据集内容
- 图像数量:500 张
- 图像类型:类石原色盲测试图像
- 特征内容:包含从 0 到 99 的数字,具有不同的颜色组合
- 挑战目标:要求视觉语言模型准确识别嵌入复杂视觉模式中的数字信息
评估方法
- 评估模型数量:9 个视觉语言模型
- 提示类型:是/否提示和开放式提示
- 对比基准:与人类参与者的表现进行比较
主要发现
- 模型在对抗性环境中解释数字的能力存在局限性
- 存在普遍的幻觉问题
- 突显了在复杂视觉环境中提高视觉语言模型鲁棒性的必要性
应用价值
- 作为基准测试工具,用于评估和提高视觉语言模型在现实应用中的可靠性
- 适用于对准确性要求较高的关键应用场景
发布信息
- 数据集上传时间:2025 年 4 月 27 日
- 学术认可:已被 ICLR 研讨会(开放科学基础模型)接受
相关资源
- 论文地址:https://github.com/ApplyU-ai/ColorBlindnessEval/
- 数据集地址:https://huggingface.co/datasets/Apply-U/ColorBlindnessEval
联系方式
- 联系人:zijian.ling@applyu.ai
- 交流内容:研究合作或相关对话




