Misbehavior-Bench
收藏资源简介:
Misbehavior-Bench 是一个用于评估大型视觉语言模型(LVLM)在四种不同类别异常行为上的基准数据集。该数据集旨在通过证据不确定性量化的方法,检测模型在幻觉(Hallucination)、越狱(Jailbreak)、对抗攻击(Adversarial Attacks)和分布外失败(Out-of-Distribution, OOD Failures)等方面的表现。数据集包含四个主要子集,分别对应上述四种异常行为,每个子集包含CSV文件和相关的图像数据。该数据集是ICLR 2026论文《通过证据不确定性量化检测大型视觉语言模型的异常行为》的官方基准,适用于模型安全性改进和不确定性量化方法的验证。数据集规模介于10K到100K之间,语言为英文,涵盖多模态任务,并涉及AI安全、对抗攻击、鲁棒性等标签。
Misbehavior-Bench is a benchmark dataset for evaluating large vision-language models (LVLMs) across four distinct categories of anomalous behaviors. This dataset aims to detect model performance in scenarios including Hallucination, Jailbreak, Adversarial Attacks, and Out-of-Distribution (OOD) Failures via evidence uncertainty quantification methods. The dataset contains four main subsets corresponding to the four aforementioned anomalous behaviors, each including CSV files and associated image data. This dataset is the official benchmark for the ICLR 2026 paper titled "Detecting Anomalous Behaviors of Large Vision-Language Models via Evidence Uncertainty Quantification", and is applicable to model safety improvement and validation of uncertainty quantification methodologies. The dataset has a scale ranging from 10K to 100K, uses English as its primary language, covers multimodal tasks, and includes labels such as AI safety, adversarial attacks, and robustness.




