HalluSegBench
收藏资源简介:
HalluSegBench是一个用于评估视觉语言分割中像素级幻觉的基准数据集。该数据集由1340个事实-反事实图像对组成,跨越281个独特的物体类别,旨在通过反事实视觉推理来评估模型对像素级幻觉的敏感性。数据集的创建过程是通过对RefCOCO数据集中的图像进行精确的对象替换,生成视觉上连贯但语义上不同的图像对。该数据集广泛应用于视觉语言模型分割模型的评估,旨在解决模型在复杂视觉场景中像素级理解时的幻觉问题。
HalluSegBench is a benchmark dataset for evaluating pixel-level hallucinations in vision-language segmentation tasks. It consists of 1,340 fact-counterfactual image pairs spanning 281 unique object categories, designed to assess a model's sensitivity to pixel-level hallucinations via counterfactual visual reasoning. The dataset is constructed by precisely replacing objects in images from the RefCOCO dataset to generate visually coherent yet semantically distinct image pairs. This benchmark is widely used for evaluating vision-language segmentation models, aiming to address the hallucination issue when models perform pixel-level comprehension in complex visual scenarios.
HalluSegBench: Counterfactual Visual Reasoning for Segmentation
概述
- 数据集名称: HalluSegBench
- 研究领域: 视觉语言分割中的幻觉评估
- 开发团队: PLAN Lab, University of Illinois Urbana-Champaign
- 主要作者: Xinzhuo Li, Adheesh Juvekar (共同第一作者), Xingyou Liu, Muntasir Wahed, Kiet A. Nguyen, Ismini Lourentzou
核心贡献
-
新基准测试:
- 首个使用反事实图像-文本对评估分割幻觉的基准
- 包含1,340个反事实实例对,涵盖281个对象类别
-
新评估指标:
- 引入4个新指标量化视觉/文本反事实下的幻觉严重程度
- 评估语义先验的过度依赖和幻觉掩码的空间合理性
-
实证发现:
- 现有视觉语言分割模型在视觉编辑下比文本编辑更容易产生幻觉
- 模型经常持续产生错误分割,凸显反事实推理诊断的必要性
数据集特点
- 数据规模: 1,340个反事实实例对
- 类别覆盖: 281个独特对象类别
- 评估维度:
- 文本和视觉IoU下降(
ΔIoUtextual,ΔIoUvisual) - 事实和反事实混淆掩码评分(
CMS) - 对比幻觉指标(
CCMS)
- 文本和视觉IoU下降(
相关资源
- 论文: arXiv:2506.21546
- BibTeX引用: bibtex @article{li2025hallusegbench, title={HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation}, author={Li, Xinzhuo and Juvekar, Adheesh and Liu, Xingyou and Wahed, Muntasir and Nguyen, Kiet A and Lourentzou, Ismini}, journal={arXiv preprint arXiv:2506.21546}, year={2025} }




