遇见数据集

Relation-Associated Instructions & Hallucination Benchmark

收藏
DataCite Commons2024-07-08 更新2024-07-13 收录
官方服务:

资源简介:

Large vision-language models (LVLMs) suffer from hallucination, generating responses that apparently contradict to the image content occasionally. The key problem lies in its weak ability to comprehend detailed content in multi-modal contexts, which can be mainly attributed its training data. The vision instruction dataset primarily focuses on global description that are highly relevant to the image, with few samples containing image details. Therefore, we construct a fine-grained vision instruction dataset, RAI-30k, by generate image-text pairs with detailed relationship annotations in panoptic scene graph dataset (PSG). These conversations pay more attention on detailed facts in the image, encouraging the model to answer questions based on multi-modal contexts. Moreover, to provide a deeper evaluation on the hallucination in LVLMs, we propose a new benchmark, RAH-Bench. It divides vision hallucination into three different types that contradicts the image with wrong categories, attributes or relations, and introduces False Positive Rate as detailed sub-metric for each type. We hope the provided dataset and benchmark will benefit the future research in large vision-language models.

大型视觉语言模型(Large Vision-Language Models,LVLMs)存在幻觉问题,偶尔会生成与图像内容明显相悖的回复。其核心症结在于模型对多模态场景下的细节内容理解能力不足,这一问题主要可归咎于其训练数据。现有视觉指令数据集多聚焦于与图像高度相关的全局描述,仅含少量涵盖图像细节的样本。为此,我们基于全景场景图数据集(Panoptic Scene Graph Dataset,PSG)构建了细粒度视觉指令数据集RAI-30k,通过生成带有详细关系标注的图像-文本对。该数据集生成的对话更关注图像中的细节事实,旨在引导模型基于多模态语境完成问答。此外,为更深入地评估大型视觉语言模型的幻觉问题,我们提出了全新基准测试集RAH-Bench:该基准将视觉幻觉划分为三类与图像事实相悖的情况——类别错误、属性错误或关系错误,并针对每一类引入假阳性率(False Positive Rate)作为细分评估指标。我们期望本数据集与基准测试集能够为大型视觉语言模型的后续研究提供助力。

提供机构:
IEEE DataPort
创建时间:
2024-07-08
搜集汇总
背景与挑战
背景概述
该数据集旨在解决大型视觉语言模型的幻觉问题,包含细粒度视觉指令数据集RAI-30k和幻觉评估基准RAH-Bench。RAI-30k通过全景场景图数据生成带有详细关系标注的图像-文本对,强调图像细节理解;RAH-Bench则将幻觉分为类别、属性和关系错误三类,并提供误报率子指标进行量化评估。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务