ViHallu-Instruction
收藏资源简介:
ViHallu-Instruction数据集由上海大学的研究团队创建,旨在解决大型视觉语言模型(LVLMs)中视觉幻觉问题。该数据集包含经过精心策划的视觉变化图像,通过引入可控的视觉变化,同时保持整体图像结构,帮助LVLMs更好地理解细粒度的视觉内容。数据集还包含了高质量的指令数据,用于指导LVLMs进行细粒度视觉语义对齐。ViHallu-Instruction数据集的创建过程结合了文本引导和分割掩码控制,生成了符合指定标题并保持原始图像全局结构的视觉变化图像。该数据集适用于LVLMs的幻觉缓解和视觉语义对齐研究,旨在提升LVLMs在视觉理解方面的性能。
The ViHallu-Instruction dataset was created by the research team from Shanghai University, aiming to address the visual hallucination issue in Large Vision-Language Models (LVLMs). This dataset includes carefully curated visually altered images, where controlled visual changes are introduced while preserving the overall image structure, to help LVLMs better understand fine-grained visual content. The dataset also contains high-quality instruction data for guiding LVLMs to achieve fine-grained visual-semantic alignment. The construction of the ViHallu-Instruction dataset combines text guidance and segmentation mask control to generate visually altered images that conform to the specified captions and maintain the global structure of the original images. This dataset is applicable to research on hallucination mitigation and visual-semantic alignment for LVLMs, with the goal of improving the visual understanding performance of LVLMs.
ViHallu数据集概述
数据集状态
- 代码和完整数据集将于下个月发布
其他信息
- 作者当前正在寻找PhD职位




