VIDHAL
收藏资源简介:
VIDHAL数据集由新加坡国立大学创建,专门用于评估视觉大语言模型(VLLMs)在视频中的幻觉问题。该数据集包含1000个视频实例,覆盖了广泛的时间维度,如实体动作和事件序列。每个视频通过自动注释生成多个带有不同幻觉级别的字幕,以捕捉细微和显著的差异。数据集的创建过程包括从现有数据集中选择视频实例,生成锚定字幕,并使用GPT-4o生成幻觉字幕。VIDHAL旨在解决视频内容中复杂的时间动态导致的幻觉问题,特别是在视频特定的时间方面,如运动方向和事件的时序。
The VIDHAL dataset was created by the National University of Singapore, specifically designed to evaluate hallucination issues in videos for Visual Large Language Models (VLLMs). This dataset comprises 1000 video instances covering a wide range of temporal dimensions, such as entity actions and event sequences. For each video, multiple subtitles with varying levels of hallucination are generated through automatic annotation to capture both subtle and significant differences. The dataset creation process includes selecting video instances from existing datasets, generating anchor subtitles, and producing hallucinatory subtitles using GPT-4o. VIDHAL aims to address hallucination problems caused by complex temporal dynamics in video content, particularly regarding video-specific temporal aspects such as motion direction and event timing.

- 1VidHal: Benchmarking Temporal Hallucinations in Vision LLMs新加坡国立大学 · 2024年



