AVHBench
收藏资源简介:
AVHBench是由韩国科学技术院的研究团队创建的音频-视觉大型语言模型幻觉基准数据集。该数据集包含5,816个问答对和1,238个音频-视觉描述,涵盖四个不同的任务:音频驱动的视频幻觉、视频驱动的音频幻觉、音频-视觉匹配和音频-视觉描述。数据集的创建过程包括从现有数据集中提取视频和音频信息,并通过半自动注释管道生成问答对。AVHBench旨在评估和提升音频-视觉LLMs在处理复杂多模态信号时的鲁棒性,特别是在减少跨模态幻觉方面。
AVHBench is a benchmark dataset for audio-visual large language model (LLM) hallucination, developed by a research team from the Korea Advanced Institute of Science and Technology (KAIST). It consists of 5,816 question-answer pairs and 1,238 audio-visual captions, spanning four distinct tasks: audio-driven video hallucination, video-driven audio hallucination, audio-visual matching, and audio-visual captioning. The dataset construction process involves extracting video and audio content from existing datasets, and generating question-answer pairs via a semi-automated annotation pipeline. AVHBench aims to evaluate and enhance the robustness of audio-visual LLMs when processing complex multimodal signals, particularly in reducing cross-modal hallucinations.




