VL-RewardBench
收藏资源简介:
VL-RewardBench是由香港大学、华南理工大学、上海交通大学、北京大学、华盛顿大学和艾伦人工智能研究院联合创建的一个综合视觉语言生成奖励模型基准数据集。该数据集包含1250个高质量样本,涵盖了多模态查询、视觉幻觉检测和复杂推理任务。数据集的创建过程结合了AI辅助的样本选择和人工验证,旨在全面评估和挑战现有的视觉语言模型。VL-RewardBench主要应用于多模态AI系统的对齐和评估,旨在解决当前评估方法中的偏见和不足,推动视觉语言生成奖励模型的发展。
VL-RewardBench is a comprehensive visual-language generation reward model benchmark dataset jointly created by The University of Hong Kong, South China University of Technology, Shanghai Jiao Tong University, Peking University, University of Washington, and Allen Institute for AI. It contains 1,250 high-quality samples covering multimodal queries, visual hallucination detection and complex reasoning tasks. The dataset construction process combines AI-assisted sample selection and manual verification, aiming to comprehensively evaluate and challenge existing visual-language models. VL-RewardBench is primarily applied to the alignment and evaluation of multimodal AI systems, and is designed to address the biases and limitations in current evaluation methods, so as to advance the development of visual-language generation reward models.




