WorldVQA
收藏资源简介:
WorldVQA 是一个精心设计的基准数据集,旨在评估多模态大语言模型(MLLMs)中的原子视觉世界知识。该数据集包含 3,000 个视觉问答对,涵盖 8 个类别,特别关注语言和文化多样性。原始基准包含 3,500 个视觉问答对和 9 个类别,但由于版权问题,本次发布移除了'人物'类别。WorldVQA 的主要目标是严格衡量模型对视觉事实的记忆能力,而非推理能力,从而为评估当前和下一代前沿模型的百科全书广度和幻觉率建立标准。数据集支持英语和中文,适用于视觉问答任务。
WorldVQA is a meticulously designed benchmark dataset aimed at evaluating atomic visual world knowledge in multimodal large language models (MLLMs). This dataset contains 3,000 visual question-answer pairs spanning 8 categories, with a particular focus on linguistic and cultural diversity. The original benchmark included 3,500 visual question-answer pairs across 9 categories, but the "Person" category was removed in this release due to copyright issues. The core objective of WorldVQA is to rigorously measure models' memorization capability of visual facts rather than their reasoning ability, thereby establishing a standard for evaluating the encyclopedic breadth and hallucination rate of current and next-generation state-of-the-art models. The dataset supports both English and Chinese, and is applicable for visual question answering tasks.




