CHARTMUSEUM
收藏资源简介:
CHARTMUSEUM 是一个图表问答基准数据集,旨在评估大型视觉语言模型(LVLMs)在真实图表上的复杂视觉和文本推理能力。该数据集由13位计算机科学研究人员创建,包含1162个(图像、问题、答案)元组,这些元组来源于928个独特的真实世界图像,跨越了184个网站。与先前基准数据集不同,CHARTMUSEUM 中的所有问题都是由研究人员手动策划的,没有使用大型语言模型(LLMs)的辅助。数据集中的问题涵盖了多种推理类型,包括文本推理、视觉推理、文本/视觉推理和综合推理。通过评估10个开源模型和11个专有模型,CHARTMUSEUM 暴露了模型和人类性能之间的巨大差距,表明现有的 LVLMs 在视觉推理方面存在显著不足。
CHARTMUSEUM is a chart question answering benchmark dataset developed to evaluate the complex visual and textual reasoning capabilities of large vision-language models (LVLMs) on real-world charts. Created by 13 computer science researchers, this dataset contains 1,162 (image, question, answer) tuples sourced from 928 unique real-world images across 184 websites. Unlike prior benchmark datasets, all questions in CHARTMUSEUM are manually curated by researchers without the assistance of large language models (LLMs). The questions in the dataset cover multiple reasoning types, including textual reasoning, visual reasoning, textual-visual hybrid reasoning, and comprehensive reasoning. By evaluating 10 open-source models and 11 proprietary models, CHARTMUSEUM uncovers a considerable gap between model performance and human performance, demonstrating that existing LVLMs have significant deficiencies in visual reasoning.




