ChartM60k
收藏资源简介:
ChartM60k 是一个用于评估多模态大语言模型(MLLMs)在图表理解任务中表现的数据集。该数据集从 MegaCQA 中提取了 60,000 个样本,覆盖了 21 种图表类型(如折线图、散点图、桑基图等)和 11 种任务分类(包括视觉理解、数值分析和逻辑推理等)。数据集旨在支持高层次开放式任务,如空间识别、多步推理和布局优化。评估指标包括关键词准确率(KAcc)、数值准确率(NAcc)、推理时间(time)、推理标记(token)和推理漂移(drift)。此外,数据集还配备了交互式视觉分析系统 ChartMLens,用于动态探索推理模式和多代理诊断框架。该数据集适用于图表问答(ChartQA)基准测试和 MLLMs 的推理诊断。
ChartM60k is a dataset for evaluating the performance of multimodal large language models (MLLMs) on chart understanding tasks. It consists of 60,000 samples extracted from MegaCQA, covering 21 chart types (e.g., line charts, scatter plots, Sankey diagrams, etc.) and 11 task categories including visual understanding, numerical analysis, logical reasoning, etc. The dataset is designed to support high-level open-ended tasks such as spatial recognition, multi-step reasoning, and layout optimization. Evaluation metrics include Keyword Accuracy (KAcc), Numerical Accuracy (NAcc), inference time, inference tokens, and inference drift. In addition, the dataset is equipped with an interactive visual analysis system ChartMLens for dynamically exploring inference patterns and multi-agent diagnostic frameworks. This dataset is suitable for chart question answering (ChartQA) benchmarking and inference diagnosis of MLLMs.




