ECD
收藏资源简介:
ECD数据集包含10k+图表图像和300k+问答对,覆盖25个主题和250+图表类型组合,具有高度的视觉复杂性。该数据集旨在提高多模态大语言模型(MLLM)在图表理解方面的性能。
The ECD dataset encompasses over 10k chart images and 300k+ question-answer pairs, spanning 25 topics and over 250 combinations of chart types, characterized by high visual complexity. This dataset is designed to enhance the performance of multi-modal large language models (MLLMs) in chart comprehension.
数据集概述
基本信息
- 数据集名称: Effective Chart Dataset (ECD)
- 发布年份: 2025
- 相关会议: IEEE/CVF International Conference on Computer Vision (ICCV)
- 数据集大小: 10k+ 图表图像,300k+ 问答对 (QA pairs)
- 覆盖主题: 25个主题
- 图表类型组合: 250+种
数据集内容
- 图表类型: 单图和多子图图表
- 视觉复杂度: 高
- 数据生成方法: 五步数据合成流程
- 分离数据和功能创建
- 多子图生成条件化
- 视觉多样化
- 低质量数据过滤
- 使用GPT-4o生成问答对
数据集用途
- 主要用途: 提升多模态大语言模型 (MLLM) 的图表理解能力
- 适用模型: 包括但不限于LLaVA-Next-Llama3-8B、MiniCPM-V2.6、Phi-3-Vision、Qwen2.5-VL-7B
数据集获取
- Hugging Face地址: https://huggingface.co/datasets/ChartFoundation/ECD-10k-Images
- 数据生成代码: 包含在GitHub仓库的
data_generation_pipeline目录中
基准测试 (ECDBench)
- 图表数量: 1,224张
- 单图图表: 364张
- 多子图图表: 860张
- 2种图表类型: 457张
- 3种图表类型: 403张
- 平均分辨率: 1378 × 968像素
- 问答对数量: 2,448对 (每图1描述性+1推理性问题)
性能提升
- LLaVA-Next-Llama3-8B: 平均性能从10.95提升至31.58
- MiniCPM-V2.6: 平均性能从27.53提升至35.17
- Phi-3-Vision: 平均性能从31.41提升至44.40
- Qwen2.5-VL-7B: 平均性能从38.19提升至50.86
引用格式
bibtex @inproceedings{yang2025effective, title={Effective Training Data Synthesis for Improving MLLM Chart Understanding}, author={Yang, Yuwei and Zhang, Zeyu and Hou, Yunzhong and Li, Zhuowan and Liu, Gaowen and Payani, Ali and Ting, Yuan-Sen and Zheng, Liang}, booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)}, year={2025} }
@article{yang2025effective, title={Effective Training Data Synthesis for Improving MLLM Chart Understanding}, author={Yang, Yuwei and Zhang, Zeyu and Hou, Yunzhong and Li, Zhuowan and Liu, Gaowen and Payani, Ali and Ting, Yuan-Sen and Zheng, Liang}, journal={arXiv preprint arXiv:2508.06492}, year={2025} }
许可信息
- 许可证类型: MIT License




