ChartInsighter Benchmark
收藏资源简介:
ChartInsighter Benchmark是由复旦大学研究团队创建的高质量时间序列图表摘要数据集,旨在解决时间序列图表摘要生成中的幻觉问题。该数据集包含75对图表和摘要,总计2693个句子,每对图表生成4个摘要,包括手动创建的金标准摘要、ChartInsighter生成的摘要、VL2NL生成的摘要和GPT-4生成的摘要。数据集通过多代理协作和自我一致性测试方法生成摘要,并在句子级别标注了幻觉类型,便于评估减少幻觉的效果。数据集的应用领域主要集中在时间序列数据的可视化摘要生成,旨在提高摘要的准确性和语义丰富性,减少幻觉现象,帮助决策者更高效地理解图表数据。
ChartInsighter Benchmark is a high-quality time-series chart summarization dataset developed by the research team at Fudan University, which targets solving the hallucination issue in time-series chart summarization generation. This dataset consists of 75 chart-summary pairs, with a total of 2693 sentences. Each chart is associated with four summaries, namely manually created gold-standard summaries, summaries generated by ChartInsighter, summaries generated by VL2NL, and summaries generated by GPT-4. The summaries are generated through multi-agent collaboration and self-consistency testing approaches, and hallucination types are annotated at the sentence level to facilitate the evaluation of hallucination mitigation effects. The main application scope of this dataset lies in visual summarization generation for time-series data, with the goals of improving the accuracy and semantic richness of summaries, reducing hallucination phenomena, and enabling decision-makers to comprehend chart data more efficiently.
ChartInsighter 数据集概述
数据集简介
ChartInsighter 是一个用于减少时间序列图表摘要生成中的幻觉(hallucination)的基准数据集。该数据集包含75对时间序列折线图及其对应的摘要,总计2,693个句子,涵盖了简单、中等和复杂三个复杂度级别。每对图表数据包括四种模态:图像、CSV文件、Vega-Lite规范、手动创建的金标准摘要、由ChartInsighter生成的摘要、VL2NL生成的摘要和GPT-4生成的摘要。所有摘要的句子级别都标注了幻觉类型,旨在评估减少幻觉的有效性。
幻觉类型
数据集总结了10种在生成时间序列数据摘要时可能出现的幻觉类型,并对每种类型进行了定义和标注。这些幻觉类型包括:
- 极值错误(Extremum Error):错误地将局部极值描述为绝对最大值或最小值。
- 数值错误(Numerical Value Error):在描述或计算定量数据时出现差异。
- 趋势方向错误(Trend Direction Error):错误地识别趋势的方向。
- 多维趋势错误(Multidimensional Trend Error):将相同或对比的趋势/关系混淆。
- 范围错误(Range Error):错误地识别趋势的开始和结束时间。
- 周期性错误(Cyclicality Error):将非周期性趋势错误地解释为周期性。
- 稳定性错误(Stability Error):将波动的趋势错误地描述为稳定,或将稳定的数据错误地描述为波动。
- 细节遗漏(Detail Omission):在特定范围内泛化数据,忽略关键波动和转折点。
- 垃圾描述(Junk Description):使用过于宽泛的描述,未能指定关键细节。
- 比例感知错误(Proportion Perception Error):在描述波动时,错误地使用“显著”等术语。
基准评估内容
该基准数据集可用于评估以下内容:
- 生成摘要的幻觉率:通过标注的幻觉类型,评估生成摘要中幻觉的出现频率。
- 生成摘要的语义丰富度:评估生成摘要的语义丰富程度。
相关论文
该数据集的相关论文《ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset》已被IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis 2025)接受,预计于2025年发表。

- 1ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset复旦大学 · 2025年



