GeoSym-Bench
收藏资源简介:
GeoSym-Bench 是一个用于定量图表推理的受控评估基准,其核心特点是支持细粒度的错误归因。该数据集是论文《GeoSym: Structured Evidence Fusion for Reliable Quantitative Chart Reasoning》的配套基准,旨在评估模型在理解和推理数学函数图表方面的能力。数据集包含 460 个高分辨率(1440 x 1440 像素)的干净渲染数学函数图图像,这些图像覆盖了 8 个不同的函数族,包括线性、二次、绝对值、三次、分段线性、指数、对数和有理函数。每个图像样本都配有完整的代数真实值标注,这些标注横跨 10 个评估维度:截距、极值点、预测值、变化率、单调性、对称性、函数表达式、定义域、值域以及推理质量(前9项为客观评估,最后1项为主观评判)。数据以图像文件(PNG格式)和对应的结构化标注文件(JSON格式)组织。JSON文件详细包含了图表渲染元数据、坐标轴范围、刻度值、精确的曲线-网格交点坐标、函数参数与表达式,以及针对10个维度的标准答案。GeoSym-Bench 专为多模态图表推理、视觉问答和图像到文本生成等任务设计,尤其侧重于定量和数学推理能力。其独特的价值在于能够将模型在图表推理过程中产生的错误,精确地归因到视觉感知、代数转换或逻辑推理等特定处理阶段,从而为模型诊断和改进提供更深入的洞察。数据集的问题文本为英语,但支持中英双语环境。
GeoSym-Bench is a controlled evaluation benchmark for quantitative chart reasoning, with its core feature being support for fine-grained error attribution. This dataset accompanies the paper GeoSym: Structured Evidence Fusion for Reliable Quantitative Chart Reasoning and aims to evaluate models capabilities in understanding and reasoning about mathematical function charts. The dataset contains 460 high-resolution (1440 x 1440 pixels) cleanly rendered mathematical function graph images, covering 8 different function families, including linear, quadratic, absolute value, cubic, piecewise linear, exponential, logarithmic, and rational functions. Each image sample comes with complete algebraic ground truth annotations spanning 10 evaluation dimensions: intercepts, extrema, predictions, rates of change, monotonicity, symmetry, function expressions, domain, range, and reasoning quality (the first 9 are objective assessments, and the last is subjective judgment). The data is organized as image files (PNG format) and corresponding structured annotation files (JSON format). The JSON files include detailed chart rendering metadata, axis ranges, tick values, precise curve-grid intersection coordinates, function parameters and expressions, and standard answers for the 10 dimensions. GeoSym-Bench is designed for tasks such as multimodal chart reasoning, visual question answering, and image-to-text generation, with a particular focus on quantitative and mathematical reasoning abilities. Its unique value lies in the ability to precisely attribute errors generated by models during chart reasoning to specific processing stages such as visual perception, algebraic transformation, or logical reasoning, thereby providing deeper insights for model diagnosis and improvement. The datasets question text is in English, but it supports both Chinese and English environments.




