RodelaG/gsm8k-rendered-vlm
收藏资源简介:
GSM8K-VL 数据集是一个多模态数学推理数据集,专为视觉语言模型评估而设计。每个示例包含四个关键部分:一个来自GSM8K的文本数学问题(question)、最终数值答案(answer)、清理后的链式思维推理过程(reasoning),以及渲染图像路径(image)。该数据集旨在支持控制实验,以比较纯文本和基于图像的推理行为。数据来源于GSM8K测试集,图像通过数字渲染生成,无额外视觉提示,仅包含问题文本。评估主要基于最终答案的精确匹配准确率,推理部分用于分析而非评分。数据集规模在1K到10K之间,适用于多模态和数学推理研究。
Rendered GSM8K-VL is a multimodal math-reasoning dataset for vision-language model evaluation. Each example includes a GSM8K word problem (question), the final numeric answer (answer), cleaned chain-of-thought reasoning (reasoning), and a rendered image path (image). The dataset is intended for controlled experiments comparing text-only and image-based reasoning behavior. It is derived from the GSM8K test split, with images rendered digitally to contain only the question text without visual cues. Evaluation focuses on exact match accuracy of the final answer, with reasoning provided for analysis. The dataset size is between 1K and 10K samples and supports multimodal and math-reasoning research.




