MathVision-Mini
收藏资源简介:
MATH-Vision(MATH-V)是由香港中文大学多媒体实验室、商汤科技和上海人工智能实验室联合创建的高质量视觉数学推理数据集,旨在全面评估多模态模型在视觉情境下的数学推理能力。该数据集包含3,040个数学问题,涵盖16个数学学科(如代数、解析几何、组合几何等)和5个难度等级,数据来源为19个真实数学竞赛。所有问题均经过专家标注和验证,确保唯一且正确的答案。数据集分为开放式和多项选择题,提供丰富的视觉情境(如函数图像、几何图形等)以支持多模态推理。MATH-V的创建过程严格筛选和整理了大量竞赛题目,确保问题的多样性和挑战性。该数据集的应用领域主要集中在评估多模态模型的数学推理能力,特别是在视觉情境下的逻辑推理、几何理解以及符号计算等任务。通过MATH-V,研究人员可以深入分析模型在不同学科和难度等级下的表现,为未来多模态模型的发展提供重要参考。
MATH-Vision (MATH-V) is a high-quality visual mathematical reasoning dataset jointly developed by the Multimedia Laboratory of The Chinese University of Hong Kong, SenseTime, and Shanghai AI Laboratory. It is designed to comprehensively evaluate the mathematical reasoning abilities of multimodal models in visual contexts. This dataset comprises 3,040 mathematical problems spanning 16 mathematical disciplines (including algebra, analytic geometry, combinatorial geometry, etc.) and 5 difficulty tiers, with data sourced from 19 real-world mathematics competitions. All problems have been annotated and verified by domain experts to guarantee unique and accurate answers. The dataset includes both open-ended and multiple-choice questions, and provides rich visual contexts such as function plots, geometric figures, and more to enable multimodal reasoning. During the creation of MATH-V, a large number of competition problems were rigorously screened and curated to ensure the diversity and challenging nature of the dataset. The primary application scope of this dataset is the evaluation of multimodal models' mathematical reasoning capabilities, particularly tasks involving logical reasoning, geometric comprehension, and symbolic computation within visual scenarios. Using MATH-V, researchers can conduct in-depth analyses of model performance across different disciplines and difficulty levels, providing critical references for the future advancement of multimodal models.




