VCBench
收藏资源简介:
VCBench是一个用于评估多模态数学推理能力的综合基准数据集,专注于评估模型解决小学数学问题的能力,这些问题依赖于明确的视觉依赖关系。数据集包含1720个问题,跨越六个认知领域,并配以6697张图片,平均每个问题包含3.9张图片。该数据集的创建过程严格筛选并翻译了中国小学数学教科书中的问题,确保每个问题都具有独特且明确的答案。VCBench旨在解决当前基准测试中缺乏对视觉-数学推理能力的充分评估的问题,为多模态数学推理的研究提供了宝贵的资源。
VCBench is a comprehensive benchmark dataset for evaluating multimodal mathematical reasoning capabilities, with a specific focus on assessing models' ability to solve elementary school mathematics problems that rely on explicit visual dependencies. The dataset includes 1,720 problems spanning six cognitive domains, paired with 6,697 images, averaging 3.9 images per problem. During its creation, problems sourced from Chinese elementary school mathematics textbooks were strictly screened and translated, ensuring that each problem has a unique and unambiguous answer. VCBench aims to address the gap in current benchmarks that lack sufficient evaluation of visual-mathematical reasoning capabilities, providing a valuable resource for multimodal mathematical reasoning research.

- 1Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency阿里巴巴集团 DAMO Academy, 湖畔实验室, 浙江大学, 新加坡科技与设计大学 · 2025年



