UCSC-VLAA/VLM-CapCurriculum-VisualReasoning-Data
收藏资源简介:
VLM-CapCurriculum-VisualReasoning(D_vis)数据集是ICML 2026论文《从看到思考:解耦感知与推理提升视觉语言模型后训练》中用于分阶段后训练的第三阶段视觉推理数据。该数据集包含16,195个样本,混合了视觉数学和基于图形的推理内容,源自四个开源语料库:Math PUMA(合成数据)、GeoQA170K、CLEVR-Math和ArxivQA(2倍下采样)。每个样本都附带源图像,并预计算了pass_rate指标,该指标基于Qwen3-VL-8B-Instruct基础模型的16次rollout评估得出,用于表示样本难度,从而支持按难度排序的课程学习实验。数据集旨在支持视觉语言模型的视觉推理能力训练,特别是在分阶段后训练中提升推理性能。
VLM-CapCurriculum-VisualReasoning (D_vis) is a Stage-3 visual-reasoning dataset for the staged post-training recipe in the ICML 2026 paper From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models. It consists of 16,195 samples mixing visual math and figure-grounded reasoning, drawn from four open-source corpora: Math PUMA (synthesis), GeoQA170K, CLEVR-Math, and ArxivQA (2× downsampled). Each sample includes source images and a precomputed pass_rate derived from 16 rollouts of the Qwen3-VL-8B-Instruct base model, which indicates sample difficulty for capability × difficulty curriculum experiments. The dataset is designed to enhance visual reasoning in vision-language models through post-training.



