遇见数据集

UCSC-VLAA/VLM-CapCurriculum-Perception-Data

收藏
Hugging Face2026-05-20 更新2026-06-14 收录
官方服务:

资源简介:

VLM-CapCurriculum-Perception (D_perc) 是一个用于视觉语言模型分阶段后训练的第一阶段视觉感知数据集,源自ICML 2026论文《从看到思考:解耦感知与推理改进视觉语言模型的后训练》。该数据集包含3,360个训练样本,每个样本是一个基于图像的4项选择题,其设计核心是问题可以通过细粒度图像描述回答,但强大的视觉语言模型仅依赖图像时无法正确回答,从而有效隔离了感知失败与推理失败。每个样本还提供了预计算的pass_rate(通过率),用于衡量样本难度,支持基于能力×难度的课程学习实验。数据图像来源于DOCCI数据集(经过2倍下采样),难度信号通过Qwen3-VL-8B-Instruct模型的16次rollouts计算得出。数据集文件包括perception_difficulty_curriculum.jsonl(包含问题、答案、图像路径、预测结果、正确性和通过率等字段)和images/目录(包含压缩的图像文件)。该数据集旨在用于视觉语言模型的感知能力强化训练,特别是通过课程学习优化模型性能。

VLM-CapCurriculum-Perception (D_perc) is the first-stage visual perception dataset for staged post-training of vision-language models, derived from the ICML 2026 paper "From Seeing to Thinking: Decoupling Perception and Reasoning to Improve Post-Training of Vision-Language Models". This dataset contains 3,360 training samples, each being an image-based 4-option multiple-choice question. The core design of these samples is that the questions can be answered via fine-grained image descriptions, but state-of-the-art vision-language models cannot correctly answer them solely relying on the input images, which effectively isolates perception failures from reasoning failures. Each sample also provides a pre-computed pass_rate to measure sample difficulty, supporting ability × difficulty-based curriculum learning experiments. The images in the dataset are sourced from the DOCCI dataset (2x downsampled), and the difficulty signals are calculated via 16 rollouts of the Qwen3-VL-8B-Instruct model. The dataset files include perception_difficulty_curriculum.jsonl (containing fields such as question, answer, image path, prediction result, correctness, and pass_rate) and the images/ directory (containing compressed image files). This dataset is intended for enhanced training of perception capabilities for vision-language models, particularly to optimize model performance through curriculum learning.

提供机构:
UCSC-VLAA
二维码
社区交流群
二维码
科研交流群
商业服务