inclusionAI/ZwZ-RL-VQA
收藏资源简介:
ZwZ-RL-VQA是一个通过区域到图像蒸馏(R2I)方法生成的合成数据集,专门用于训练多模态大语言模型(MLLMs)进行细粒度感知任务,而无需在测试时使用工具。该数据集包含37,000个样本,源图像来自SA-1B、LAION、MetaCLIP、Visual Genome、CC12M和STPLS3D等多个数据集,图像分辨率大多高于1000×1000像素,裁剪区域通常小于完整图像面积的10%。数据集包含多种问题类型,如计数、OCR、颜色、结构、材料和识别等。数据生成过程使用了强大的教师模型(Qwen3-VL-235B和GLM-4.5V)进行问题生成和回答生成,并通过严格的共识过滤(>75%教师模型一致同意)和质量控制步骤确保数据质量。数据集主要用于多模态大语言模型的强化学习研究,以及将工具使用能力蒸馏到单次传递模型的研究。
ZwZ-RL-VQA is a synthetic dataset generated via the Region-to-Image (R2I) distillation method, specifically tailored for training multimodal large language models (MLLMs) on fine-grained perception tasks without relying on tools during inference. This dataset consists of 37,000 samples, with source images sourced from multiple datasets including SA-1B, LAION, MetaCLIP, Visual Genome, CC12M and STPLS3D. Most of the source images have a resolution exceeding 1000×1000 pixels, and the cropped regions typically account for less than 10% of the total image area. The dataset encompasses a wide range of question types, such as counting, OCR, color, structure, material and recognition-related tasks. During the data generation pipeline, two powerful teacher models (Qwen3-VL-235B and GLM-4.5V) are utilized for both question generation and answer generation. Strict consensus filtering (requiring >75% agreement across all teacher models) and quality control procedures are adopted to guarantee data quality. This dataset is mainly applied to reinforcement learning research on multimodal large language models, as well as studies on distilling tool-use capabilities into single-pass models.




