Tennet7/ZwZ-RL-VQA
收藏资源简介:
ZwZ-RL-VQA数据集包含74K个通过区域到图像蒸馏(R2I)生成的高质量视觉问答(VQA)对,用于训练多模态大语言模型(MLLMs)进行细粒度感知任务,而无需在测试时使用工具。该数据集采用“无需缩放的缩放(ZwZ)”方法,将“缩放”从推理时工具转变为训练时原语,包括微裁剪区域的放大合成和将区域监督蒸馏回完整图像的缩小蒸馏。数据集经过共识过滤、难度过滤和视觉 grounding 等质量控制,适用于MLLMs的强化学习(如DAPO/GRPO)和将工具使用能力蒸馏到单次推理模型的研究。
The ZwZ-RL-VQA dataset contains 74K high-quality VQA pairs generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. The dataset employs the Zooming without Zooming (ZwZ) method to transform zooming from an inference-time tool into a training-time primitive, involving zoom-in synthesis on micro-cropped regions and zoom-out distillation of region-grounded supervision back to full images. It undergoes quality control measures like consensus filtering, difficulty filtering, and visual grounding, and is intended for reinforcement learning on MLLMs (e.g., with DAPO/GRPO) and research on distilling tool-use capabilities into single-pass models.




