MMPR-Tiny
收藏资源简介:
MMPR-Tiny是一个用于视觉问答任务的数据集,包含图像、问题、选择的答案和拒绝的答案等特征。它是InternVL3.5模型在在线强化学习阶段的训练数据,能够显著提高模型在不同规模下的推理能力。
MMPR-Tiny is a dataset for visual question answering (VQA) tasks, which includes features such as images, questions, selected answers, and rejected answers. It serves as the training data for the InternVL3.5 model during its online reinforcement learning stage, and can significantly enhance the model's reasoning capabilities across different scales.
MMPR-Tiny数据集概述
数据集基本信息
- 许可证:MIT
- 任务类别:视觉问答
- 语言:英语
- 数据集名称:MMPR-Tiny
- 数据规模:100万到1000万样本之间
数据特征
- 图像:字符串类型
- 问题:字符串类型
- 选择答案:字符串类型
- 拒绝答案:字符串类型
数据集描述
该数据集基于MMPR-v1.2数据集构建,通过计算每个查询的准确率并筛选模型准确率在0.2到0.8之间的样本用于在线强化学习。为进一步增强多样性,还扩展了最新的多模态数据集。
应用场景
该训练数据用于InternVL3.5模型的在线强化学习阶段,显著提升了InternVL3.5所有规模模型的整体性能。具体应用于:
- InternVL3.5-MPO模型:基于InternVL3.5-Instruct初始化,使用MPO方法在MMPR-v1.2上微调
- InternVL3.5-CascadeRL模型:基于InternVL3.5-MPO初始化,使用GSPO方法在MMPR-Tiny上进一步微调
相关资源
- 训练代码:https://github.com/Weiyun1025/verl-internvl
- 技术文档:https://internvl.readthedocs.io/en/latest/internvl3.0/preference_optimization.html
- 相关论文:https://huggingface.co/papers/2508.18265
引用信息
BibTeX @article{wang2024mpo, title={Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization}, author={Wang, Weiyun and Chen, Zhe and Wang, Wenhai and Cao, Yue and Liu, Yangzhou and Gao, Zhangwei and Zhu, Jinguo and Zhu, Xizhou and Lu, Lewei and Qiao, Yu and Dai, Jifeng}, journal={arXiv preprint arXiv:2411.10442}, year={2024} }




