MRCL
收藏资源简介:
#### Introduction of MRCL Benchmark To rigorously evaluate continual reinforcement learning while reducing the interference of pre-training data contamination, we construct a new benchmark comprising datasets released in 2025 or later which called MRCL(Multimodal Reasoning Continual Learning). The benchmark is selected to remain challenging for state-of-the-art VLMs and to cover diverse forms of multimodal reasoning, including MedBookVQA, Navigation, We-Math2.0, Puzzle, FinMME. The figure below presents the query and answer formats. We invite you to explore our work on the MRCL benchmark: [RL Forgets! Towards Continual Policy Optimization](https://arxiv.org/abs/2607.04364). Code is available at: [https://github.com/MaolinLuo/CPO](https://github.com/MaolinLuo/CPO) #### MRCL数据集介绍 为了严格评估持续强化学习,同时减少预训练数据污染的干扰,我们构建了一个名为MRCL[1](Multimodal Reasoning Continual Learning)的全新基准测试集,其中包含的数据集均发布于2025年或之后。该基准测试集旨在对当前最先进的视觉语言模型(VLMs)仍具挑战性,并涵盖多种形式的多模态推理任务,包括MedBookVQA[2], Navigation[3], We-Math2.0[4], Puzzle[3], FinMME[6]。下图展示了查询与回答的格式。 欢迎关注我们在MRCL数据集上的工作:[RL Forgets! Towards Continual Policy Optimization](https://arxiv.org/abs/2607.04364). 代码地址:[https://github.com/MaolinLuo/CPO](https://github.com/MaolinLuo/CPO)  #### 下载方法 :modelscope-code[]{type="sdk"} :modelscope-code[]{type="git"} #### 引用 [1] Luo M L, Wang Z X, Zhou Z H, et al. RL Forgets! Towards Continual Policy Optimization[J]. arXiv preprint arXiv:2607.04364, 2026. [2] Yip S L, He S, Nie Y, et al. MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book[J]. arXiv preprint arXiv:2506.00855, 2025. [3] Gu J, Hao Y, Wang H W, et al. Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning[J]. arXiv preprint arXiv:2510.27492, 2025. [4] Qiao R, Tan Q, Yang P, et al. We-math 2.0: A versatile mathbook system for incentivizing visual mathematical reasoning[J]. arXiv preprint arXiv:2508.10433, 2025. [5] Luo J, Kou Z, Yang L, et al. Finmme: Benchmark dataset for financial multi-modal reasoning evaluation[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025: 29465-29489.



