iieycx/rlsd-train-MMFineReason-123K
收藏资源简介:
该数据集是RLSD(自蒸馏RLVR)的训练数据,来源于MMFineReason数据集,旨在通过开放数据中心方法缩小多模态推理差距。原始数据经过处理,仅保留每个响应中`</think>`标签后的最终结论,并移除该标签前的冗长推理内容,从而使其更适合作为RLSD训练中的结论监督数据。数据集包括主要数据文件(如MMFineReason_data_with_conclusion.json)和图像资源(MMFineReason_images/),以及验证集文件(如vl_math_val_mini_data.json)和对应的图像存档(val_data_image.zip)。验证集包含来自12个数据源的700个示例,例如val_MMMU_Pro_10c、val_MathVerse_Text_Dominant等。
This dataset is the training data for RLSD (Self-Distilled RLVR), derived from the MMFineReason dataset, and aims to narrow the multimodal reasoning gap through open data center methods. The original raw data was processed to only retain the final conclusion after the `</think>` tag in each response, and remove the lengthy reasoning content before this tag, making it more suitable as conclusion supervision data for RLSD training. The dataset includes main data files (e.g., MMFineReason_data_with_conclusion.json), image resources (MMFineReason_images/), validation set files (e.g., vl_math_val_mini_data.json) and their corresponding image archive (val_data_image.zip). The validation set contains 700 examples from 12 data sources, such as val_MMMU_Pro_10c, val_MathVerse_Text_Dominant, etc.



