遇见数据集

two-tiger/VRPRM3.6K

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

VisualPRM400K Train for SFT Rollout H 是一个多模态推理数据集,专门用于过程奖励模型(PRM)和监督微调(SFT)的rollout训练。该数据集基于VisualPRM400K-v1.1的候选版本(VisualPRM400K-derived)构建,包含了从原始VisualPRM400K-v1.1-Raw/images中提取的图像。数据以对话式架构组织,每个样本包括多轮对话消息(messages)和相关的图像路径(images),其中图像路径已从本地绝对路径转换为相对于仓库的路径。数据集包含训练集,共有4,294个样本和3,821个独特图像,适用于视觉问答和图像文本到文本任务,支持多模态推理和过程奖励模型的研究与应用。

VisualPRM400K Train for SFT Rollout H is a multimodal reasoning dataset specifically designed for Process Reward Model (PRM) and Supervised Fine-Tuning (SFT) rollout training. Derived from the VisualPRM400K-derived release candidate, it incorporates images from VisualPRM400K-v1.1-Raw/images. The data is organized in a conversation-style schema, with each sample containing multi-turn conversation messages and corresponding repo-relative image paths. The dataset includes a train split with 4,294 samples and 3,821 unique images, targeting visual question answering and image-text-to-text tasks to support research and applications in multimodal reasoning and process reward modeling.

提供机构:
two-tiger
二维码
社区交流群
二维码
科研交流群
商业服务