SRM (Step-wise, Multi-dimensional Reward Model)
收藏资源简介:
SRM数据集是虚拟代理人领域中首个针对逐步、多维奖励模型训练和评估的基准。该数据集由两部分组成:用于训练的SRMTrain和用于评估的SRMEval。SRMTrain包含78,000个自动注释的数据点,而SRMEval包含32,000个经过精心挑选的测试数据点。数据集涵盖了Web、Linux、Windows和Android等多个平台,通过自动收集和注释的方式,为研究虚拟代理人奖励模型提供了丰富的资源。
The SRM dataset is the first benchmark for step-by-step and multi-dimensional reward model training and evaluation in the field of virtual agents. This dataset consists of two parts: SRMTrain for training and SRMEval for evaluation. SRMTrain contains 78,000 automatically annotated data points, while SRMEval includes 32,000 carefully curated test data points. The dataset covers multiple platforms including Web, Linux, Windows, and Android, and provides abundant resources for research on virtual agent reward models through automatic collection and annotation.

- 1Boosting Virtual Agent Learning and Reasoning: A Step-wise, Multi-dimensional, and Generalist Reward Model with Benchmark浙江大学, 中国杭州 · 2025年



