OPENRM
收藏资源简介:
OPENRM是一个利用外部工具对知识密集型长文本进行评价的工具增强型奖励模型。该模型通过调用外部工具来收集相关证据,对开放式的长文本回答进行系统性的评价。OPENRM在超过27K个合成的成对示例上进行了训练,这些示例是通过可控的数据合成框架生成的。训练目标同时监督中间工具的使用和最终结果的准确性,激励奖励模型学习基于证据的有效判断策略。在三个新收集的数据集和两个广泛使用的基准上进行的广泛实验表明,OPENRM明显优于现有的奖励模型。
OPENRM is a tool-augmented reward model that leverages external tools to evaluate knowledge-intensive long texts. This model collects relevant evidence by invoking external tools to conduct systematic evaluations of open-ended long-text answers. OPENRM is trained on over 27K synthetic paired examples, which are generated via a controllable data synthesis framework. The training objective supervises both the usage of intermediate tools and the accuracy of final outputs, incentivizing the reward model to learn evidence-based effective judgment strategies. Extensive experiments conducted on three newly collected datasets and two widely used benchmarks demonstrate that OPENRM significantly outperforms existing reward models.




