RAG-RewardBench
收藏资源简介:
RAG-RewardBench是由中国科学院自动化研究所复杂系统认知与决策智能实验室创建的,用于评估检索增强生成(RAG)场景中奖励模型的基准数据集。该数据集包含1485个高质量的偏好对,涵盖了18个子集、6种检索器和24种RAG模型,旨在提高数据源的多样性。数据集的创建过程包括设计四个关键的RAG特定场景,并通过LLM-as-a-judge方法提高偏好标注的效率和有效性。RAG-RewardBench主要应用于检索增强语言模型的偏好对齐,旨在解决现有模型在偏好对齐方面的不足,推动模型向偏好对齐训练的转变。
RAG-RewardBench is a benchmark dataset developed by the Laboratory of Complex System Cognition and Decision Intelligence, Institute of Automation, Chinese Academy of Sciences, for evaluating reward models in retrieval-augmented generation (RAG) scenarios. This dataset comprises 1,485 high-quality preference pairs, covering 18 subsets, 6 retrievers and 24 RAG models, with the goal of enhancing the diversity of data sources. The construction process of the dataset includes designing four key RAG-specific scenarios, and utilizing the LLM-as-a-judge approach to improve the efficiency and validity of preference annotation. RAG-RewardBench is primarily designed for preference alignment of retrieval-augmented language models, aiming to address the limitations of existing models in preference alignment and promote the shift of model training toward preference-aligned training.




