RAGTruth
收藏资源简介:
RAGTruth数据集是由华中科技大学计算机科学与技术学院创建的,用于辅助检索增强生成模型(RAG)的训练。该数据集包含了模型的历史对话记录,旨在通过主动学习方法筛选出最有信息量的样本,进而构建偏好数据集,帮助模型学会拒绝可能导致虚构回答的查询,同时提高对有能力回答的查询的稳定性。数据集的具体内容和创建过程未在论文中详细描述,但提到了其在精炼大型语言模型(LLM)方面的应用,特别是在减少虚构回答和提高回答准确性方面的作用。
The RAGTruth dataset was developed by the School of Computer Science and Technology, Huazhong University of Science and Technology, to support the training of retrieval-augmented generation (RAG) models. This dataset includes historical conversation logs of models, with the goal of screening out the most informative samples via active learning approaches to build a preference dataset. It is designed to help models learn to reject queries that may trigger hallucinatory responses, while enhancing the stability of responses to properly answerable queries. The specific contents and creation procedures of the dataset are not elaborated in the associated paper, but its application in refining large language models (LLMs) is noted, particularly its efficacy in reducing hallucinatory outputs and improving response accuracy.




