RAGTruth
收藏资源简介:
RAGTruth是一个专为分析大型语言模型(LLM)在标准RAG框架应用中各个领域和任务的词级幻觉而设计的数据集。该数据集由NewsBreak和伊利诺伊大学厄巴纳-香槟分校创建,包含近18,000个来自不同LLM的自然生成响应,这些响应经过细致的人工标注,涵盖了幻觉强度的评估。RAGTruth不仅用于基准测试不同LLM的幻觉频率,还用于评估现有幻觉检测方法的有效性。该数据集主要应用于开发和评估在RAG设置下防止幻觉的策略,旨在提高LLM在实际应用中的可靠性和准确性。
RAGTruth is a dataset specifically designed for analyzing token-level hallucinations of Large Language Models (LLMs) across diverse domains and tasks in standard Retrieval-Augmented Generation (RAG) framework applications. Developed by NewsBreak and the University of Illinois Urbana-Champaign, this dataset comprises nearly 18,000 naturally generated responses from various LLMs, which have undergone meticulous manual annotation covering evaluations of hallucination intensity. RAGTruth can be used not only to benchmark the hallucination frequency of different LLMs but also to assess the effectiveness of existing hallucination detection methods. This dataset is primarily applied to develop and evaluate hallucination prevention strategies under RAG settings, aiming to improve the reliability and accuracy of LLMs in real-world applications.




