HoH
收藏资源简介:
HoH是第一个大规模基准测试,旨在评估检索增强生成(RAG)在过时信息影响下的鲁棒性。该基准测试利用令牌级差异算法结合LLM管道,高效创建了一个大规模QA数据集,准确捕捉了现实世界事实中的时间知识演变。
HoH is the first large-scale benchmark designed to evaluate the robustness of Retrieval-Augmented Generation (RAG) against outdated information. By combining token-level difference algorithms with LLM pipelines, this benchmark efficiently constructs a large-scale QA dataset that accurately captures the temporal evolution of factual knowledge within real-world facts.
HoH 数据集概述
基本信息
- 数据集名称: HoH (How Outdated Information Harms Retrieval-Augmented Generation)
- 论文标题: HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation
- 会议/年份: ACL 2025
- 论文链接: https://arxiv.org/abs/2503.04800
- 数据集链接: https://huggingface.co/datasets/russwest404/HoH-QAs
研究背景
- 检索增强生成(RAG)是解决大语言模型(LLM)知识过时问题的有效方法,但其面临知识库中过时信息的关键挑战。
数据集目标
- 评估RAG在过时信息影响下的鲁棒性。
- 揭示过时信息如何显著降低RAG性能(如降低回答准确性并可能导致有害输出)。
数据集特点
- 规模: 首个大规模基准测试。
- 构建方法: 结合token-level diff算法和LLM pipeline,高效创建大规模QA数据集。
- 核心特征: 准确捕捉现实世界事实中的时间知识演变。
引用格式
bibtex @misc{ouyang2025hohdynamicbenchmarkevaluating, title={HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation}, author={Jie Ouyang and Tingyue Pan and Mingyue Cheng and Ruiran Yan and Yucong Luo and Jiaying Lin and Qi Liu}, year={2025}, eprint={2503.04800}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2503.04800}, }




