frames-benchmark
收藏资源简介:
FRAMES数据集是一个综合评估数据集,旨在测试检索增强生成(RAG)系统在事实性、检索准确性和推理方面的能力。该数据集包含824个具有挑战性的多跳问题,这些问题需要从2到15篇维基百科文章中获取信息。问题涵盖了历史、体育、科学、动物、健康等多个主题,并且每个问题都标有推理类型,如数值、表格、多重约束、时间性和后处理。数据集还提供了每个问题的黄金答案和相关的维基百科文章。FRAMES数据集的主要特点包括测试端到端的RAG能力、需要整合来自多个来源的信息、包含复杂的推理和时间性消歧,并设计为对最先进的语言模型具有挑战性。该数据集可用于评估RAG系统性能、基准测试语言模型的事实性和推理能力,以及开发和测试多跳检索策略。
The FRAMES dataset is a comprehensive evaluation dataset designed to test the capabilities of retrieval-augmented generation (RAG) systems in terms of factuality, retrieval accuracy, and reasoning. This dataset includes 824 challenging multi-hop questions that require information retrieved from 2 to 15 Wikipedia articles. The questions cover multiple topics such as history, sports, science, animals, and health, and each question is annotated with reasoning types including numerical, tabular, multi-constraint, temporal, and post-processing. The dataset also provides the gold standard answers and the associated Wikipedia articles for every question. Key features of the FRAMES dataset are as follows: it evaluates end-to-end RAG capabilities, requires integrating information from multiple sources, contains complex reasoning and temporal disambiguation, and is designed to be challenging for state-of-the-art language models. This dataset can be used to evaluate RAG system performance, benchmark the factuality and reasoning abilities of language models, as well as develop and test multi-hop retrieval strategies.




