长格式问答数据集
收藏资源简介:
该数据集旨在评估FRANQ和其他UQ技术在RAG上的表现。数据集包含76个问题,每个问题都有相应的答案,其中包含从Llama 3B和Falcon 3B模型输出中提取的1,782个断言。数据集通过自动标注和手动验证相结合的方式创建,并标注了断言的真实性和忠实性。该数据集可用于训练和测试UQ方法,以评估长格式问答中RAG生成的回答的真实性。
This dataset is intended to evaluate the performance of FRANQ and other Uncertainty Quantification (UQ) techniques when applied to Retrieval-Augmented Generation (RAG) systems. It comprises 76 questions, each paired with a corresponding answer, and a total of 1,782 assertions extracted from the outputs of Llama 3B and Falcon 3B models. The dataset was constructed through a combined workflow of automatic annotation and manual verification, with the factuality and faithfulness of each assertion annotated. This dataset can be utilized for training and testing UQ methods to assess the factuality of responses generated by RAG systems in long-form question answering scenarios.

- 1Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation瑞士苏黎世联邦理工学院, 阿布扎比MBZUAI · 2025年



