FARAD
收藏资源简介:
FARAD数据集由苏黎世联邦理工学院的研究团队创建,专门用于评估RAG-DI(RAG数据集推理)方法。该数据集包含3591个组,每组包含4篇独立撰写的虚构文章,这些文章共享一个主题和大量信息,但由不同的LLM作者撰写。数据集的创建过程包括从RepLiQA数据源中提取信息,并通过GPT4O等先进的LLM生成文章。FARAD数据集旨在模拟现实世界中RAG系统的数据冗余情况,主要应用于检测RAG系统中未经授权的数据使用问题。
The FARAD dataset was developed by a research team at ETH Zurich, specifically for evaluating the RAG-DI (RAG Dataset Inference) method. This dataset comprises 3591 groups, each containing four independently written fictional articles that share a unified theme and substantial overlapping information, but are authored by distinct LLMs. The creation pipeline of the FARAD dataset includes extracting information from the RepLiQA data source and generating articles via advanced LLMs such as GPT-4o. The FARAD dataset is intended to simulate data redundancy scenarios in real-world RAG systems, and is primarily employed to detect unauthorized data usage issues within RAG systems.




