Gutenberg dataset
收藏资源简介:
本文涉及的数据集是Gutenberg数据集,用于评估大型语言模型(LLMs)上的成员推理攻击(MIAs)。该数据集由研究者用于实验,旨在通过消除已知偏差来创建“非偏见”和“不可分类”的数据集,以实现更公正的MIA评估。实验结果表明,即使消除了已知偏差,MIAs的评估仍然具有挑战性。该数据集的应用领域主要集中在LLMs的版权和伦理问题评估上,特别是用于检测模型训练数据中是否包含未经授权的受保护内容。
The dataset discussed in this paper is the Gutenberg Dataset, which is used to evaluate Membership Inference Attacks (MIAs) against Large Language Models (LLMs). This dataset is employed by researchers in experiments aiming to create "unbiased" and "unclassifiable" datasets by eliminating known biases, so as to enable more equitable evaluation of MIAs. Experimental results demonstrate that the evaluation of MIAs remains challenging even after known biases have been eliminated. The application scenarios of this dataset mainly focus on the evaluation of copyright and ethical issues related to LLMs, particularly for detecting whether unauthorized protected content is included in the model's training data.

- 1Nob-MIAs: Non-biased Membership Inference Attacks Assessment on Large Language Models with Ex-Post Dataset ConstructionUniversidad Carlos III de Madrid · 2024年



