Gandalf
收藏资源简介:
Gandalf数据集由Lakera公司创建,旨在为大语言模型(LLM)的安全防御提供多样化的自适应攻击数据。该数据集包含27.9万条提示攻击数据,通过众包红队平台Gandalf生成,涵盖了多种攻击类型,如越狱攻击、系统提示泄露和间接注入攻击等。数据集的创建过程通过游戏化的方式激励用户生成真实且多样化的攻击数据,并自动标记攻击的成功与否。该数据集的应用领域主要集中在LLM的安全防御研究,旨在帮助开发者设计既能有效防御攻击又不影响用户体验的防御策略。
The Gandalf dataset was developed by Lakera Inc. to provide diverse adaptive attack data for the security defense of Large Language Models (LLMs). This dataset contains 279,000 prompt attack samples, generated through the Gandalf crowdsourced red team platform, covering various attack types including jailbreak attacks, system prompt leakage, indirect prompt injection attacks, and more. The dataset was constructed via a gamified mechanism to incentivize users to produce authentic and diverse attack data, with automatic labeling of whether an attack was successful. Its primary application lies in LLM security defense research, aiming to help developers design defense strategies that can effectively defend against attacks while preserving user experience.

- 1Gandalf the Red: Adaptive Security for LLMsLakera · 2025年



