Gandalf Ignore Instruction
收藏arXiv2025-09-30 收录
官方服务:
资源简介:
该数据集包含了在教育游戏中收集的提示,该游戏旨在告知人们关于大型语言模型(LLMs)在提示攻击下可能出现的AI泄露风险。提示内容通过角色扮演的方式揭示游戏中的秘密密码。该数据集的规模为1000条提示,其任务是绕过模型的对齐防御机制。
This dataset contains prompts collected from an educational game developed to inform the public about potential AI leakage risks faced by Large Language Models (LLMs) during prompt attacks. These prompts are used to reveal the secret in-game passwords through role-playing scenarios. Comprising a total of 1000 prompts, the dataset targets the task of bypassing the alignment defense mechanisms of the models.
提供机构:
Lakera AI搜集汇总
背景与挑战
背景概述
该数据集是一个包含1000条提示的教育游戏数据集,旨在通过角色扮演方式模拟提示攻击,揭示大型语言模型在面临攻击时可能出现的AI泄露风险,并帮助用户理解如何绕过模型的对齐防御机制。
以上内容由遇见数据集搜集并总结生成



