MuhammadAnas1657/llmail-inject-challenge
收藏资源简介:
该数据集包含大量攻击提示,这些提示是已关闭的LLMail-Inject: Adaptive Prompt Injection Challenge挑战中收集的。挑战旨在模拟LLM集成电子邮件客户端(LLMail服务)中的提示注入防御规避,攻击者需通过精心设计的电子邮件绕过防御,使LLM执行未授权的操作(如发送电子邮件)。数据集涵盖两个阶段(Phase1和Phase2),包括不同场景(如无检索的电子邮件汇总、基于检索的问答和数据外泄)、多种防御机制(如Spotlighting、PromptShield、LLM-as-a-judge、TaskTracker及其组合)和不同LLM模型(如Phi-3-medium-128k-instruct和GPT-4o mini)。数据格式包括原始提交记录、带标签的唯一提交、元数据(如场景描述、系统提示)和测试用电子邮件,用于研究自适应提示注入攻击的检测和防御。
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge. The challenge simulates prompt injection defense evasion in an LLM-integrated email client (the LLMail service), where attackers craft emails to bypass defenses and cause the LLM to perform unauthorized actions (e.g., sending emails). The dataset covers two phases (Phase1 and Phase2) with various scenarios (e.g., email summarization without retrieval, retrieval-based Q&A, and data exfiltration), multiple defense mechanisms (e.g., Spotlighting, PromptShield, LLM-as-a-judge, TaskTracker, and their combination), and different LLM models (e.g., Phi-3-medium-128k-instruct and GPT-4o mini). The data format includes raw submissions, labeled unique submissions, metadata (e.g., scenario descriptions, system prompts), and emails for false positive testing, aimed at studying adaptive prompt injection attacks for detection and defense research.




