NotInject
收藏资源简介:
NotInject数据集由威斯康星大学麦迪逊分校的研究团队创建,旨在评估提示防护模型中的过度防御问题。该数据集包含339个精心设计的良性输入样本,这些样本中嵌入了常见于提示注入攻击中的触发词,以进行细粒度评估。数据集的创建过程包括触发词识别、精炼和语料生成,确保样本的语义一致性和无害性。NotInject数据集主要应用于检测和缓解提示注入攻击,旨在提高大型语言模型在面对恶意输入时的鲁棒性和准确性。
The NotInject dataset was created by a research team from the University of Wisconsin-Madison, aiming to assess the over-defense problem in prompt guard models. This dataset includes 339 meticulously designed benign input samples, which embed trigger words commonly found in prompt injection attacks to support fine-grained evaluation. The dataset's creation process covers trigger word identification, refinement, and corpus generation, ensuring the semantic consistency and harmlessness of the samples. The NotInject dataset is mainly applied to detect and mitigate prompt injection attacks, with the objective of improving the robustness and accuracy of large language models when facing malicious inputs.
InjecGuard 数据集概述
数据集名称
- NotInject
数据集描述
- NotInject 数据集旨在评估现有防护模型中的过度防御问题。数据集包含339个良性样本,这些样本中嵌入了常见于提示注入攻击中的触发词。
- 数据集分为三个子集,每个子集包含具有一个、两个或三个触发词的句子。
- 每个子集包含113个良性句子,涵盖四个主题:常见查询、技术查询、虚拟创建和多语言查询。
数据集结构
- 子集划分:
- 一个触发词的句子
- 两个触发词的句子
- 三个触发词的句子
- 主题分类:
- 常见查询
- 技术查询
- 虚拟创建
- 多语言查询
数据集用途
- 用于评估提示防护模型在处理包含触发词的良性输入时的性能,特别是检测过度防御问题。
数据集下载
- 数据集可通过 Hugging Face 下载。
相关资源
- 论文:arXiv
- 代码:GitHub
- 演示页面:InjecGuard 演示
引用
@articles{InjecGuard, title={InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models}, author={Hao Li and Xiaogeng Liu and Chaowei Xiao}, journal = {arXiv preprint arXiv:2410.22770}, year={2024} }

- 1InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models威斯康星大学麦迪逊分校 · 2024年



