JailFlipBench
收藏资源简介:
我们提出的JailFlipBench可以分为三种场景:单模态、多模态和事实扩展。完整的多模态子集和其他子集的实例包含在`data`文件夹和[huggingface](https://huggingface.co/datasets/JailFlip/JailFlipBench)中。JailFlipBench的完整版本将在我们的论文被接受后发布。
The JailFlipBench proposed by us can be divided into three scenarios: single-modal, multi-modal, and fact extension. Complete instances of the multi-modal subset and other subsets are included in the 'data' folder and [huggingface](https://huggingface.co/datasets/JailFlip/JailFlipBench). The full version of JailFlipBench will be released after our paper is accepted.
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
数据集概述
- 名称: JailFlipBench
- 类型: 多模态与单模态数据集
- 场景分类: 单模态、多模态和事实扩展
- 存储位置:
- GitHub仓库的
data文件夹 - Hugging Face平台: https://huggingface.co/datasets/JailFlip/JailFlipBench
- GitHub仓库的
数据集内容
- 多模态子集: 完整版本已包含
- 其他子集: 实例化版本已提供
- 完整版本: 待论文接受后发布
相关资源
- 论文: Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- 项目网页: https://jailflip.github.io/
实验方法
- 攻击类型:
- 直接查询(Direct Query)
- 直接攻击(Direct Attack)
- 提示攻击(Prompting Attack)
- 高级攻击:
- LLM作为攻击者(llm-as-an-attacker)
- 对抗性后缀攻击(adversarial suffix attack)
引用
bibtex @article{zhou2025beyond, title={Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures}, author={Zhou, Yukai and Yang, Sibei and Wang, Wenjie}, journal={arXiv preprint arXiv:2506.07402}, year={2025} }
许可
- 许可证: MIT
- 许可证链接: https://opensource.org/licenses/MIT




