Awesome-Jailbreak-on-LLMs
收藏资源简介:
该合集是一个关于大语言模型越狱方法的资源集合,涵盖越狱攻击、防御、评估和分析等领域。它收集了包括数据集、论文、代码和评估工具在内的多种资源,主题聚焦于大语言模型的安全性和越狱技术。
This collection is a curated resource set dedicated to jailbreak methods for large language models (LLMs), spanning domains including jailbreak attacks, defensive strategies, evaluation methodologies, and related analyses. It gathers a wide range of resources such as datasets, academic papers, code, and evaluation tools, with its core focus on the security of LLMs and jailbreak-related technologies.
数据集概述
项目名称:Awesome-Jailbreak-on-LLMs
项目类型:文献与资源合集
核心主题:大语言模型(LLM)的越狱攻击与防御方法
主要内容:该合集收录了针对大语言模型的最新、最前沿的越狱方法,涵盖论文、代码、数据集、评估与分析。
分类目录
-
Jailbreak Attack(越狱攻击)
- Attack on LRMs(对推理模型攻击)
- Black-box Attack(黑盒攻击)
- White-box Attack(白盒攻击)
- Multi-turn Attack(多轮攻击)
- Attack on RAG-based LLM(对基于检索增强生成模型的攻击)
- Multi-modal Attack(多模态攻击)
-
Jailbreak Defense(越狱防御)
- Learning-based Defense(基于学习的防御)
- Strategy-based Defense(基于策略的防御)
- Guard Model(守卫模型)
- Moderation API(审核API)
-
Evaluation & Analysis(评估与分析)
-
Application(应用)
收录论文示例
-
Attack on LRMs:
- BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models(AAAI26)
- ExtendAttack: Attacking Servers of LRMs via Extending Reasoning(AAAI26)
- H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models(arXiv, 2025)
-
Black-box Attack:
- FlipAttack: Jailbreak LLMs via Flipping(ICML25)
- Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection(ICML25)
- Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy(CVPR25)
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs(ACL24)
- Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses(NeurIPS24)
引用论文
该仓库推荐引用以下五篇相关论文:
- GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video(arXiv, 2026)
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning(arXiv, 2025)
- GuardReasoner: Towards Reasoning-based LLM Safeguards(arXiv, 2025)
- FlipAttack: Jailbreak LLMs via Flipping(arXiv, 2024)
- Safety in Large Reasoning Models: A Survey(arXiv, 2025)
获取方式
- 数据集详情页:https://github.com/yueliu1999/Awesome-Jailbreak-on-LLMs
- 论文链接可在仓库中各表格中获取(均为arXiv或会议官方地址)
- 代码链接可在仓库中各表格中获取(部分为GitHub或Hugging Face链接)
联系方式
如有问题,可联系 yliu@u.nus.edu。




