JailbreakBench
收藏资源简介:
JailbreakBench是一个开放的鲁棒性基准,用于评估大型语言模型(LLMs)的越狱攻击。该数据集包含100种行为,旨在与OpenAI的使用政策保持一致。数据集的创建过程涉及收集和标准化先进的对抗性提示,以及开发一个评估框架,该框架包括明确的威胁模型、系统提示、聊天模板和评分函数。JailbreakBench的应用领域主要集中在提高LLMs对越狱攻击的抵抗力,确保模型在安全关键领域的部署更为安全。
JailbreakBench is an open-access robustness benchmark for evaluating jailbreak attacks against large language models (LLMs). This dataset contains 100 behaviors that are aligned with OpenAI's usage policies. The creation of the dataset involves collecting and standardizing state-of-the-art adversarial prompts, as well as developing an evaluation framework that includes a clear threat model, system prompts, chat templates, and scoring functions. The primary use cases of JailbreakBench center on enhancing the resilience of LLMs against jailbreak attacks, and ensuring safer deployment of the models in safety-critical domains.

- 1JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models宾夕法尼亚大学 · 2024年



