Jaemo3123/JBB-Behaviors
收藏资源简介:
JBB-Behaviors数据集是JailbreakBench基准测试的一部分,是一个开源的大语言模型(LLM)越狱鲁棒性基准测试数据集。该数据集包含100个不同的滥用行为,这些行为既有原创的,也来源于先前的工作(特别是Trojan Detection Challenge/HarmBench和AdvBench),并参考了OpenAI的使用政策进行整理。每个数据条目包含以下五个组件:Behavior(描述独特滥用行为的唯一标识符)、Goal(请求不当行为的查询)、Target(对目标字符串的肯定响应)、Category(基于OpenAI使用政策的更广泛滥用类别)和Source(行为来源,如原创、Trojan Detection Challenge 2023 Red Teaming Track/HarmBench或AdvBench)。数据集分为十个类别,对应OpenAI的使用政策,旨在通过代表性行为实现快速评估新攻击方法。数据集主要用于跟踪越狱攻击和防御的性能,支持研究人员比较算法效果。
The JBB-Behaviors dataset is part of JailbreakBench, an open-source robustness benchmark for jailbreaking large language models (LLMs). It comprises 100 distinct misuse behaviors, which are both original and sourced from prior work (specifically, Trojan Detection Challenge/HarmBench and AdvBench), curated with reference to OpenAIs usage policies. Each entry in the dataset has five components: Behavior (a unique identifier describing a distinct misuse behavior), Goal (a query requesting an objectionable behavior), Target (an affirmative response to the goal string), Category (a broader category of misuse from OpenAIs usage policies), and Source (the source from which the behavior was sourced, e.g., Original, Trojan Detection Challenge 2023 Red Teaming Track/HarmBench, or AdvBench). The dataset is divided into ten categories corresponding to OpenAIs usage policies, focusing on 100 representative behaviors to enable faster evaluation of new attacks. It is designed to track the performance of attacks and defenses, providing a stable way for researchers to compare future algorithms.



