GuidedBench
收藏资源简介:
GuidedBench是由香港科技大学创建的一个评估基准,旨在通过提供一套包含特定恶意问题的数据集,以及详细的评估指南,对越狱方法进行更标准化和公平的评估。该数据集分为核心集和附加集,核心集包含180个所有受害语言模型都会拒绝的问题,附加集包含20个因特定安全政策而只有部分语言模型会拒绝的问题。这些问题设计为直接、简洁的文本指令,而非情景混合案例,以避免耦合。此外,针对每个有害问题案例,都编写了详细的评分指南,关注攻击者实现有害目标所需的关键实体和功能。
GuidedBench is an evaluation benchmark developed by The Hong Kong University of Science and Technology. It aims to enable more standardized and fair evaluations of jailbreaking methods by providing a dataset containing specific malicious prompts and detailed evaluation guidelines. This benchmark is split into a core set and an additional set: the core set includes 180 prompts that all victim language models will refuse to respond to, while the additional set contains 20 prompts that only a portion of language models will decline to address due to specific security policies. These prompts are designed as direct and concise text instructions rather than scenario-mixed cases to avoid confounding variables. Furthermore, detailed scoring guidelines have been formulated for each harmful prompt case, focusing on the key entities and functions required for an attacker to accomplish their malicious objectives.

- 1GuidedBench: Equipping Jailbreak Evaluation with Guidelines香港科技大学 · 2025年



