MarkrAI/ko-jailbreak
收藏资源简介:
ko-jailbreak是一个韩国语jailbreak(越狱)评估基准数据集,它整合了多个公开的英语jailbreak基准,并将其翻译和精炼成韩语,同时去除了源之间的重复项,形成了一个统一的评估工具。该数据集旨在测量韩国语大型语言模型(LLM)和安全防护措施对实际jailbreak攻击的鲁棒性。它包含两个主要部分:behavior(模型应拒绝的有害请求本身,共10,303个)和template(用于包装有害请求以绕过安全防护的包装器,如DAN等,共1,138个),总规模为11,441个条目。数据集基于允许重新分发的许可证(如MIT/Apache-2.0)从多个来源(如SALAD-Bench、HarmBench等)收集,并经过翻译、质量过滤和去重处理。它主要用于红队测试和安全评估,以帮助提高LLM的安全性。
ko-jailbreak is a Korean jailbreak evaluation benchmark that integrates multiple public English jailbreak benchmarks, translating and refining them into Korean while removing duplicates across sources to form a unified evaluation tool. It is designed to measure the robustness of Korean large language models (LLMs) and safety guards against actual jailbreak attacks. The dataset consists of two axes: behavior (the harmful requests that models should reject, totaling 10,303 entries) and template (wrappers like DAN that bypass safety guards by enclosing harmful requests, totaling 1,138 entries), with an overall size of 11,441 entries. It is collected from sources with redistribution-permissive licenses (e.g., MIT/Apache-2.0), such as SALAD-Bench and HarmBench, and undergoes translation, quality filtering, and deduplication. It is intended for red-teaming and safety evaluation purposes to enhance LLM security.




