McGill-NLP/tukabench
收藏资源简介:
TukaBench是一个用于评估大型语言模型在非洲语言中是否会被越狱的安全基准数据集。它旨在测试模型在阿姆哈拉语、豪萨语、伊博语、奇切瓦语、斯瓦希里语、科萨语或约鲁巴语等语言中,是否会被诱导生成有害内容(如欺诈指南、选举虚假信息或暴力指令)。数据集包含三个部分:afri-jbb-harm(JailbreakBench的有害提示,经人工忠实翻译,保留西方背景)、afri-jbb-benign(JailbreakBench的良性控制提示,用于测量过度拒绝)和afri-jbb-culture(相同的有害提示,但在翻译前被重写为非洲文化背景,包含本地相关的命名实体和场景)。每种语言共有300个提示(每个部分100个),支持8种语言。数据集适用于多语言安全对齐、攻击鲁棒性和低资源环境下LLM-as-a-judge可靠性的研究。
TukaBench is a safety benchmark dataset for evaluating whether large language models (LLMs) can be jailbroken when interacting in African languages. It is designed to test if models can be induced to generate harmful content—including fraud guidance, election disinformation, and violent instructions—across languages such as Amharic, Hausa, Igbo, Chichewa, Swahili, Xhosa, and Yoruba. The dataset consists of three core components: afri-jbb-harm (harmful prompts sourced from JailbreakBench, manually and faithfully translated while preserving the original Western contextual background), afri-jbb-benign (benign control prompts from JailbreakBench, utilized to measure model over-rejection), and afri-jbb-culture (identical harmful prompts rewritten with African cultural contexts before translation, integrating locally relevant named entities and scenarios). Each supported language contains 300 prompts, with 100 prompts allocated to each of the three components, and the dataset supports a total of 8 languages. This dataset is applicable to research on multilingual safety alignment, adversarial attack robustness, and the reliability of LLM-as-a-judge frameworks in low-resource language environments.




