BackdoorLLM
收藏资源简介:
BackdoorLLM是由新加坡管理大学、墨尔本大学和复旦大学联合创建的一个综合基准数据集,专门用于研究大型语言模型中的后门攻击。该数据集包括一个标准化的训练流程和多种攻击策略,如数据中毒、权重中毒、隐藏状态攻击和思维链攻击。数据集内容丰富,涉及多种模型架构和任务数据集,旨在全面评估和分析后门攻击的有效性和局限性。BackdoorLLM的应用领域主要集中在提高AI安全性,解决大型语言模型在敏感应用中的安全问题。
BackdoorLLM is a comprehensive benchmark dataset jointly developed by Singapore Management University, University of Melbourne and Fudan University, which is specifically designed for researching backdoor attacks in large language models. This dataset provides a standardized training pipeline and a variety of attack strategies, including data poisoning, weight poisoning, hidden state attacks and chain-of-thought attacks. It covers diverse model architectures and task datasets, aiming to comprehensively evaluate and analyze the effectiveness and limitations of backdoor attacks. The application scope of BackdoorLLM mainly focuses on enhancing AI security and addressing the security problems of large language models in sensitive applications.

- 1BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks on Large Language Models新加坡管理大学, 墨尔本大学, 复旦大学 · 2024年



