AutoAdvExBench
收藏资源简介:
AutoAdvExBench是一个用于评估大型语言模型是否能自主利用对抗样本防御的基准。该数据集包含了51个真实世界的对抗样本防御实现,由ETH Zurich的研究人员创建。数据集涵盖了从arXiv抓取的论文中筛选出的与对抗机器学习相关的防御方法,并通过手动筛选确保了防御方法的多样性和可复现性。数据集旨在解决机器学习安全领域中的实际问题,为评估AI模型在对抗机器学习研究中的应用提供了一种新的、直接的度量方式。
AutoAdvExBench is a benchmark for evaluating whether large language models can autonomously utilize adversarial sample defenses. This dataset contains 51 real-world adversarial sample defense implementations created by researchers from ETH Zurich. It covers adversarial machine learning-related defense methods selected from papers scraped from arXiv, and ensures the diversity and reproducibility of these defense methods through manual screening. This benchmark aims to address practical issues in the field of machine learning security, and provides a novel and direct metric for evaluating the application of AI models in adversarial machine learning research.




