CentificAIResearch/email-classifier-adversarial-security
收藏资源简介:
该基准测试评估对抗条件下电子邮件分类系统的安全性和鲁棒性,旨在评估模型正确分类安全相关电子邮件(如钓鱼邮件、垃圾邮件和良性邮件)的能力,以及评估这些分类的基于大语言模型(LLM)的分级器的可靠性。基准测试包含两个互补的数据集:1) 电子邮件分类数据集,评估模型将电子邮件分类为预定义类别(如钓鱼、垃圾邮件和良性邮件)的能力,涵盖多个行业和攻击目标的现实对抗性电子邮件场景;2) 分级器评估数据集,评估基于LLM的分级器在评估电子邮件分类输出正确性时的性能,测量其在审查分类器预测和推理时的分级一致性、准确性和鲁棒性。这些数据集共同提供了一个框架,用于在安全聚焦的环境中基准测试电子邮件分类性能和分级可靠性。
This benchmark evaluates the safety and robustness of email classification systems under adversarial conditions. It is designed to assess both the ability of models to correctly classify security-related emails (such as phishing, spam, and benign emails) and the reliability of LLM-based graders that evaluate those classifications. The benchmark consists of two complementary datasets: 1) the Email Classification Dataset, which evaluates whether a model can correctly classify emails into predefined categories (e.g., phishing, spam, benign) and includes realistic adversarial email scenarios spanning multiple industries and attack objectives; and 2) the Grader Evaluation Dataset, which evaluates the performance of LLM-based graders in assessing the correctness of email classification outputs, measuring grading consistency, accuracy, and robustness when reviewing classifier predictions and reasoning. Together, these datasets provide a framework for benchmarking both email classification performance and grading reliability in security-focused environments.




