AEGIS2.0
收藏资源简介:
AEGIS2.0 是由英伟达团队创建的一个多样化的大语言模型(LLM)安全数据集,旨在解决LLM在商业应用中的内容安全问题。该数据集包含34,248条人类与LLM交互的样本,涵盖了12个核心风险类别和9个细粒度风险类别。数据集的生成过程结合了人工标注和多LLM“陪审团”系统,确保了数据的多样性和质量。AEGIS2.0 的数据来源包括真实世界交互中的有害提示和LLM生成的响应,特别关注对抗性攻击、文化背景和关键风险。数据集的应用领域主要集中在LLM的安全防护,帮助模型更好地识别和处理新兴的安全风险,确保其在商业应用中的安全性和可靠性。
AEGIS2.0 is a diverse large language model (LLM) safety dataset developed by the NVIDIA team, which aims to resolve content security challenges faced by LLMs in commercial deployments. This dataset includes 34,248 human-LLM interaction samples, spanning 12 core risk categories and 9 fine-grained risk categories. The construction of AEGIS2.0 integrates manual annotation and a multi-LLM jury system, guaranteeing the diversity and high quality of the dataset. The data sources of AEGIS2.0 cover harmful prompts extracted from real-world human-LLM interactions and LLM-generated responses, with a special emphasis on adversarial attacks, cultural contexts and critical risks. The main application scope of this dataset is LLM security protection, assisting models in better identifying and addressing emerging security risks, thereby ensuring the safety and reliability of LLMs in commercial applications.




