Aegis-AI-Content-Safety-Dataset-2.0
收藏资源简介:
Aegis AI Content Safety Dataset 2.0 包含33,416条人类与LLM之间的注释交互,分为30,007条训练样本、1,445条验证样本和1,964条测试样本。该数据集是之前发布的Aegis 1.0内容安全数据集的扩展。数据集通过使用HuggingFace版本的人类偏好数据(来自Anthropic HH-RLHF)进行策划,仅提取提示,并从Mistral-7B-v0.1中引出响应。数据集遵循一个全面且可适应的安全风险分类法,分为12个顶级危险类别和9个细粒度子类别。数据集采用混合数据生成管道,结合了全对话级别的人类注释和多LLM“陪审团”系统来评估响应的安全性。
Aegis AI Content Safety Dataset 2.0 contains 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This dataset is an extension of the previously released Aegis 1.0 Content Safety Dataset. The dataset is curated using HuggingFace-hosted human preference data sourced from Anthropic HH-RLHF, where only the prompts are extracted, and responses are elicited from Mistral-7B-v0.1. It follows a comprehensive and adaptable safety risk taxonomy, categorized into 12 top-level hazardous categories and 9 fine-grained subcategories. The dataset adopts a hybrid data generation pipeline that combines full-conversation-level human annotations and a multi-LLM "jury" system to evaluate the safety of model responses.




