amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
收藏资源简介:
该数据集名为“对抗竞技场:可信AI挑战数据集”,包含通过对抗竞技场框架生成的多轮对抗对话。该框架是一个交互式竞赛,攻击者机器人试图从防御者机器人处诱导出不安全的代码或网络攻击辅助。数据集收集自亚马逊Nova AI挑战赛——可信AI项目,专注于大型语言模型(LLM)的网络安全对齐。数据集包含25,600个会话,总计250,906个对话轮次,平均每个会话9.8轮,最大轮次为10轮。数据生成涉及10个学术团队(5个攻击者、5个防御者)之间的多轮锦标赛,每个匹配包括自动多轮对话:攻击者试图诱导出易受攻击的代码、恶意代码或网络攻击辅助;防御者旨在提供有用的编码响应,同时拒绝不安全请求。对话限制在最多10轮。数据集经过两步评估:使用静态分析检测易受攻击代码,并由3名网络安全专家人工标注安全事件。数据集结构包括会话元数据、代码与漏洞信息、标注信息和对话内容。该数据集用于安全研究目的,包含可能涉及恶意代码的对抗性提示,需负责任使用。
This dataset, named Adversarial Arena: Trusted AI Challenge Dataset, contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. It was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs. The dataset includes 25,600 sessions with a total of 250,906 conversation turns, averaging 9.8 turns per session and a maximum of 10 turns. Data generation involved tournaments among 10 academic teams (5 attackers, 5 defenders), with automated multi-turn conversations: attackers tried to elicit vulnerable code, malicious code, or cyberattack assistance; defenders aimed to provide helpful coding responses while refusing unsafe requests. Conversations were limited to a maximum of 10 turns. Evaluation was conducted through a two-step process: vulnerable code detection using static analysis and security event detection by a panel of 3 expert human annotators. The dataset structure comprises session metadata, code and vulnerability information, annotation information, and conversation content. It is intended for safety research purposes and contains adversarial prompts that may reference malicious code, requiring responsible use.




