AI-Secure/DecodingTrust-Agent-Platform
收藏资源简介:
DecodingTrust-Agent Platform (DTAP) 是一个可控且交互式的AI代理红队测试平台数据集,涵盖14个真实世界领域(如浏览器、代码、CRM、客户服务、金融、法律、macOS、医疗、操作系统文件系统、研究、电信、旅行、Windows和工作流)和50多个模拟环境,复现了广泛使用的系统如Google Workspace、PayPal、Slack、Salesforce、Snowflake和Databricks。该数据集包含2,806个良性任务和3,876个恶意任务,恶意任务进一步按威胁模型(直接与间接)和风险类别(如数据窃取、危险操作、无效同意、错误信息幻觉、欺诈冒充、通用AI限制、拒绝用户请求、操纵性使用等)划分。每个任务提供配置、环境设置和结果验证文件,用于评估AI代理的安全性和性能。
DecodingTrust-Agent Platform (DTAP) is a controllable and interactive red-teaming platform for AI agents, spanning 14 real-world domains (e.g., browser, code, CRM, customer-service, finance, legal, macOS, medical, os-filesystem, research, telecom, travel, windows, workflow) and 50+ simulation environments that replicate widely-used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. The dataset includes 2,806 benign tasks and 3,876 malicious tasks, with malicious tasks further partitioned by threat model (direct vs. indirect) and risk category (e.g., data-exfiltration, dangerous-actions, invalid-consent, misinformation-hallucination, fraud-impersonation, general-ai-restrictions, deny-user-requests, manipulative-use). Each task ships configuration, environment seeding, and outcome verification files for evaluating AI agent safety and performance.




