AI-Secure/DTap-Bench-Agent-Trajectories
收藏资源简介:
DecodingTrust-Agent Platform(DTAP-BENCH)是一个用于AI代理的可控交互式红队测试平台。该数据集完整收集了从评估DTap-Bench产生的代理轨迹,涵盖14个现实世界领域和50多个模拟环境,这些环境复制了广泛使用的系统,如Google Workspace、PayPal、Slack、Salesforce、Snowflake和Databricks。每个任务都提供了评估者所需配置,以启动沙箱、运行代理并验证结果——包括`config.yaml`(任务规范+MCP服务器绑定)、`setup.sh`(每个任务的环境初始化)和`judge.py`(可验证的结果检查)。数据集包含良性任务和恶意任务,恶意任务进一步按威胁模型(直接与间接)和风险类别(如数据外泄、危险操作、无效同意、错误信息幻觉、欺诈冒充、通用AI限制、拒绝用户请求、操纵使用等)划分。
DecodingTrust-Agent Platform (DTAP-BENCH) is a controllable interactive red teaming test platform for AI Agents. This dataset comprehensively collects agent trajectories generated from evaluations performed on DTAP-BENCH, covering 14 real-world domains and over 50 simulated environments that replicate widely used systems such as Google Workspace, PayPal, Slack, Salesforce, Snowflake, and Databricks. Each task provides all necessary configurations for evaluators to launch sandboxes, run agents and validate results, including `config.yaml` (task specification + MCP server binding), `setup.sh` (environment initialization for each individual task), and `judge.py` (verifiable result checking). The dataset contains both benign and malicious tasks. Malicious tasks are further categorized by threat models (direct vs. indirect) and risk categories including data exfiltration, dangerous operations, invalid consent, misinformation hallucination, fraudulent impersonation, general AI limitations, refusing user requests, and manipulative usage, among others.




