AI45Research/ATBench-Codex
收藏资源简介:
ATBench-Codex是一个面向Codex的代理轨迹安全基准数据集,源自ATBench,作为AI代理安全和安全诊断护栏框架AgentDoG的基准配套。它设计用于可执行编码代理设置中的轨迹级安全评估,重点关注在诸如shell执行、工作区变更、仓库变更、MCP工具调用或长时工具链等操作实际执行之前必须做出安全决策的关键点。与原始ATBench相比,此版本围绕Codex特定的操作语义构建,包括多工具编码工作流、结构化展开事件、仓库和工件操作、MCP服务器供应链表面以及仅在针对实时工作区执行操作时才可见的指令遵循失败。该数据集包含500个样本,每个样本包含用户对话、结构化的Codex执行轨迹、二进制安全判断和细粒度分类标签,适用于对话级安全检测和工具驱动代理轨迹的深入分析。
ATBench-Codex is a proxy trajectory safety benchmark dataset tailored for Codex, derived from ATBench, and serves as the supporting benchmark for AgentDoG, an AI agent safety and safety diagnostic guardrail framework. It is designed for trajectory-level safety assessment in executable coding agent scenarios, focusing on critical points where security decisions must be made prior to the actual execution of operations such as shell execution, workspace modification, repository modification, MCP tool invocation, or long-duration toolchains. Compared with the original ATBench, this version is built around Codex-specific operational semantics, including multi-tool coding workflows, structured unfolding events, repository and artifact operations, MCP server supply chain surfaces, and instruction-following failures that are only observable when executing operations against live workspaces. This dataset contains 500 samples, each including user dialogues, structured Codex execution trajectories, binary safety judgments, and fine-grained classification labels, which is suitable for dialogue-level security detection and in-depth analysis of tool-driven agent trajectories.



