SecureAgentBench
收藏资源简介:
SecureAgentBench是一个包含105个编码任务的数据集,旨在严格评估代码代理在安全代码生成方面的能力。每个任务都包括真实的任务设置,需要在大型的代码库中进行多文件编辑,基于真实世界的开源漏洞构建的上下文,以及功能测试、通过概念验证漏洞进行的漏洞检查和静态分析检测新引入漏洞的全面评估。该数据集旨在模拟软件开发过程中人类开发者引入漏洞的情境,并提供了真实且符合实际软件演变的评估场景。
SecureAgentBench is a dataset consisting of 105 coding tasks, designed to rigorously evaluate the capabilities of code agents in secure code generation. Each task includes realistic task settings, requires multi-file edits within large-scale codebases, uses contexts constructed from real-world open-source vulnerabilities, and features comprehensive evaluations covering functional tests, vulnerability checks via proof-of-concept (PoC) exploits, and static analysis for detecting newly introduced vulnerabilities. This dataset aims to simulate the scenario where human developers introduce vulnerabilities during software development, and provides realistic assessment scenarios that align with actual software evolution practices.
SecureAgentBench 数据集概述
基本信息
- 数据集名称:SecureAgentBench
- 托管地址:https://github.com/iCSawyer/SecureAgentBench
当前状态
- 开发状态:代码库准备中
- 可用性说明:暂未发布,请保持关注




