EVMbench
收藏资源简介:
EVMbench是由OpenAI、Paradigm和OtterSec联合发布的首个智能合约安全领域的大规模AI代理基准测试数据集,包含来自40个Code4rena审计仓库的120个精选漏洞样本。该数据集通过隔离的Docker环境评估代理在漏洞检测、修复和利用三方面的能力,其核心价值在于为自动化AI审计提供标准化测试框架。数据来源主要为2025年8月前的历史审计竞赛报告,可能存在模型训练数据污染风险。研究团队额外构建了包含22个2026年2月后真实安全事件的纯净子集,以验证模型在真实场景中的泛化能力。该数据集主要应用于区块链安全领域,旨在评估AI代理在智能合约漏洞挖掘方面的有效性,推动自动化审计技术的发展。
EVMbench is the first large-scale AI Agent benchmark dataset in the field of smart contract security, jointly released by OpenAI, Paradigm and OtterSec. It contains 120 curated vulnerability samples from 40 Code4rena audit repositories. This dataset evaluates AI Agents' capabilities across three core aspects: vulnerability detection, repair and exploitation, via isolated Docker environments. Its core value lies in providing a standardized testing framework for automated AI-powered auditing. The dataset is primarily sourced from historical audit contest reports prior to August 2025, which may carry the risk of training data contamination for AI models. The research team additionally constructed a clean subset containing 22 real-world security incidents that occurred after February 2026, to validate the generalization ability of models in real-world scenarios. This dataset is mainly applied in the field of blockchain security, aiming to evaluate the effectiveness of AI Agents in smart contract vulnerability mining and promote the development of automated auditing technologies.



