TermiBench
收藏资源简介:
TermiBench是一个针对真实世界、细粒度和面向代理的渗透测试评估基准。它包含510个主机,跨越25个不同的服务和30个CVE,涵盖了2015年至2025年间的漏洞。每个主机配置了多达七个无漏洞的服务和一个在2015年至2025年间的一个漏洞服务。TermiBench旨在更真实地模拟现实世界的渗透测试场景,要求代理在没有预先信息的情况下自主进行侦察,并能够区分可利用和不可利用的服务。该基准为评估代理在真实世界中的渗透测试性能提供了一个更准确的平台。
TermiBench is a real-world, fine-grained, agent-oriented penetration testing evaluation benchmark. It contains 510 hosts spanning 25 distinct services and 30 CVEs, covering vulnerabilities from 2015 to 2025. Each host is configured with up to seven non-vulnerable services and one vulnerable service dating within the 2015 to 2025 period. TermiBench is designed to more authentically simulate real-world penetration testing scenarios, requiring agents to conduct autonomous reconnaissance without prior information and distinguish between exploitable and non-exploitable services. This benchmark provides a more accurate platform for evaluating agents' penetration testing performance in real-world scenarios.
数据集概述:Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
基本信息
- 标题:Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
- 作者:Wuyuao Mai, Geng Hong, Qi Liu, Jinsong Chen, Jiarun Dai, Xudong Pan, Yuan Zhang, Min Yang
- 提交日期:2025年9月11日(v1),2025年9月15日修订(v2)
- arXiv标识符:arXiv:2509.09207v2
- DOI:https://doi.org/10.48550/arXiv.2509.09207
- 所属学科:Cryptography and Security (cs.CR)
摘要
渗透测试对于识别和缓解安全漏洞至关重要,但传统方法仍然昂贵、耗时且依赖专家人力。近期研究探索了AI驱动的渗透测试代理,但评估依赖于过度简化的夺旗(CTF)设置,这些设置嵌入了先验知识并降低了复杂性,导致性能估计远离实际实践。
本研究通过引入第一个真实世界、面向代理的渗透测试基准TermiBench来弥补这一差距,该基准将目标从“找旗”转变为实现全系统控制。基准涵盖25个服务和30个CVE的510个主机,具有需要自主侦察、区分良性和可利用服务以及稳健漏洞利用执行的现实环境。使用此基准,发现现有系统在现实条件下几乎无法获得系统shell。
为解决这些挑战,提出了TermiAgent,一个多代理渗透测试框架。TermiAgent通过定位记忆激活机制减轻长上下文遗忘,并通过结构化代码理解而非简单检索构建可靠的漏洞利用库。在评估中,该工作优于最先进的代理,表现出更强的渗透测试能力,减少执行时间和财务成本,并展示了即使在笔记本电脑规模部署上的实用性。该工作提供了第一个用于真实世界自主渗透测试的开源基准和一个新颖的代理框架,为AI驱动的渗透测试建立了里程碑。
相关资源
- 论文PDF:https://arxiv.org/pdf/2509.09207v2
- HTML版本:https://arxiv.org/html/2509.09207v2
- TeX源码:https://arxiv.org/src/2509.09207v2
- 其他格式:https://arxiv.org/format/2509.09207v2

- 1Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing复旦大学 · 2025年



