AIPot Paper 1 Dataset: 400-Run Empirical Study of LLM Penetration Testing Consistency
收藏官方服务:
资源简介:
This dataset accompanies the paper "How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency" (Erdem, 2026). It contains 400 raw run logs (100 each from Claude Sonnet 4, Gemini 2.5 Flash-Lite, GPT-4o-mini, and qwen2.5-coder:14b) collected against a fixed multi-service honeypot, the orchestrator and analysis code, the Terraform configuration for the target infrastructure, and the complete statistical report. See README.md for layout and reproduction instructions.
提供机构:
Zenodo创建时间:
2026-05-28



