PSStrikes
收藏资源简介:
PSStrikes是由那不勒斯费德里科二世大学等研究机构精心构建的、首个面向真实世界PowerShell恶意软件评估的标注数据集,旨在支持生成式人工智能在网络安全威胁中的实证研究。该数据集包含817个经过严格去混淆处理的恶意脚本样本,每个样本均配有手工撰写的自然语言描述,总数据量约250行代码以内,主要来源于GitHub公开仓库、恶意软件分析平台(如Malware Bazaar)以及渗透测试框架(如Nishang)。数据集的创建过程经历了多阶段筛选、自动化与人工去混淆以及语义标注,确保了样本的可执行性与语义清晰度。该数据集的核心应用在于训练和评估大型语言模型生成恶意代码的能力,助力安全分析师深入理解AI驱动的网络攻击模式,并为入侵检测与威胁狩猎提供基准测试支持。
PSStrikes is the first annotated dataset for real-world PowerShell malware evaluation, meticulously constructed by research institutions including the University of Naples Federico II, and designed to support empirical research on generative AI in cybersecurity threats. This dataset contains 817 rigorously deobfuscated malicious script samples, each paired with a handcrafted natural language description, with a total code volume of approximately under 250 lines. The samples are mainly sourced from public GitHub repositories, malware analysis platforms (e.g., Malware Bazaar), and penetration testing frameworks (e.g., Nishang). The dataset was developed through multi-stage screening, automated and manual deobfuscation, and semantic annotation, ensuring the executability and semantic clarity of the samples. The core applications of this dataset lie in training and evaluating the capability of Large Language Models (LLMs) to generate malicious code, assisting security analysts in gaining in-depth understanding of AI-driven cyber attack patterns, and providing benchmark support for intrusion detection and threat hunting.




