Dataset for AI-Driven Cloud Misconfiguration Risk Classification in SOC Environments
收藏资源简介:
This dataset was generated as part of the research project titled “Enabling AI-Driven Cloud SOC Resilience”. The dataset contains 30,000 synthetically generated cloud configuration samples designed to support supervised machine learning-based risk classification of cloud misconfigurations within Security Operations Center (SOC) environments. Each sample represents a cloud resource configuration and includes features associated with common cloud security misconfigurations, including:- Publicly accessible storage- Over-permissive IAM policies- Open network ports- Disabled encryption- Unsecured endpoints- Disabled logging and monitoring A numerical risk scoring mechanism ranging from 1 to 5 was used during dataset generation to represent the severity and combination of misconfigurations. These scores were subsequently mapped into three risk classes:- Low Risk (0): score 1- Medium Risk (1): scores 2–3- High Risk (2): scores 4–5 The dataset was structured as flat tabular data in CSV format to facilitate preprocessing and compatibility with machine learning workflows. A balanced class distribution was maintained across the three risk categories. This dataset was developed to address the limited availability of publicly accessible cloud misconfiguration datasets and to support reproducibility of the experiments presented in the associated research work. It is intended for academic and research purposes related to cloud security, AI-driven SOC operations, risk scoring, and cybersecurity analytics.
本数据集是为完成题为“赋能AI驱动的云安全运营中心韧性(Enabling AI-Driven Cloud SOC Resilience)”的研究项目而生成的。 该数据集包含30000条人工合成的云资源配置样本,旨在支持安全运营中心(Security Operations Center, SOC)环境内云配置错误的基于监督式机器学习的风险分类任务。 每条样本对应一条云资源配置,包含与常见云安全配置错误相关的特征,具体包括: - 公开可访问的存储资源 - 权限过度宽松的身份与访问管理(Identity and Access Management, IAM)策略 - 开放的网络端口 - 加密功能已禁用 - 未设防的端点 - 日志与监控功能已禁用 数据集生成过程中采用了1至5分的数值化风险评分机制,用以表征配置错误的严重程度及其组合影响。随后将上述评分映射为三类风险等级: - 低风险(0):评分1 - 中风险(1):评分2–3 - 高风险(2):评分4–5 本数据集采用CSV格式的扁平化表格数据结构,以适配预处理流程并兼容各类机器学习工作流。三类风险类别间保持了均衡的样本分布。 本数据集旨在解决当前公开可用的云配置错误数据集匮乏的问题,并支撑相关研究工作中实验的可复现性。其适用于与云安全、AI驱动的SOC运营、风险评分以及网络安全分析相关的学术与研究用途。



