遇见数据集

ZTA-RAD - Zero Trust Architecture – Risk Assessment Dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Insider threats constitute one of the main risks to information security, particularly in corporate environments that handle sensitive data. Unlike external threats, these attacks exploit valid credentials and prior knowledge of the infrastructure, rendering traditional perimeter-based security models insufficient. In this context, Zero Trust Architecture (ZTA) emerges as an essential paradigm, adopting the principle of “never trust, always verify” through continuous verification, microsegmentation, and real-time monitoring. Motivated by the lack of datasets specifically designed for this scenario, this work proposes ZTA-RAD (Zero Trust Architecture – Risk Assessment Dataset), derived and expanded from CERT, incorporating logon, device, and HTTP access metrics. The dataset was labeled through two approaches: (i) expert curation and (ii) risk analysis performed by five LLMs. For evaluation, training and validation experiments were conducted with three machine learning algorithms (MLP, Random Forest, and SVM), under both balanced and imbalanced scenarios. The results indicate that MLP and Random Forest achieved superior performance, particularly after balancing with SMOTE, while SVM proved more sensitive to imbalance. Furthermore, it was found that expert labeling enabled the construction of more consistent classifiers compared to LLMs, highlighting the importance of human curation. In summary, this work contributes by providing a novel dataset for the study of insider threats in ZTA, offering methodological and practical support for the advancement of machine learning–based solutions.

内部威胁是信息安全面临的主要风险之一,尤其在处理敏感数据的企业环境中更为突出。与外部威胁不同,此类攻击会利用合法凭证与对基础设施的先验知识,使得传统基于边界的安全模型难以发挥作用。在此背景下,零信任架构(Zero Trust Architecture, ZTA)作为至关重要的安全范式应运而生,其秉持“永不信任、始终验证”的原则,通过持续验证、微分段与实时监控实现安全防护。鉴于当前缺乏针对该场景的专用数据集,本研究提出ZTA-RAD(零信任架构-风险评估数据集,Zero Trust Architecture – Risk Assessment Dataset),该数据集源自CERT并进行了扩展,纳入了登录、设备与HTTP访问相关指标。该数据集通过两种方式完成标注:(i) 专家人工审核标注;(ii) 由五个大语言模型(Large Language Model, LLM)开展风险分析实现标注。为开展评估,本研究针对平衡与非平衡两种数据集场景,使用三种机器学习算法(多层感知机MLP、随机森林Random Forest与支持向量机SVM)开展了训练与验证实验。实验结果显示,多层感知机与随机森林的性能更为优异,尤其是在通过SMOTE进行数据平衡处理后;而支持向量机则对数据非平衡状态更为敏感。此外,研究发现相较于大语言模型标注,专家标注能够构建出一致性更强的分类器,凸显了人工审核标注的重要价值。综上,本研究为零信任架构下的内部威胁研究提供了全新的专用数据集,为基于机器学习的安全解决方案的迭代升级提供了方法论与实践层面的支撑。

创建时间:
2025-09-03
二维码
社区交流群
二维码
科研交流群
商业服务