Benign and malicious domains based on DNS logs
收藏资源简介:
The dataset is meant for supervised machine learning based analysis of malicious and non-malicious domain names. The dataset was created from scratch, using publicly DNS logs of both malicious and non-malicious domain names. Using the domain name as input, 34 features were obtained. Features like the domain name, entropy, number of strange characters and domain name length were obtained directly from the domain name. Other features like, domains name creation date, IP, open ports, geolocation were obtained from data enrichment processes (e.g. OSINT). This dataset consists of data from 90000 domains names and it is balanced between 50% non-malicious and 50% of malicious domain names.
本数据集旨在用于恶意与非恶意域名的监督式机器学习分析。本数据集从零构建,采集了源自公开渠道的恶意与非恶意域名的域名系统(DNS)日志。以域名为输入时,可提取得到34项特征:其中域名本身、信息熵、特殊字符数量、域名长度等特征可直接从域名文本中获取;而域名创建日期、IP地址、开放端口、地理位置等其余特征,则需通过数据富集流程(如开源情报(Open-Source Intelligence,OSINT))获取。本数据集涵盖90000个域名的相关数据,且类别均衡,恶意域名与非恶意域名各占50%。



