遇见数据集

williegeodev/dga-prediction-muticlass-dataset

收藏
Hugging Face2026-04-25 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含约57,000个域名记录,从二级域名字符串中提取了20个手工设计的词法特征,用于DGA(域名生成算法)僵尸网络家族归属任务,覆盖27个恶意软件家族和一个良性类别。恶意域名通过执行开源算法生成,每个家族样本上限为3,000个以限制类别主导;良性域名从Tranco排名前一百万的列表中采样,以匹配恶意样本总数。数据集支持分层分类设计:第二阶段区分15个主要家族(样本数≥1,000),第三阶段区分11个次要家族(样本数在100至703之间)。数据集已按80/20的比例进行分层训练-测试分割,并在所有相关实验中设置随机种子为42以确保可复现性。

This dataset contains approximately 57,000 domain name records with 20 handcrafted lexical features extracted from second-level domain strings, constructed for DGA botnet family attribution across 27 malware families and one benign class. DGA domain strings were generated by executing the open-source family implementations in the baderj/domain_generation_algorithms repository, with each family capped at 3,000 samples to limit class dominance. Benign domains were sampled from the Tranco top-one-million list to match the total malicious sample count. The dataset supports a hierarchical classification design in which 15 major families with 1,000 or more samples are distinguished at Stage 2, and 11 minor families with between 100 and 703 samples are attributed at Stage 3. The dataset was partitioned using an 80/20 stratified train-test split with random_state=42 in all associated experiments.

提供机构:
williegeodev
二维码
社区交流群
二维码
科研交流群
商业服务