CIRCL/vulnerability-attack-techniques-llm-ollama-qwen3.5-122b
收藏资源简介:
该数据集是一个通过大语言模型(ollama/qwen3.5:122b)自动生成的扩展数据集,用于增强漏洞攻击技术分析。它基于MITRE CTID的将ATT&CK映射到CVE以评估影响方法,扩展了人工标注的基准数据集(CIRCL/vulnerability-attack-techniques)。数据集包含297个CVE条目,使用MITRE ATT&CK框架版本19.1进行标注,每行的标签来源(label_sources)均为[llm],表示由大语言模型生成,并记录具体的模型信息。与人工标注数据集的验证结果显示,在121个CVE的测试集上,F1微平均分数为0.392。数据集设计用于与人工标注数据结合,以提升训练效果,通过保留label_sources字段可区分自动生成和人工标注数据。数据集特征包括漏洞ID、标题、描述、利用技术、主要影响、次要影响、技术、衍生技术、标签来源、ATT&CK版本、大语言模型模型和模型评论等。
This dataset is a machine-generated expansion labeled by the large language model (ollama/qwen3.5:122b), not analyst-curated. It follows the MITRE CTID Mapping ATT&CK to CVE for Impact methodology as an expansion of the curated gold dataset (CIRCL/vulnerability-attack-techniques). The dataset includes 297 CVEs, uses ATT&CK version 19.1, and every row has label_sources as [llm], with the llm_model column recording the exact model per row. Validation against the gold set shows an f1_micro score of 0.392 on the full 121-CVE gold test split. It is intended for training augmentation alongside the analyst-curated gold set, and the label_sources column is kept to allow recovery of gold-only data. Features include id, title, description, exploitation_techniques, primary_impact, secondary_impact, techniques, techniques_derived, label_sources, attack_version, llm_model, and llm_comment.




