eiphuggincve/cve-cwe-consensus
收藏资源简介:
CVE到CWE共识数据集是一个多标签数据集,用于将CVE漏洞描述映射到其CWE弱点类型,专为微调指令调优的大型语言模型(例如使用Unsloth)而构建。每个标签都是共识分配:仅当NVD和CVE编号机构(CNA)独立同意,并将两者汇总到CWE View-1003(约130个弱点的“简化已发布漏洞映射弱点”)时,才保留CWE。数据集包含116,793个示例,格式为对话式消息(系统/用户/助手),适用于聊天模板监督微调。标签基于两个独立官方来源(NVD和CNA)的交集构建,从而排除了单源错误。数据集用于CVE描述分类为CWE类型的模型微调或评估,支持漏洞分类、丰富和优先级辅助任务。
A multi-label dataset mapping CVE vulnerability descriptions to their CWE weakness type(s), built for fine-tuning instruction-tuned LLMs (e.g. with Unsloth). Each label is a consensus assignment: a CWE is kept only when NVD and the CVE Numbering Authority (CNA) independently agree on it, after rolling both up to CWE View-1003 (the ~130-weakness Weaknesses for Simplified Mapping of Published Vulnerabilities). The dataset contains 116,793 examples in a conversational messages format (system/user/assistant), ready for chat-template supervised fine-tuning. Labels are the intersection of two independent official sources (NVD and CNA), excluding single-source errors by construction. It is intended for fine-tuning or evaluating models that classify CVE descriptions into CWE types, supporting vulnerability triage, enrichment, and prioritization tasks.




