遇见数据集

mjbommar/linux-cve-dossiers

收藏
Hugging Face2026-04-22 更新2026-04-26 收录
官方服务:

资源简介:

Linux CVE Dossier Corpus是一个经过范围审核的语料库,包含11个Linux基础系统包的每个CVE研究档案:Linux内核、glibc、musl、systemd、util-linux、coreutils、BusyBox、OpenSSL、curl、Node.js和CPython。每个档案都包含摘要、时间线、补丁谱系、漏洞利用说明和参考收集,以及结构化导出,供下游消费者使用而无需重新解析markdown。语料库分为in_scope(1,418条记录)和hard_negatives(1,138条记录)两部分,分别包含目标包的审核档案和NVD CPE过匹配案例。数据集还包含详细的模式(列)描述,如身份和位置、NVD信息、生命周期、补丁谱系、证据和来源等。

The Linux CVE Dossier Corpus is a scope-audited corpus of per-CVE research dossiers for 11 Linux base-system packages: Linux kernel, glibc, musl, systemd, util-linux, coreutils, BusyBox, OpenSSL, curl, Node.js, and CPython. Each dossier carries a summary, dated timeline, patch lineage, exploit notes, and reference harvest, plus a structured export that downstream consumers can use without re-parsing the markdown. The corpus is split into in_scope (1,418 records) for the target packages and hard_negatives (1,138 records) for NVD CPE-overmatch cases. The dataset includes detailed schema (columns) such as identity and placement, NVD information, lifecycle, patch lineage, evidence, and provenance.

提供机构:
mjbommar
二维码
社区交流群
二维码
科研交流群
商业服务