Voidreaper2026/qwen3-4b-training
收藏资源简介:
这是一个大规模网络安全指令调优数据集,采用ShareGPT对话格式,从多个权威开源数据源(包括NVD、OSV、GitHub Advisory Database、Cybersec Causal Reasoning、Security Stack Exchange、ExploitDB、MITRE ATT&CK、CISA KEV、Kali Linux Tools和Vulners)汇编而成,并经过去重处理,包含1,807,941条去重记录。该数据集旨在用于微调大型语言模型(LLM),以支持网络安全推理、漏洞分析、威胁情报、安全运营中心(SOC)助手模型、安全感知编码助手以及CTF/渗透测试知识库等应用。数据格式为JSON,包含human和gpt角色的对话,部分记录还包含系统提示。去重过程使用基于源、源ID和对话内容MD5哈希的智能键,以确保跨源CVE覆盖的同时移除重复项。许可证为Apache 2.0,各源数据集保留其原始开放许可证。
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources (including NVD, OSV, GitHub Advisory Database, Cybersec Causal Reasoning, Security Stack Exchange, ExploitDB, MITRE ATT&CK, CISA KEV, Kali Linux Tools, and Vulners) and deduplicated, with 1,807,941 deduplicated records. The dataset is intended for fine-tuning large language models (LLMs) for cybersecurity reasoning, vulnerability analysis, threat intelligence, SOC assistant models, security-aware coding assistants, and CTF/penetration testing knowledge bases. The data format is JSON, featuring conversations with human and gpt roles, and some records include a system prompt. Deduplication is performed using a smart key combining source, source_id, and MD5 hash of conversation content to retain cross-source CVE coverage while removing true duplicates. The license is Apache 2.0, with individual source datasets retaining their original open licenses.



