quantum-cryptography-and-post-quantum-security
收藏资源简介:
Neura Parse 量子密码学与后量子安全数据集是一个专注于量子密码学和抗量子攻击经典密码学领域的深度垂直数据集。它涵盖了量子密钥分发(包括 BB84、B92、六态、SARG04、E91、BBM92、诱饵态、MDI-QKD、TF-QKD、CV-QKD 等协议)、设备无关协议、可组合与有限密钥安全性证明、量子黑客攻击及对策、经典后处理(协调、隐私放大、认证)、量子随机数生成与认证随机性,以及量子货币、抛币、比特承诺不可行定理和量子数字签名等密码学原语。在后量子密码学方面,数据集覆盖了 NIST 标准化算法(如 FIPS 203 ML-KEM、FIPS 204 ML-DSA、FIPS 205 SLH-DSA、草案 FIPS 206 FN-DSA 以及 2025 年选定的 HQC)、基于格/编码/哈希/同源/多元的密码家族、现在收获,以后解密威胁,以及密码敏捷性迁移(混合密钥交换、TLS/PKI、NIST IR 8547 和 CNSA 2.0 时间线)。数据集共包含 16,548 条记录,采用多格式混合结构,包含 `code`、`concept`、`corpus`、`instruction`、`qa_mcq`(多项选择问答)和 `qa_open`(开放式问答)六种记录类型,并按难度分为入门、本科、研究生和研究四个级别。数据通过专家策划的分类法、LLM 合成以及基于 2025-2026 年 arXiv 预印本和官方文档的确定性 Codex 生成方法构建,并经过严格的质量门控验证。该数据集适用于量子计算感知 AI 系统的研究与发展,可用于监督微调、评估基准测试和持续预训练。但需注意,数据集包含模型生成的合成记录,可能存在错误,不应作为权威科学参考。
The Neura Parse Quantum Cryptography and Post-Quantum Security Dataset is a deep vertical dataset focused on the fields of quantum cryptography and quantum attack-resistant classical cryptography. It covers quantum key distribution (including protocols such as BB84, B92, six-state, SARG04, E91, BBM92, decoy state, MDI-QKD, TF-QKD, CV-QKD, etc.), device-independent protocols, composable and finite-key security proofs, quantum hacking attacks and countermeasures, classical post-processing (coordination, privacy amplification, authentication), quantum random number generation and certified randomness, as well as cryptographic primitives including quantum money, coin flipping, the no-go theorem for bit commitment, and quantum digital signatures. In terms of post-quantum cryptography, the dataset covers NIST standardized algorithms (such as FIPS 203 ML-KEM, FIPS 204 ML-DSA, FIPS 205 SLH-DSA, draft FIPS 206 FN-DSA, and HQC selected in 2025), cryptographic families based on lattice/coding/hash/isogeny/multivariate, "harvest now, decrypt later" threats, and cryptographic agility migration (hybrid key exchange, TLS/PKI, NIST IR 8547 and CNSA 2.0 timeline). The dataset contains a total of 16,548 records, adopting a multi-format hybrid structure with six record types: `code`, `concept`, `corpus`, `instruction`, `qa_mcq` (multiple-choice question answering) and `qa_open` (open-ended question answering), and is categorized into four difficulty levels: beginner, undergraduate, graduate, and research. The data is constructed through expert-curated taxonomy, LLM synthesis, and deterministic Codex generation methods based on 2025-2026 arXiv preprints and official documents, and has undergone strict quality gate validation. This dataset is suitable for the research and development of quantum-aware AI systems, and can be used for supervised fine-tuning, benchmark evaluation, and continuous pre-training. It should be noted that the dataset contains synthetic records generated by models, which may contain errors and should not be used as an authoritative scientific reference.
数据集概述
这是一个专注于量子密码学与后量子安全的深度垂直领域数据集,由 Neura Parse 构建,旨在为量子计算相关的监督微调、评估、检索增强生成和继续预训练提供支持。
核心信息
| 属性 | 内容 |
|---|---|
| Hub ID | Neura-parse/quantum-cryptography-and-post-quantum-security |
| 发布版本 | v3.1.0 |
| 数据规模 | 106,488 行 |
| 数据划分 | train(训练集)、test(测试集) |
| 许可证 | CC BY 4.0 |
| 语言 | 英语 |
数据覆盖范围
数据涵盖了量子密码学和后量子安全领域的核心主题:
- 量子密钥分发 (QKD):包括 BB84、B92、六态协议、SARG04、E91、BBM92、诱骗态、MDI-QKD、TF-QKD、CV-QKD 等协议。
- 设备无关协议:设备无关密码学与自测试。
- 安全性证明与攻击:组合和有限密钥安全证明、量子黑客攻击与对策、经典后处理(协商、隐私放大、认证)。
- 量子随机数生成:量子随机数生成与认证随机性。
- 量子原语:量子货币、抛币、比特承诺、量子数字签名。
- 后量子算法:NIST 标准算法(ML-KEM、ML-DSA、SLH-DSA、FN-DSA、HQC)以及基于格、编码、哈希、同源、多元的密码学家族。
- 威胁模型与迁移:Harvest-Now-Decrypt-Later 威胁、密码敏捷性(混合密钥交换、TLS/PKI 集成、NIST IR 8547 和 CNSA 2.0 时间线)。
记录类型与用途
数据集包含多种记录格式,适用于不同任务:
| 记录类型 | 数量 | 说明 | 最佳用途 |
|---|---|---|---|
qa_mcq |
35,956 | 多项选择题 | 基准测试、评分、对比评估 |
qa_open |
34,636 | 开放式问答题 | 推理评估、RAG 答案生成、辅导 |
instruction |
24,320 | 指令与回答对 | 监督微调、助手行为塑造 |
concept |
11,402 | 结构化概念条目 | 术语表、检索、课程构建 |
corpus |
171 | 预训练风格的技术段落 | 继续预训练、上下文来源 |
code |
3 | 可执行代码示例 | 示例检查(非核心基准) |
数据组成
- 按难度分布:入门级(2 条)、本科级(21,317 条)、研究生级(66,162 条)、研究级(19,007 条)。
- 来源验证:所有行均包含
source_url溯源,标记为source=neura-parse-research,并经过模式有效性、去重、活跃 URL、arXiv ID 检查等质量门控。
模式结构
每条记录共享通用字段(如 id, domain, record_type, topic, source, source_url 等),并根据记录类型包含特定字段:
- qa_mcq:
question,choices,answer,answer_index - qa_open:
question,answer - instruction:
prompt,response - concept:
term,definition - corpus:
text - code:
prompt,code,expected_output
质量保证
数据集实施了严格的质控门,确保事实准确性(如符合 NIST 标准状态)、QKD 安全性声明清晰、代码可执行、MCQ 选项唯一正确且干扰项合理、内容为教育性质。




