fault-tolerant-quantum-computing
收藏资源简介:
Neura Parse — Fault-Tolerant Quantum Computing 是一个专注于容错量子计算领域的多格式数据集。它涵盖了量子纠错(QEC)码族、解码器、容错门构造以及从物理到逻辑的完整资源估计流程,基于Stim框架进行模拟。数据集旨在扩展纠错主题至研究级覆盖,包括2024-2026年的里程碑成果,如低于阈值的表面码、qLDPC/双变量自行车码存储器和魔法态培育。数据集包含12,551条英语记录,采用CC-BY-4.0许可证,版本为2.0.0。它混合了指令/响应对、开放式和多选题问答、可运行代码任务以及百科全书式概念条目,统一模式适用于监督微调、评估/基准测试和持续预训练。记录类型包括:开放式问答(3,458条)、多选题问答(3,153条)、概念条目(2,975条)、指令对(2,155条)、语料段落(786条)和代码示例(24条)。难度分布涵盖本科(33条)、研究生(7,187条)和研究级别(5,331条)。内容按五大分类组织:1) 稳定子与拓扑QEC码;2) 量子LDPC与低开销存储器;3) 解码器与探测器错误模型;4) 容错逻辑与魔法态;5) 阈值、噪声与资源估计。每条记录共享公共元数据字段(如ID、领域、记录类型、类别、主题、难度、语言、来源、许可证等),并具有类型特定字段。数据来源为混合式:初始版本基于专家策划的研究分类法,后续版本通过确定性Codex生成,并融合了来自2025-2026年arXiv预印本和IBM/Google/Microsoft等官方量子计算文档的合成种子,所有记录均经过质量门验证。数据集适用于量子计算感知AI系统的研究和开发,但需注意其合成记录虽经验证仍可能包含错误,不应视为权威科学参考,关键事实需对照原始文献核实。
Neura Parse — Fault-Tolerant Quantum Computing is a multi-format dataset focused on the field of fault-tolerant quantum computing. It covers quantum error correction (QEC) code families, decoders, fault-tolerant gate construction, and complete resource estimation pipelines from physical to logical qubits, with simulations conducted using the Stim framework. The dataset aims to extend coverage of error correction topics to the research level, including milestone achievements from 2024 to 2026 such as sub-threshold surface codes, qLDPC/bivariate bicycle code memories, and magic state distillation. The dataset contains 12,551 English-language records, is licensed under CC-BY-4.0, and is at version 2.0.0. It combines instruction-response pairs, open-ended and multiple-choice question answering tasks, runnable code tasks, and encyclopedic concept entries, with a unified schema suitable for supervised fine-tuning, evaluation/benchmarking, and continual pre-training. Record types include: open-ended question answering (3,458 entries), multiple-choice question answering (3,153 entries), concept entries (2,975), instruction-response pairs (2,155), corpus paragraphs (786), and code examples (24). The difficulty distribution spans undergraduate-level (33 entries), graduate-level (7,187 entries), and research-level (5,331 entries). The content is organized into five major categories: 1) Stabilizer and topological QEC codes; 2) Quantum LDPC and low-overhead memories; 3) Decoders and detector error models; 4) Fault-tolerant logic and magic states; 5) Thresholds, noise, and resource estimation. Each record shares common metadata fields (e.g., ID, domain, record type, category, topic, difficulty, language, source, license, etc.) and includes type-specific fields. The dataset employs hybrid data sourcing: the initial version is based on expert-curated research taxonomies, while subsequent versions are generated via deterministic Codex and incorporate synthetic seeds from 2025–2026 arXiv preprints and official quantum computing documentation from organizations including IBM, Google, and Microsoft. All records have undergone quality gate validation. This dataset is suitable for research and development of quantum computing-aware AI systems. Note that although its synthetic records have been validated, they may still contain errors and should not be treated as authoritative scientific references; critical factual claims must be cross-verified against original literature.
数据集概述
数据集名称: Neura Parse — Fault-Tolerant Quantum Computing: QEC Codes, Decoders, Magic States & Resource Estimation
发布版本: v3.1.0
语言: 英语 (en)
许可证: CC BY 4.0
数据规模: 109,594 行 (100K < n < 1M)
数据集ID: Neura-parse/fault-tolerant-quantum-computing
数据划分与格式
- 划分 (Splits):
train和test - 存储格式: Parquet
- 记录类型 (Record Types):
qa_mcq(37,487 条): 多项选择题,附带答案草图qa_open(35,718 条): 开放式量子问题instruction(25,190 条): 指令与答案对concept(11,054 条): 结构化概念条目corpus(145 条): 预训练风格的技术段落
难度分布
| 难度 | 数量 |
|---|---|
| 本科 (undergrad) | 8 |
| 研究生 (graduate) | 64,111 |
| 研究 (research) | 45,475 |
主题分类 (Taxonomy)
数据集涵盖 14 个主题,分为五大领域:
- 稳定子与拓扑 QEC 码 (4 个主题): 表面/环面码、颜色码、Floquet/蜂窝码、子系统与 Bacon-Shor 码
- 量子 LDPC 与低开销存储器 (2 个主题): 双变自行车码、超图/提升/平衡积码
- 解码器与检测器错误模型 (3 个主题): MWPM/稀疏 blossom、并查集、置信传播+OSD、张量网络/相关解码器
- 容错逻辑与魔法态 (3 个主题): 横向门与 Eastin-Knill、代码切换/形变、晶格手术与编织、魔法态蒸馏与培育
- 阈值、噪声与资源估计 (3 个主题): 阈值定理与电路级噪声、Stim/Sinter 逻辑错误基准测试、物理到逻辑的资源估计管线
数据模式 (Schema)
每条记录包含公共字段 (id, domain, record_type, category, topic, subtopics, difficulty, language, source, source_url, license, tags, provenance, quality, metadata) 以及记录类型特定的字段:
| 记录类型 | 特定字段 |
|---|---|
qa_mcq |
question, choices, answer, answer_index |
qa_open |
question, answer |
instruction |
prompt, response |
concept |
term, definition |
corpus |
text |
来源验证 (Source Verification)
- 所有行均带有
source_url溯源,标签为source=neura-parse-research - 审查通过标准: 模式有效性、分类匹配、去重、活跃源 URL、arXiv-ID 检查、代码编译与执行
- 验证结果: 697 个源 URL 无坏链,513 个 arXiv ID 验证无伪造,177,532 条代码记录编译通过
推荐工作流
- 用于量子计算助手的监督微调
- 量子推理的多项选择与开放式评估
- 基于溯源量子主题的检索增强生成
- 检索、解释与评估工作流
- 在结构化、有源的技术文本上进行持续预训练




