quantum-information-and-complexity-theory
收藏资源简介:
Neura Parse — 量子信息与复杂性理论:信道、熵、类与优势结构数据集是一个专注于量子信息理论和量子复杂性理论交叉领域的大规模、多格式文本数据集,旨在为量子计算感知的人工智能系统提供研究和发展基础。数据集包含17,211条记录,涵盖六种记录类型:开放式问答(4,372条)、多项选择问答(4,348条)、概念条目(3,747条)、指令/响应对(2,861条)、语料文本(1,857条)和代码任务(26条)。内容按难度分为入门级(1条)、本科级(2,140条)、研究生级(10,129条)和研究级(4,941条)。知识体系分为五个核心领域:1)量子态、信道与操作资源;2)熵与可区分性;3)纠缠理论与量子香农理论;4)量子复杂性类与哈密顿量复杂性;5)量子优势的结构。数据集采用混合来源,结合专家策划和基于2025-2026年arXiv及官方量子源文件的确定性代码合成,并经过严格质量控制,确保模式有效性、引用完整性等。适用于监督微调、评估/基准测试和持续预训练,但需注意合成记录可能包含错误,不应视为权威科学参考。
Neura Parse — Quantum Information and Complexity Theory: Channels, Entropy, Classes, and Advantage Structures Dataset is a large-scale, multi-format text dataset focusing on the intersection of quantum information theory and quantum complexity theory. It aims to provide a research and development foundation for quantum computing-aware AI systems by systematically covering topics from generalized datasets in information theory and complexity classes. The dataset contains 17,211 records across six types: open-ended questions (4,372), multiple-choice questions (4,348), concept entries (3,747), instruction/response pairs (2,861), corpus texts (1,857), and code tasks (26). Content is categorized by difficulty into beginner (1 record), undergraduate (2,140), graduate (10,129), and research (4,941) levels. The knowledge framework is divided into five core areas: 1) Quantum states, channels, and operational resources; 2) Entropy and distinguishability; 3) Entanglement theory and quantum Shannon theory; 4) Quantum complexity classes and Hamiltonian complexity; 5) Structures of quantum advantage. The dataset is sourced from a hybrid approach, combining expert curation and deterministic code synthesis based on arXiv and official quantum source files from 2025-2026, with rigorous quality control for pattern validity, citation integrity, mathematical consistency, theorem accuracy, multiple-choice completeness, code executability, deduplication, and license compliance. It is suitable for supervised fine-tuning, evaluation/benchmarking, and continued pre-training, but note that synthetic records may contain errors and should not be considered authoritative scientific references.
数据集概述
基本信息
- 数据集名称: Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage
- Hub ID:
Neura-parse/quantum-information-and-complexity-theory - 发布版本: v3.1.0
- 数据行数: 108,984
- 许可证: CC BY 4.0
- 语言: 英语
- 数据集划分:
train、test - 数据格式: Parquet
记录类型与数量
| 记录类型 | 数量 | 说明 |
|---|---|---|
qa_mcq |
36,937 | 多项选择题,含答案草图 |
qa_open |
35,459 | 开放式问答 |
instruction |
24,829 | 指令与回答对 |
concept |
11,583 | 结构化概念条目 |
corpus |
175 | 预训练风格技术文本 |
code |
1 | 可执行代码示例 |
| 总计 | 108,984 |
按难度划分
| 难度 | 数量 |
|---|---|
| 本科生 | 16,188 |
| 研究生 | 57,929 |
| 研究级 | 34,867 |
主题分类
- 量子态、信道与操作资源 — 密度算子、CPTP信道、Kraus/Stinespring/Choi表示、噪声信道、不可克隆/不可广播定理、量子隐形传态与超密编码(4个主题)
- 熵与可区分性 — von Neumann熵、条件熵、互信息、相对熵、Rényi熵及其不等式,以及状态与信道的可区分性度量(3个主题)
- 纠缠理论与量子Shannon理论 — LOCC、PPT/可分离性、纠缠度量、量子Shannon编码定理(2个主题)
- 量子复杂性类与Hamiltonian复杂性 — BQP、QMA、QCMA、QIP、PostBQP=PP等复杂性类及其关系,以及基态能量估计的复杂性(2个主题)
- 量子优势的结构 — 采样优势、反集中与XEB、验证、查询/通信下界、伪随机态与酉算子、去量子化(4个主题)
数据模式(Schema)
每条记录共享公共字段:id、domain、record_type、category、topic、subtopics、difficulty、language、source、source_url、license、tags、provenance、quality、metadata
各记录类型特有字段:
qa_mcq:question、choices、answer、answer_indexqa_open:question、answerinstruction:prompt、responseconcept:term、definitioncorpus:textcode:prompt、code、expected_output
来源验证
- 来源已验证的发布版本 v3.1.0
- 每行数据均带有
source_url溯源信息,标签为source=neura-parse-research - 历经完整性检查:697个来源URL,0个无效;513个arXiv ID,0个伪造
- 177,532条代码记录通过编译,0次编译失败
推荐应用场景
- 面向量子计算助手的监督微调
- 量子推理的多项选择和开放式评估
- 基于检索增强生成(RAG)的量子与量子AI主题
- 基于结构化、有源可查技术文本的继续预训练




