quantum-machine-learning-theory
收藏资源简介:
Neura Parse — 量子机器学习理论数据集是一个专注于量子模型与量子数据学习理论的研究深度、证明导向的垂直领域数据集。它涵盖了参数化量子电路为何能够或无法训练(贫瘠高原问题)、其表达能力、泛化性能以及何时能在理论上超越经典模型等核心议题。对于量子数据,它还探讨了如何通过少量测量(如经典影子、影子层析成像)预测未知量子态或通道的性质,以及量子内存何时能带来指数级学习优势。数据集包含14,823条记录,采用多格式混合结构,具体包括指令/回复对、开放式与多项选择题问答、可运行代码任务以及百科全书式的概念条目。这些内容被组织在五个主要分类下:可训练性与贫瘠高原、表达能力与泛化、量子核与学习分离、从量子数据中学习(影子与层析)、量子内存优势与下界。每条记录都包含通用元数据(如ID、领域、记录类型、类别、主题、难度、来源、许可证等)和特定于记录类型的字段。数据通过混合方式生成,结合了专家策划和基于2025-2026年arXiv预印本及官方量子计算文档的确定性LLM合成,并经过严格的质量验证流程(包括范围控制、引用验证、代码执行、数学一致性检查等)。该数据集旨在用于量子计算感知人工智能系统的研究与开发,但请注意,其合成记录虽经验证,仍可能包含错误,不应被视为权威的科学参考。
Neura Parse — Quantum Machine Learning Theory Dataset is a proof-focused vertical research dataset dedicated to the in-depth theoretical research of quantum models and quantum data learning. It covers core topics including why parameterized quantum circuits (PQCs) can or cannot be trained (the barren plateaus problem), their expressive power, generalization performance, and under what conditions they can theoretically outperform classical models. For quantum data, it also explores how to predict properties of unknown quantum states or channels via limited measurements such as classical shadows and shadow tomography, as well as when quantum memory can confer exponential learning advantages. The dataset contains 14,823 records with a mixed multi-format structure, specifically including instruction-response pairs, open-ended and multiple-choice question answering tasks, runnable code tasks, and encyclopedic concept entries. These contents are organized under five main categories: Trainability and Barren Plateaus, Expressive Power and Generalization, Quantum Kernels and Learning Separation, Learning from Quantum Data (Shadows and Tomography), and Quantum Memory Advantages and Lower Bounds. Each record includes general metadata (such as ID, domain, record type, category, topic, difficulty level, source, license, etc.) and fields specific to the record type. The dataset is generated via a hybrid approach, combining expert curation and deterministic LLM synthesis based on 2025-2026 arXiv preprints and official quantum computing documentation, and has undergone strict quality validation procedures including scope control, citation verification, code execution, mathematical consistency checks, and more. This dataset is intended for research and development of quantum-computing-aware artificial intelligence systems. Please note that although its synthesized records have been validated, they may still contain errors and should not be regarded as authoritative scientific references.
数据集概述
- 数据集名称: Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data
- Hub ID:
Neura-parse/quantum-machine-learning-theory - 发布版本: v3.1.0
- 总行数: 109,283
- 数据集规模: 100K < n < 1M
- 语言: 英语 (en)
- 许可证: CC BY 4.0
- 数据集分割:
train、test - 数据类型: 包括
code、concept、corpus、instruction、qa_mcq、qa_open等多种格式 - 主要来源字段:
source_url
数据集内容与主题
本数据集是一个深度、以证明为导向的垂直领域数据集,专注于量子模型和量子数据的学习理论。涵盖以下五大主题类别:
- 可训练性与贫瘠高原 (Trainability & Barren Plateaus): 解释参数化量子电路为何能或不能训练(贫瘠高原),包括其分类、通过动力学李代数的精确方差缩放定律、缓解策略,以及即使没有高原也持续存在的更深层障碍(陷阱、NP-hardness)。包含4个子主题。
- 表达能力、容量与泛化 (Expressivity, Capacity & Generalization): 探讨参数化量子电路模型能表示什么以及如何从少量数据中泛化:编码的通用性和傅里叶图像、可表达性/纠缠能力与t-design度量、以及基于门计数/有效维度/覆盖数的泛化界限。包含3个子主题。
- 量子核、数据与学习分离 (Quantum Kernels, Data & Learning Separations): 涉及量子核理论(特征映射、指数级集中、归纳偏置、经典估计的困难)、数据的威力、经典替代模型与去量子化(dequantization),以及严格的可证明的量子vs经典学习分离。包含3个子主题。
- 从量子数据中学习:影子与层析成像 (Learning From Quantum Data: Shadows & Tomography): 关注从少量测量中预测未知状态、通道和哈密顿量的性质:经典影子(随机Clifford/Pauli、中位数均值)、影子层析成像与温和测量、状态的可PAC学习,以及Pauli/噪声通道学习。包含3个子主题。
- 量子记忆优势与下界 (Quantum Memory Advantages & Lower Bounds): 讨论纠缠的多副本测量和量子记忆何时能带来可证明的、通常是指数级的学习优势("从实验中学习"),以及匹配的信息论样本复杂度下界和学习困难结果。包含2个子主题。
记录类型与数量
| 记录类型 | 数量 | 载荷内容 | 最佳用途 |
|---|---|---|---|
qa_mcq |
37,423 | 多项选择题及解答草图 | 基准测试、评分、对比评估 |
qa_open |
35,245 | 开放式量子问题 | 推理评估、RAG答案生成、辅导 |
instruction |
25,112 | 指令与答案对 | 监督微调、助理行为塑造、任务跟随 |
concept |
11,357 | 结构化概念条目 | 词汇表、检索、课程构建 |
corpus |
144 | 预训练风格的技术段落 | 持续预训练和来源支持的上下文 |
code |
2 | 小型可执行范例集 | 抽查和示例 |
| 总计 | 109,283 |
按难度划分
| 难度 | 数量 |
|---|---|
| 本科 (undergrad) | 724 |
| 研究生 (graduate) | 75,174 |
| 研究级 (research) | 33,385 |
数据模式 (Schema)
每条记录共享一个通用信封(包含 id、domain、record_type、category、topic、subtopics、difficulty、language、source、source_url、license、tags、provenance、quality、metadata),此外还有特定于其记录类型的字段:
| 记录类型 | 特定字段 |
|---|---|
qa_mcq |
question, choices, answer, answer_index |
qa_open |
question, answer |
instruction |
prompt, response |
concept |
term, definition |
corpus |
text |
code |
prompt, code, expected_output |
来源验证与质量控制
- 来源验证: 该数据集是 v3.1.0 来源验证版本。每一条发布的行都携带
source_url出处,并标记为source=neura-parse-research。 - 质量门禁: 包含多道质量关卡,例如:范围强制(每条记录映射到特定主题)、引用完整性(arXiv ID / DOI 必须真实)、MCQ有效性、代码执行、语料格式、数学正确性、难度校准和去重。
推荐工作流程
- 用于量子计算感知助手的监督微调。
- 量子推理的多项选择和开放式回答评估。
- 基于有源量子及量子AI主题的检索增强生成。
- 需要基于量子研究记录进行检索、解释和评估的工作流程。
- 在结构化的、有源的技术文本上进行持续预训练。
引用
bibtex @misc{neuraparse_quantum_machine_learning_theory, title = {Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data}, author = {Neura Parse}, year = {2026}, url = {https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory} }





