quantum-machine-learning-models
收藏资源简介:
Neura Parse量子机器学习模型数据集是一个专注于量子计算与机器学习交叉领域的综合性、多格式数据集。该数据集旨在为量子计算感知的AI系统的研究和开发提供实践性、代码优先的垂直领域资源。数据集包含16,123条记录,涵盖六种记录类型:概念条目(3,704条)、开放问答(3,409条)、可运行代码任务(3,136条)、多项选择问答(2,833条)、指令/响应对(2,143条)以及语料文本(898条)。内容按难度分为入门级(4条)、本科级(2,848条)、研究生级(10,390条)和研究级(2,881条)。数据集系统化地组织了量子机器学习模型的核心主题,包括:数据编码与特征映射、变分分类器与量子神经网络、量子核与量子支持向量机、生成与能量基量子模型、序列/视觉/强化学习/光子架构,以及训练机制与端到端流水线。该数据集混合了指令/响应对、开放和多项选择问答、可执行代码任务和百科全书式概念条目,采用统一模式,因此既适用于监督微调,也适用于评估/基准测试和持续预训练。数据来源于混合途径,包括专家策划的研究分类法、LLM合成,以及基于arXiv预印本和官方文档的确定性合成种子,并经过严格的质量门控和验证流程。数据集附有详细的验证措施,确保代码可执行性、答案准确性、引用真实性,并明确标注了经典基线或无声称量子优势的声明。数据集主要用于研究和开发目的,其合成记录虽经验证但仍可能包含错误,不应被视为权威的科学参考。
The Neura Parse Quantum Machine Learning Model Dataset is a comprehensive, multi-format dataset focused on the intersection of quantum computing and machine learning. It aims to provide practical, code-first vertical domain resources for the research and development of quantum computing-aware AI systems. The dataset contains 16,123 records, covering six record types: concept entries (3,704), open Q&A (3,409), runnable code tasks (3,136), multiple-choice Q&A (2,833), instruction/response pairs (2,143), and corpus text (898). Content is categorized by difficulty into beginner (4), undergraduate (2,848), graduate (10,390), and research (2,881) levels. It systematically organizes core topics in quantum machine learning models, including: data encoding and feature mapping (how to embed classical data into quantum states), variational classifiers and quantum neural networks (supervised models based on parameterized quantum circuits), quantum kernels and quantum support vector machines (QSVMs), generative and energy-based quantum models (such as quantum GANs, circuit Born machines), sequence/vision/reinforcement learning/photonic architectures (e.g., quantum convolutional networks, quantum attention/transformers), and training mechanisms and end-to-end pipelines. The dataset mixes instruction/response pairs, open and multiple-choice Q&A, executable code tasks, and encyclopedic concept entries in a unified schema, making it suitable for supervised fine-tuning, evaluation/benchmarking, and continued pre-training. Data comes from mixed sources, including expert-curated research taxonomies, LLM synthesis, and deterministic, source-grounded synthetic seeds based on arXiv preprints from 2025-2026 and official documentation from IBM/Google/Microsoft/Quantinuum/NIST/OpenQASM/QIR, with rigorous quality gating and validation processes. The dataset includes detailed verification measures to ensure code executability, answer accuracy, citation authenticity, and clear labeling of classical baselines or no quantum advantage claimed statements. It is primarily intended for research and development purposes; although synthetic records are verified, they may still contain errors and should not be considered authoritative scientific references.
数据集概述
数据集名称: Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures
Hub ID: Neura-parse/quantum-machine-learning-models
许可证: CC-BY-4.0
语言: 英语
数据集规模: 100K < n < 1M(共计 115,429 行)
数据集版本: v3.1.0
数据集划分: train 和 test,数据以 Parquet 格式存储。
记录格式: 包含六种格式:code、concept、corpus、instruction、qa_mcq、qa_open。
主要来源字段: source_url,所有行均带有来源验证标签 source=neura-parse-research。
适用场景: 监督微调、评估/基准测试、检索增强生成(RAG)以及继续预训练。
数据集内容与主题
该数据集是一个 多格式、来源可验证的研究数据集,专注于量子机器学习模型。涵盖以下主题领域:
- 数据编码与特征映射:基础编码、振幅编码、角度编码、IQP/ZZ 编码、数据重上传等。
- 变分分类器与量子神经网络:基于参数化量子电路的监督模型,包括 EstimatorQNN/SamplerQNN、混合 Torch/Keras 层、迁移学习、量子自编码器。
- 量子核与 QSVM:基于特征映射电路的保真度/重叠核,核目标对齐,以及对真实数据集的评估。
- 生成式与基于能量的量子模型:量子 GAN、电路波恩机、量子玻尔兹曼机、量子扩散模型与归一化流模型。
- 序列、视觉、强化学习与光子学架构:量子卷积网络、量子/混合注意力与 Transformer、量子强化学习智能体、连续变量/光子学神经网络。
- 训练机制与端到端流水线:参数偏移与伴随梯度、预算分配、小批量训练、编码感知初始化、可重复的端到端流水线及经典基线。
记录类型与用途
| 记录类型 | 数量 | 数据负载 | 最佳用途 |
|---|---|---|---|
qa_open |
34,527 | 开放答案的量子问题 | 推理评估、RAG 答案生成、辅导 |
instruction |
23,979 | 指令与答案对 | 监督微调、助手行为塑造、任务跟随 |
code |
23,753 | 可执行的量子/软件任务 | 代码生成、代码审查、工具使用评估 |
qa_mcq |
23,554 | 带答案草稿的多选题 | 基准测试、评分、对比评估 |
concept |
9,486 | 结构化概念条目 | 词汇表、检索、课程构建 |
corpus |
130 | 预训练风格的技术段落 | 继续预训练和来源支持的上下文 |
数据组成
按难度划分:
| 难度 | 数量 |
|---|---|
| intro | 1 |
| undergrad | 16,083 |
| graduate | 77,838 |
| research | 21,507 |
总计: 115,429 行
数据模式
每行都包含公共字段:id、domain、record_type、category、topic、subtopics、difficulty、language、source、source_url、license、tags、provenance、quality、metadata。此外,每种记录类型有特定的字段:
qa_open:question、answerinstruction:prompt、responsecode:prompt、code、expected_outputqa_mcq:question、choices、answer、answer_indexconcept:term、definitioncorpus:text
数据质量与来源验证
- 每个
code记录均可端到端执行,并输出所述指标。 - 每个
qa_mcq记录包含四个选项并唯一正确。 - 所有 arXiv ID、API/类名称均经过验证,无伪造引用。
- 分类器、核和生成记录均包含诚实的经典基线或明确的“无声称量子优势”声明。
- 编码和成本声明经过数值检查。
- 概念和词汇的数学符号准确,每个记录均可追溯至所列来源。
- 所有行均携带
source_url来源证明,并经过模式有效性、分类法匹配、去重、活跃源 URL、arXiv ID 验证和代码编译/执行检查。




