quantum-simulation-chemistry-materials
收藏资源简介:
Neura Parse — 量子模拟化学与材料:编码、VQE/QPE与动力学数据集是一个专注于量子计算在化学和材料科学模拟应用领域的专业数据集。该数据集旨在为研究和开发量子计算感知的人工智能系统提供支持。数据内容深度覆盖量子模拟物质的核心主题,包括:电子结构与费米子到量子比特的编码(将化学/材料问题转化为量子比特哈密顿量,涉及二次和一次量子化电子结构哈密顿量、经典PySCF预处理、各种费米子-量子比特编码如Jordan-Wigner、parity、Bravyi-Kitaev等,以及基于Z2对称性的量子比特缩减);哈密顿量分解与容错资源估计(通过单/双/张量超收缩分解压缩双电子张量及其对哈密顿量1-范数和块编码成本的影响,针对FeMoco、催化、阴极等案例的端到端容错资源估计,以及对量子方法必须超越的经典竞争对手的客观评估);基态与激发态算法(提取本征态和性质的算法,如VQE变体与化学拟设、量子相位估计、量子子空间/Krylov和虚时方法、测量分组与采样预算,以及激发态、格林函数、响应和有限温度方法);动力学、凝聚态模型与模拟仿真(模拟量子物质在时间和晶格上的行为,包括Trotter和后Trotter实时动力学、淬火模拟、晶格规范理论与核/高能模型、凝聚态晶格模型,以及在中性原子、囚禁离子和超导硬件上的模拟/可编程模拟器)。数据集规模为213条记录,采用多格式混合结构,包含六种记录类型:概念条目(66条)、开放问答(63条)、多项选择问答(29条)、代码任务(24条)、语料(18条)和指令对(13条)。每条记录都按难度分级:本科生级别(24条)、研究生级别(93条)和研究级别(96条)。每条记录共享一个通用信封结构(包含id、domain、record_type、category、topic等元数据字段),并包含其特定记录类型的字段。数据来源是混合的,初始版本(v0.1)源自专家策划的研究分类法,并集成了策划和大型语言模型合成以进行扩展。数据集实施了严格的质量控制门限,包括:所有代码种子在固定环境中端到端执行并验证数值输出;所有引用的arXiv标识符均被验证;多项选择问答的答案草稿格式和干扰项均经过检查;费米子编码声明通过OpenFermion或Qiskit-Nature进行符号验证;资源估计数据均注明出处;化学惯例被明确定义;并通过主题范围分类器和审核确保内容聚焦于既定领域。该数据集适用于监督微调、评估/基准测试和持续预训练等多种机器学习任务。然而,需注意其局限性:合成记录虽然是经过验证的模型生成内容,但仍可能包含错误,因此不应将其视为权威的科学参考文献,关键事实需对照原始资料进行核实。
Neura Parse — Quantum Simulation in Chemistry and Materials: Encoding, VQE/QPE and Dynamics Dataset is a professional dataset focusing on the application of quantum computing in chemistry and materials science simulation. This dataset aims to support the research and development of quantum computing-aware artificial intelligence systems. The dataset covers core topics of quantum-simulated matter in depth, including: Electronic structure and fermion-to-qubit encoding (converting chemical/material problems into qubit Hamiltonians, involving second- and first-quantized electronic structure Hamiltonians, classical PySCF preprocessing, various fermion-to-qubit encoding schemes such as Jordan-Wigner, parity, Bravyi-Kitaev, etc., as well as qubit reduction based on Z2 symmetry); Hamiltonian decomposition and fault-tolerant resource estimation (compressing two-electron tensors and their impacts on the 1-norm of Hamiltonians and block encoding costs via single/two/tensor hypercontraction decomposition, end-to-end fault-tolerant resource estimation for cases including FeMoco, catalysis, cathodes, etc., and objective evaluation of classical competitors that quantum methods must surpass); Ground-state and excited-state algorithms (algorithms for extracting eigenstates and properties, such as VQE variants and chemical ansätze, quantum phase estimation, quantum subspace/Krylov and imaginary-time methods, measurement grouping and sampling budgets, as well as excited-state, Green's function, response, and finite-temperature methods); Dynamics, condensed matter models and simulation (simulating the behavior of quantum matter in time and lattice, including Trotter and post-Trotter real-time dynamics, quenching simulations, lattice gauge theory and nuclear/high-energy models, condensed matter lattice models, and simulation/programmable simulators on neutral atom, trapped ion, and superconducting hardware). The dataset contains 213 records in total, with a mixed multi-format structure, including six record types: Concept entries (66), Open-ended questions (63), Multiple-choice questions (29), Coding tasks (24), Corpora (18), and Instruction pairs (13). Each record is graded by difficulty: Undergraduate level (24), Graduate level (93), and Research level (96). Each record shares a universal envelope structure (including metadata fields such as id, domain, record_type, category, topic, etc.), and contains fields specific to its record type. The data sources are mixed; the initial version (v0.1) is derived from expert-curated research taxonomies, and integrated with curated content and large language model synthesis for expansion. The dataset implements strict quality control thresholds, including: All code seeds are executed end-to-end in a fixed environment and their numerical outputs are verified; All cited arXiv identifiers are validated; The draft answer formats and distractors of multiple-choice questions are checked; Fermion encoding claims are symbolically verified via OpenFermion or Qiskit-Nature; Resource estimation data are all cited with their sources; Chemical conventions are clearly defined; and content is ensured to focus on the established domain via topic scope classifiers and audits. This dataset is applicable to various machine learning tasks such as supervised fine-tuning, evaluation/benchmarking, and continued pre-training. However, its limitations should be noted: Although synthetic records are generated by validated models, they may still contain errors, so they should not be regarded as authoritative scientific references, and key facts need to be verified against original sources.
数据集详情总结
基本信息
- 数据集名称: Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
- 领域: 量子模拟-化学-材料
- 语言: 英语
- 记录总数: 263条
- 记录类型: code, concept, corpus, instruction, qa_mcq, qa_open
- 许可证: CC-BY-4.0
- 版本: 0.7.0
数据构成
按记录类型分布
| 记录类型 | 数量 |
|---|---|
| qa_open (开放式问答) | 86 |
| concept (概念) | 75 |
| qa_mcq (多选题问答) | 35 |
| code (代码) | 28 |
| corpus (语料) | 23 |
| instruction (指令) | 16 |
| 总计 | 263 |
按难度分布
| 难度 | 数量 |
|---|---|
| 本科 | 24 |
| 研究生 | 103 |
| 研究 | 136 |
主题分类
-
电子结构及费米子-量子比特编码(5个主题)
- 化学/材料问题转化为量子比特哈密顿量
- 一阶和二阶量子化电子结构哈密顿量
- PySCF预处理(积分、基组、活性空间、嵌入)
- 费米子-量子比特编码(Jordan-Wigner、parity、Bravyi-Kitaev、ternary-tree、locality-preserving)
- Z2对称性量子比特缩减
-
哈密顿量分解与容错资源估计(3个主题)
- 双电子张量压缩(单重/双重/张量超收缩分解)
- 端到端容错资源估计(FeMoco、催化、阴极材料)
- 经典竞争对手的诚实评估(CCSD(T)、DMRG、QMC、张量网络)
-
基态与激发态算法(4个主题)
- VQE变体与化学ansatze(UCCSD、k-UpCCGSD、hardware-efficient、ADAPT)
- 量子相位估计
- 量子子空间/Krylov与虚时间方法
- 测量分组与射击预算、激发态/格林函数/响应/有限温度方法
-
动力学、凝聚态模型与模拟仿真(3个主题)
- Trotter与后Trotter实时动力学
- 淬火模拟
- 格点规范理论、凝聚态模型(Fermi-Hubbard、自旋格点)、模拟/可编程模拟器
数据模式
每条数据包含统一信封字段:id, domain, record_type, category, topic, subtopics, difficulty, language, source, source_url, license, tags, provenance, quality, metadata,以及记录类型特有的字段。
数据来源与方法
- 混合来源:由专家策划的研究分类体系生成(方法=curated)
- 策划与LLM合成结合用于规模化
- 每条记录携带 provenance 对象(方法、生成器、流水线版本)和可选的 quality 对象(事实性/清晰度评分)
质量保证
- 每个代码种子在固定环境中端到端执行,输出结果与参考值误差< 1 mHa
- 所有引用的arXiv ID均经过验证
- 多选题有四个选项、一个正确答案和一行理由说明
- 费米子编码声明使用OpenFermion或Qiskit-Nature符号检查
- 资源估计数据均附有具体论文和年份来源
预期用途与限制
- 用于量子计算感知AI系统的研发
- 合成记录虽经验证但可能包含错误
- 不应将此数据集视为权威科学参考,关键事实请查证原始出处




