E3-miu-GNN
收藏资源简介:
Neo多元素混合粒度原子数据集是一个专为原子机器学习、表示学习和可重复基准测试设计的大规模规范训练语料库。它是为三层混合粒度E(3)等变图神经网络(GNN)量身定制的,独特地融合了复杂的周期性材料、分子电响应数据、周期性密度泛函微扰理论(DFPT)响应以及共线/非共线磁性监督数据,且所有物理标签均源自真实计算,未进行人工填补。该数据集内容涵盖广泛的物理性质,包括能量、原子力、应力、电荷(赫什菲尔德电荷)、原子偶极矩、分子偶极矩、极化率、原子极化率、C6系数、玻恩有效电荷(BEC)、自旋、局域磁矩和有效自旋场。数据来源于多个权威计算材料学数据库,包括MPtrj(Materials Project轨迹)、JARVIS-DFT、QM7-X、SO3LR、SCFNN以及DeepSPIN等,并经过统一的处理流程和严格的泄漏安全分组划分(确保同一组的数据不会出现在训练、验证和测试的不同集合中)。数据集以分层级(Tier)的HDF5文件形式发布,提供了从快速开发到大规模训练的不同规模版本:Tiny(5,575个结构,20.31 MiB)和Small(15,221个结构,50.31 MiB)用于快速原型开发和中等规模实验的便携子集,完整保留了所有13个数据源和14种目标标签类型;Standard/Mixed(46,414个结构,135.11 MB)是标准的便携语料库,包含37,192个训练、4,541个验证和4,681个测试结构,所有结构均为多元素体系;Large(613,267个结构)是富含分子动力学轨迹数据的大规模语料库,侧重于力/势能面覆盖;Plus(25,819,271个结构,40.63 GB)和Max(101,283,549个结构,134.39 GB)包含海量OMat24数据与完整Large语料库嵌入的自包含大规模课程训练层级。数据集采用e3mu-hdf5-v1模式(Plus/Max为e3mu-composite-hdf5-v1)组织,清晰分离存储几何信息(原子种类、位置、晶胞)、密集的物理标签数组、标识标签有效性的掩码以及丰富的元数据(来源、方法、系统、分组、划分等)。它适用于开发和评估用于预测材料与分子多种物理性质的等变图神经网络,特别是那些需要同时处理周期性边界条件和分子体系、并联合预测能量、力、响应性质和磁性的模型。数据集附带详细的来源说明、处理流程、数据模式文档和多个针对不同训练场景的预设配置文件。
The Neo multi-element mixed-granularity atomic dataset is a large-scale canonical training corpus designed for atomic machine learning, representation learning, and reproducible benchmarking. It is tailored for three-layer mixed-granularity E(3)-equivariant graph neural networks (GNNs), uniquely integrating complex periodic materials, molecular electric response data, periodic density functional perturbation theory (DFPT) responses, and collinear/non-collinear magnetic supervision data, with all physical labels derived from real calculations and no artificial filling. The dataset covers a wide range of physical properties, including energy, atomic forces, stress, charge (Hirshfeld charge), atomic dipole moments, molecular dipole moments, polarizability, atomic polarizability, C6 coefficients, Born effective charges (BEC), spin, local magnetic moments, and effective spin fields. Data is sourced from multiple authoritative computational materials science databases, including MPtrj (Materials Project trajectories), JARVIS-DFT, QM7-X, SO3LR, SCFNN, and DeepSPIN, and undergoes a unified processing pipeline with strict leakage-safe grouping (ensuring data from the same group does not appear in different sets of training, validation, and testing). The dataset is released in tiered HDF5 files, providing different scale versions from rapid development to large-scale training: Tiny (5,575 structures, 20.31 MiB) and Small (15,221 structures, 50.31 MiB) are portable subsets for quick prototyping and medium-scale experiments, retaining all 13 data sources and 14 target label types; Standard/Mixed (46,414 structures, 135.11 MB) is a standard portable corpus, containing 37,192 training, 4,541 validation, and 4,681 test structures, all of which are multi-element systems; Large (613,267 structures) is a large-scale corpus rich in molecular dynamics trajectory data, focusing on force/potential energy surface coverage; Plus (25,819,271 structures, 40.63 GB) and Max (101,283,549 structures, 134.39 GB) are self-contained large-scale curriculum training tiers that include massive OMat24 data embedded with the complete Large corpus. The dataset is organized in the e3mu-hdf5-v1 mode (e3mu-composite-hdf5-v1 for Plus/Max), clearly separating storage of geometric information (atom types, positions, unit cells), dense physical label arrays, masks indicating label validity, and rich metadata (sources, methods, systems, groups, splits, etc.). It is suitable for developing and evaluating equivariant graph neural networks for predicting various physical properties of materials and molecules, particularly those requiring simultaneous handling of periodic boundary conditions and molecular systems, and joint prediction of energy, forces, response properties, and magnetism. The dataset comes with detailed source descriptions, processing workflows, data schema documentation, and multiple preset configuration files for different training scenarios.
Neo Multi-Element Mixed-Granularity Atomistic Dataset 概述
数据集基本信息
- 名称: Neo Multi-Element Mixed-Granularity Atomistic Dataset
- 许可协议: other(需参考 LICENSES_AND_ATTRIBUTION.md 了解具体义务)
- 规模: 100K < n < 1M 结构
- 主要用途: 原子机器学习、表示学习、可重复基准测试
- 支持模型: 三层混合粒度 E(3)-mu GNN
数据集构成
Neo 数据集结合了周期性材料、分子电响应数据、周期性 DFPT 响应以及共线/非共线磁性监督数据,不虚构缺失的物理标签。
主要文件及层级
| 文件 | 结构数 | 大小 | 主要标签 |
|---|---|---|---|
| neo_tiny_l1_l2_l3.h5 | 5,575 | 20.31 MiB | 所有标准目标类型 |
| neo_small_l1_l2_l3.h5 | 15,221 | 50.31 MiB | 所有标准目标类型 |
| neo_mixed_l1_l2_l3.h5 | 46,414 | 135.11 MB | 能量、力、电荷/响应、BEC、自旋、矩、有效场 |
| neo_large_l1_l2_l3.h5 | 613,267 | - | 能量、力、响应、BEC、自旋和矩 |
| neo_plus_l1_l2_l3.h5 | 25,819,271 | 40.63 GB | OMat24 能量/力/应力 + 所有 Large L1/L2/L3 标签 |
| neo_max_l1_l2_l3.h5 | 101,283,549 | 134.39 GB | 完整 L1 基础 + 所有 Large L1/L2/L3 标签 |
便携式层级(Tiny / Small / Standard)
层级之间为精确嵌套关系:每个 Tiny 样本 ID 出现在 Small 中,每个 Small 样本 ID 出现在 Standard 中。
| 层级 | 结构数 | 原子数 | 组数 | 大小 | 训练/验证/测试 |
|---|---|---|---|---|---|
| Tiny | 5,575 | 371,803 | 4,991 | 20.314 MiB | 4,361 / 610 / 604 |
| Small | 15,221 | 964,550 | 10,973 | 50.312 MiB | 12,094 / 1,552 / 1,575 |
| Standard | 46,414 | 2,316,736 | 28,445 | 135.109 MB | 37,192 / 4,541 / 4,681 |
所有 85 种标准元素在训练、验证和测试集中均出现,保留所有 13 个标准来源和所有 14 种活跃 HDF5 目标类型。
默认混合语料库组成(Standard)
- 化学复杂性分布: 二元 9,794 结构、三元 11,807 结构、四元及以上 24,813 结构
- 原子数分布: 1-24 原子 20,449 结构、25-64 原子 15,976 结构、65+ 原子 9,989 结构
- 固定划分: 37,192 训练、4,541 验证、4,681 测试结构(组不跨划分边界)
大型语料库组成(Large)
- 613,267 结构、17,760,024 原子、152,519 个防泄漏组、89 种化学元素
- 划分: 492,759 训练、59,813 验证、60,695 测试
- 化学复杂性: 单质 4 结构、二元 81,257 结构、三元 244,104 结构、四元及以上 287,902 结构
- 标签覆盖: 能量/力 505,736 结构、自旋/磁矩 73,029 结构、有效场 100 结构、场/总电荷 107,431 结构、电荷/原子偶极 101,993 结构、偶极 106,769 结构、极化率/C6 43,430 结构、Born 有效电荷 662 结构
标签覆盖范围
| 标签 | 结构数 | 单位 |
|---|---|---|
| 能量/力 | 22,761 | eV / eV per angstrom |
| 自旋/磁矩 | 12,100 | dimensionless / Bohr magneton |
| 有效自旋场 | 100 | eV per spin |
| 电荷/原子偶极 | 18,130 | e / e angstrom |
| 偶极 | 22,891 | e angstrom |
| 极化率/原子极化率/C6 | 4,060 | angstrom³ / eV·angstrom⁶ |
| Born 有效电荷 | 662 | e |
缺失标签: J_effective、Di、DMI 目标(其掩码为零,且 DMI 架构开关必须禁用)
数据来源
| 来源 | 许可 | 描述 |
|---|---|---|
| MPtrj | MIT | 官方完整存档,DOI 10.6084/m9.figshare.23713842.v2 |
| JARVIS-DFT | CC BY 4.0 | 官方 2025 3D 存档,DOI 10.6084/m9.figshare.6815699.v11 |
| QM7-X | CC BY 4.0 | DOI 10.5281/zenodo.4288677 |
| SO3LR | CC BY 4.0 | DOI 10.5281/zenodo.14779793 |
| SCFNN | CC BY 4.0 | DOI 10.5281/zenodo.5521328 |
| 局部 BEC | 待审查 | H2O、MAPbI3、二聚体 DFPT 结构,版权状态待定 |
| DeepSPIN | GPL-3.0 | Git 提交 526ade353906f21cf4d8cb32db3d53ce83e1ed53 |
HDF5 模式
所有规范文件使用 e3mu-hdf5-v1 模式:
structures/: 存储几何信息(atom_ptr、atomic_numbers、positions、cell、pbc)labels/: 存储密集类型化标签数组masks/: 明确标注哪些结构拥有物理标签metadata/: 存储来源、方法、系统、组、划分、样本、父级、域、能量参考和出处标识符
Plus 和 Max 使用 e3mu-composite-hdf5-v1,额外包含 selection、sources/omat24、sources/neo_large 和 atomic_reference 组。
预设配置
- mixed_qeq_spin_film.json: 推荐默认联合架构,启用 QEq、自旋、FiLM,禁用 PME/D4/DMI
- large_qeq_spin_film.json: 613,267 结构轨迹丰富语料库的 200 轮配置
- large_smoke_mps.json: 两轮 MPS 功能检查
- portable_tier_smoke_mps.json: 两轮全物理 MPS 检查
- periodic_pme_experimental.json: 周期性静电实验
- molecular_d4_qm7x.json: 分子仅 D4 实验
- smoke_cpu.json / smoke_mps.json: 可重复短功能检查
重要说明
- 数据仅供研究用途,非经验验证结构的认证来源
- 用户必须独立验证单位、参考能量、收敛性、化学适用性、不确定性和下游预测
- 数据集和元数据按“原样”提供,不提供任何保证
- OMat24 衍生组件(Plus 和 Max)以 CC BY 4.0 条款分发
- MLFF_and_BEC 补充存档的单独分发权审查尚在进行中
- 大规模预训练正在运行,验证的预训练检查点将在后续项目版本中发布




