colabfit/Massive_Atomic_Diversity_MAD-1.5_r2SCAN_Train
收藏资源简介:
MAD-1.5(大规模原子多样性版本1.5)数据集的训练分割,这是一个高度策划的集合,专为训练跨整个周期表的广泛适用的原子机器学习模型而设计。MAD-1.5扩展了原始MAD数据集,通过有针对性的富集策略覆盖102种化学元素(所有半衰期超过一天的同位素)。所有216,803个结构均使用FHI-aims(版本250806)中的r2SCAN meta-GGA功能,通过标准化全电子DFT工作流程计算,具有紧基组、8 Å^-1 k点密度、0.05 eV的高斯展宽,以及1e-6 eV(能量)、1e-4 eV/Å(力)和1e-5 e*a0^-3(电子密度)的SCF收敛阈值。数据集涵盖分子(单体、二聚体、三聚体、分子晶体)、块体晶体、表面、纳米团簇和低维结构,分为14个子集。质量通过两步异常值去除确保:启发式过滤力大于100 eV/Å的结构,然后基于LLPR不确定性进行过滤。训练分割(约占清理数据的83%)包括所有单体、二聚体和三聚体,以锚定低体序相互作用。在模型训练期间,使用了一个伴随的PBE功能数据集(Massive_Atomic_Diversity_MAD-1.5_PBE),并具有单独的预测头。
Training split of the MAD-1.5 (Massive Atomic Diversity version 1.5) dataset, a highly curated collection designed for training broadly applicable atomistic machine-learning models across the full periodic table. MAD-1.5 extends the original MAD dataset with targeted enrichment strategies covering 102 chemical elements (all isotopes with half-life above one day). All 216,803 structures are computed with a single standardized all-electron DFT workflow using the r2SCAN meta-GGA functional in FHI-aims (version 250806), with tight basis sets, 8 Angstrom^-1 k-point density, Gaussian smearing of 0.05 eV, and SCF convergence thresholds of 1e-6 eV (energy), 1e-4 eV/Angstrom (forces), and 1e-5 e*a0^-3 (electron density). The dataset spans molecules (monomers, dimers, trimers, molecular crystals), bulk crystals, surfaces, nanoclusters, and low-dimensional structures organized into 14 subsets. Quality is ensured by two-step outlier removal: heuristic filtering of structures with forces >100 eV/Angstrom, followed by LLPR uncertainty-based filtering. The training split (~83% of cleaned data) includes all monomers, dimers, and trimers to anchor low-body-order interactions. A companion PBE-functional dataset (Massive_Atomic_Diversity_MAD-1.5_PBE) was used during model training with separate prediction heads.



