Information-entropy-driven generation of material-agnostic datasets for machine-learning interatomic potentials
收藏资源简介:
This dataset contains the following information for the 8 elements from the paper, namely Be, C, Al, Sb, Te, W, Re and Os: About 30k structures (atomic positions, cell size) generated using the entropy-maximization method described in the paper. These are stored in the standard ASE atoms format, along with their respective DFT energies and forces and made available as pandas dataframes. The subset of structures used for training Atomic Cluster Expansion (ACE) models, i.e., the precise training datasets for all the above elements, and the corresponding test sets. The input files used to train potentials using the open-source Pacemaker code for all the elements. These are text files specifying the parameters used to train the models (basis set size, body-order of the expansion). The resulting trained potentials. Details of the DFT calculations and ACE parameterizations are provided in the paper.
本数据集包含论文中涉及的8种元素(即Be、C、Al、Sb、Te、W、Re和Os)的如下相关信息: 约3万组结构(包含原子位置与晶胞尺寸)通过论文所述的最大熵方法生成。上述结构以标准ASE原子(ASE atoms)格式存储,同时附带各自的密度泛函理论(DFT)能量与原子受力数据,并以pandas数据框形式对外发布。 包含用于训练原子团簇展开(Atomic Cluster Expansion, ACE)模型的结构子集,即上述所有元素的精准训练数据集与对应的测试集。 包含采用开源Pacemaker代码为所有元素训练势函数所需的输入文件。此类文本文件会指定模型训练所用的各项参数,包括基组规模与展开体阶数。 包含最终训练得到的势函数。 DFT计算细节与ACE参数化方案均可在论文中查阅。



