Au38Q MBTR-K3
收藏资源简介:
<strong>Purpose</strong><br> The purpose of Au38Q MBTR-K3 is to test the scalability of a machine learning regression model when the number of observations and the number of features change. <strong>Background</strong><br> The Au38Q MBTR-K3 was created from a trajectory file regarding the density functional theory simulation of Au38Q hybrid nanoparticle performed by Juarez-Mosqueda et al. in their paper <em>Ab initio molecular dynamics studies of Au38(SR)24 isomers under heating</em> using the MBTR descriptor by Himanen et al. as presented in paper <em>DScribe: Library of descriptors for machine learning in materials science</em>. The MBTR was used with the default parameters for K=3 (angles between atoms) presented at the website of Dscribe version 0.4.0. The dataset was first used to probe the properties of Minimal Learning Machine in paper <em>Do Randomized Algorithms Improve the Efficiency of<br> Minimal Learning Machine?</em> by Linja et al. <strong>Description</strong><br> The dataset contains nine variants of the same idea. In each, an observation refers to a MBTR description of the structural angles of the Au38Q hybrid nanoparticle of a single timestep in a DFT simulation and the potential energy of the said nanoparticle at the timestep. The input space is the MBTR description and the output space is the potential energy. Features refer to the output of the MBTR descriptor, here used as the input. We used three different numbers of observations and three different numbers of descriptor accuracies. Regarding the the number of observations, we used <em>RS-maximin</em> to find out the most different observations available and used the first 4000 and first 8000 as the selections in 4k and 8k variants. Regarding the number of features, we used different descriptor accuracy values [2,10,100] that produced descriptors of lengths [80,400,4000]. This allowed the number of features to represent the data description resolution. Downsampling of the number of features from 4000 to lower numbers was not used. Further details are presented in paper <em>Do Randomized Algorithms Improve the Efficiency of<br> Minimal Learning Machine?</em> by Linja et al.
**研究目的** Au38Q MBTR-K3数据集旨在测试机器学习回归模型在观测样本数量与特征维度发生变化时的可扩展性。 **研究背景** 本数据集源自Juarez-Mosqueda等人发表于论文*Ab initio molecular dynamics studies of Au38(SR)24 isomers under heating*中针对Au38Q杂化纳米粒子开展的密度泛函理论(Density Functional Theory, DFT)模拟轨迹文件,并采用Himanen等人于论文*DScribe: Library of descriptors for machine learning in materials science*中提出的多体张量表示描述符(Many-body Tensor Representation, MBTR)生成。本次实验采用DScribe 0.4.0版本官网提供的K=3(原子间夹角)默认参数。该数据集最初由Linja等人在论文*Do Randomized Algorithms Improve the Efficiency of Minimal Learning Machine?*中用于探究极小学习机(Minimal Learning Machine)的性能特性。 **数据集描述** 本数据集包含9种基于同一核心思路的变体。每个观测样本对应一次DFT模拟时间步下Au38Q杂化纳米粒子的MBTR结构夹角描述,以及该时间步下纳米粒子的势能。其中输入空间为MBTR描述符输出,输出空间为势能值。特征即本次实验中作为输入的MBTR描述符结果。我们设置了三种不同的观测样本数量,以及三种不同的描述符精度等级。针对观测样本数量,我们采用*RS-maximin*算法筛选出差异最大的样本,并分别选取前4000条与前8000条样本作为4k与8k变体的数据集。针对特征维度,我们使用了[2,10,100]三组不同的描述符精度参数,对应生成的描述符长度分别为[80,400,4000],以此通过特征数量体现数据描述的分辨率。本次实验未对特征数量进行从4000向下的降采样操作。更多细节可参考Linja等人的论文*Do Randomized Algorithms Improve the Efficiency of Minimal Learning Machine?*。



